Commit graph

2825 commits

Author SHA1 Message Date
Peter Steinberger
8a981770fe
fix(tooling): verify Darwin zombie groups after EPERM
Darwin killpg excludes zombies and returns EPERM when no signalable group
member remains. Strict normal-exit cleanup mistook this terminal state for
a surviving process group, even when later drainage observed termination.

Reuse the existing zombie census for Darwin, including BSD state flags,
and reconcile a group reaped during the census with a fresh ESRCH probe.
Both observation and termination require positive completion evidence;
live, mixed, and uninspectable groups retain their cleanup failures.
Keep snapshot work inside the existing escalation and drainage deadlines.

Native same-uid Z and ZN groups reproduce EPERM for signal 0 and SIGKILL.
The original owner fails four regression cases; the repair passes all 13.
Shared owner/output tests pass (146 passed, one platform skip). The full
mac-elevation-host shard passes all 123 cases in 383.63s; final metadata
replay passes all 13 selected cases. Independent review is scoped-clean.
Production delta: +44 lines for bounded Darwin termination evidence.

Test cost: node scripts/run-vitest.mjs test/scripts/managed-child-process.termination.test.ts --maxWorkers=1: 5.83s wall; node scripts/run-vitest.mjs test/scripts/managed-child-process.tree.test.ts --maxWorkers=1: 5.05s wall.
2026-09-30 18:04:04 -07:00
Peter Steinberger
dac3c15a0d
test(core,plugins): remove low-value tests (batch d119) (#162202)
* test(pr): deslop t0431 tests

* test(plugins): deslop t0430 tests

* test(mattermost): deslop t0415 tests

* test(telegram): deslop t0427 tests

* test(process): deslop t0424 tests

* test(state): deslop t0434 tests

* test(voice-call): deslop t0437 tests

* test(telegram): deslop t0428 tests

* test(gateway): deslop t0443 tests

* test(agents): deslop t0436 tests

* test: preserve lifecycle and authorization contracts in d119
2026-09-30 23:33:21 +00:00
Peter Steinberger
bfdb432570
fix(cli): release plugin resources when help finishes (#160181)
* fix(cli): release plugin resources when help finishes

Uncached CLI metadata loads now use the existing inspection acquisition when an executable invocation owns them. The source-plugin regression verifies that invocation release joins module disposal. Caller-owned programs retain their existing lifetime.

* fix(plugins): release retired registry preparation scope

Create the aggregate retirement observer outside the registry preparation
scope so it retains only child waiters. Preserve per-observation options,
cleanup ordering, failure identity, and deferred consumer results.

The unchanged cold registry-retention regression failed on the CI Bun
runtime before this change and passes after it. The original seven-file
shard passes all 172 cases, and Node retirement coverage passes 20 cases.

* fix(ci): carry the chat attachment lint repair

Carry the exact two-file fix from 2e95cdba84
(#160159). Move the unchanged image decoder into its existing helper
to restore the 700-line limit without changing attachment behavior.

Targeted lint and all 21 attachment tests pass on this candidate.
Independent review found no actionable defects.

* test: retain complete lifecycle helpers in lease fixtures

* test: isolate local command fixtures and join recovery cleanup

* test(ci): preserve CLI runtime ownership and harness staging

Keep the moved CLI command fixture in its runtime prerequisite group and
preserve complete agent coverage across the core and CLI process owners.

Carry the test-only trusted-harness fixture fix from #160024.

Co-authored-by: Vincent Koc <vincentkoc@ieee.org>

* fix(tests): scope Vitest preload to test workers

Apply the jsdom adapter only at Vitest worker entry points while preserving
package-local runtime resolution. Ordinary Workers may inherit the preload
without having a Vitest dependency. Exercise default native inheritance and
await worker termination.

Copy the static-check evidence helper into lint fixtures so the real oxlint
entry point can load its current import closure.

Validation: reproduced both original failure causes before repair; all 76
affected tests pass (351.44 seconds in Vitest), as do canonical changed-file
checks and independent review.

* test: isolate public runtime surface planning fixtures

Verify public runtime entries and packaged assets with a synthetic publishable plugin after the Diffs forwarding APIs were retired. Preserve the existing planner assertions without depending on one bundled plugin layout.

* fix(ci): carry shared lint repairs into CLI preparation

* test(tooling): repair extracted PR and npm fixtures

* test: retain gateway cleanup and isolate Windows temp paths

Close the prewarm gateway when client connection fails and retain cleanup
errors during orphaned recovery. Keep Windows-under-Bun coverage on
host-owned temporary directories, matching the fixture fix in #160985.

Validation: 13 focused fixture tests and canonical changed-file checks.

* fix(ci): keep extracted manifest within script contracts

Resolve inherited manifest lint errors without changing job selection. Move its closed true/1 flag parsing to the dependency-free argument owner and include that helper in trusted fixture copies. Validation passed 718 tests with eight existing skips, canonical changed checks, fixture lint, and independent review. The 74-case argument helper file took 1.79 seconds at one worker with warm inputs.

* test(state): apply upstream snapshot custody fixture isolation

Reuse the fixture repair from #161598 (af16b80a22). Keep the real exit callback and every custody assertion while isolating the synthetic cleanup owner from preceding files.

* test(pr): drain fake GitHub streams before exit

---------

Co-authored-by: Vincent Koc <vincentkoc@ieee.org>
2026-09-30 15:28:57 -07:00
Peter Steinberger
eb65a6c5f1
ci: overlap iOS simulator preparation with builds
Select the simulator before compilation, then boot and slim that exact
simulator alongside the app build. Join preparation before XCTest, retain
its failure log, and bound a hung join without changing test selection or
build settings.

Validate both focused groups on cold and warm native runs (139 tests each),
plus preparation failure and timeout controls. Linux proof includes full
test types, the whole tooling config, changed checks, source contracts,
and both preflight harness shapes for PR and scheduled-main inputs.
2026-09-30 14:48:59 -07:00
Peter Steinberger
59b8097d75
test(core,plugins): remove low-value tests (batch d117) (#162104)
* test(cli): deslop t0406 tests

* test(agents): deslop t0407 tests

* test(onboarding): deslop t0404 tests

* test(infra): deslop t0411 tests

* test(openai): deslop t0412 tests

* test(telegram): deslop t0418 tests

* test(tool-call-repair): deslop t0413 tests

* test(plugins): deslop t0419 tests

* test(codex): deslop t0408 tests

* test(agents): deslop t0417 tests

* test: retain distinct lifecycle and transport contracts in d117

* test(infra): apply update opt-out in restored fixture
2026-09-30 21:48:00 +00:00
Peter Steinberger
be25a24f10
test(core,plugins): remove low-value tests (batch d118) (#162138)
* test(scripts): deslop t0421 tests

* test(node-host): deslop t0420 tests

Remove repeated policy-drift inputs, generic route checks, and replays of
pager, wrapper-depth, and approval-plan owner tests. Consolidate durable
approval fixtures and lifecycle cleanup; infer mock result types instead
of casting partial copies of production contracts. Preserve Vitest routing.

The 50% campaign target is not reached. Retention ledger below names each
remaining case (table rows expanded), using the requested rules:
1 = unique covered execution path; 2 = historical regression; 3 = distinct
security boundary. No unrelated behaviors were combined to reduce counts.

- Native cwd refusal: 2, d88e10a752, preserve preparation-refusal classification.
- Lost companion response: 3, disabled fallback must never replay locally.
- Already-cancelled command: 1, entry cancellation return.
- Cancelled local completion: 1, post-run cancellation return and signal forwarding.
- Pending Mac cancellation: 1, post-bridge cancellation return without replay/publication.
- Medium-risk auto review: 2, 3b26947669, ordinary audit-related argv must reach review and execution.
- Auto review without a plan: 3, unbound commands cannot obtain model authority.
- Login-shell startup: 3, startup effects cannot bypass review through a forged plan.
- Reviewer asks: 1, defer-to-human switch branch.
- Reviewer denies: 3, explicit model refusal cannot execute.
- Approved env wrapper argv: 3, preserve approved positional execution semantics locally and over the bridge.
- Transparent versus semantic env wrappers: 3, only transparent carriers may unwrap into an allowlisted executable.
- Nested safe-bin chains: 3, bind rewritten shell payloads to canonical safe-bin paths.
- PowerShell wrappers: 1, non-POSIX transport must bypass POSIX rewriting.
- PATH-token allowlist execution: 3, launch the canonical approved executable rather than a mutable PATH alias.
- Live policy revocation at commit: 3, recheck authority after awaited persistence before real launch.
- Live policy revocation at callback: 3, the launch adapter must recheck immediately before spawning.
- Live policy unchanged at callback: 3, positive control proves authorized real execution succeeds.
- Unbindable auto executable identity: 3, model review cannot authorize an unbound dispatch chain.
- Human executable identity drift at commit: 3, changed executable resolution invalidates the approved command.
- Script drift after commit: 3, recheck mutable operands after the persistence await.
- Runtime script binding at dispatch: 3, reject changed operands and omitted required bindings.
- Relative skill-bin invocation: 3, skill trust cannot authorize an arbitrary relative executable.
- Unsafe environment inputs: 3, block dangerous overrides, argv assignments, and invalid env keys.
- Shell executable env allowlist: 3, filter shell environments even without an inline payload.
- Revoked rule during allow-always commit: 3, do not resurrect revoked authority through grant persistence.
- Unprompted ask tightening during commit: 3, current-policy authority must not survive a newly required prompt.
- Auto-review ask tightening during commit: 3, forwarded model authority remains bound to its prepared policy snapshot.
- Exact-plan forwarded strict inline review: 3, authenticated one-shot authority remains scoped to its bound plan.
- Unavailable screen recording: 3, deny before execution or durable grant mutation.
- Tightened timeout fallback policy: 3, reevaluate fallback authority at commit.
- Fallback without canonical plan: 3, reject unbound fallback provenance.
- Explicit approval without canonical plan: 3, reject unbound explicit authority.
- Delayed approval without policy snapshot: 3, require the prepared policy identity.
- Allowlist revocation after prepare: 3, reject stale authority before committing.
- Mixed fallback and explicit provenance: 3, refuse contradictory authority sources.
- Durable allowlist timeout fallback: 3, positive control for an exact durable grant.
- Revoked durable timeout source: 3, source downgrade removes fallback authority.
- Mac fallback bridge: 3, forward fallback provenance without inventing explicit approval.
- Fallback versus strict inline review: 3, timeout fallback cannot satisfy mandatory review.
- Unknown approval provenance: 3, reject unsupported authority markers.
- Benign awk durable grant: 3, a file-based grant must not reopen inline programs.
- Make inline grant: 3, explicitly approved runtime payloads remain one-shot.
- Windows cmd transport: 3, unwrap transport before deciding shell approval requirements.
- Windows durable trust downgrade: 3, revalidate the exact durable source at commit.
- Durable cwd replacement: 3, directory identity remains bound through execution.
- Safe builtin with redundant exact grant: 1, builtin authorization must survive revocation of an unused grant.
- Socket response loss: 2, adf5c67382, real completed execution must never replay after its reply is lost.

Measurements and validation:
- Executed tests: 90 -> 48, down 46.67%.
- Shard test LOC: 3431 -> 2045, down 40.40%.
- Covered statements: 357 -> 354 (-0.84%).
- Covered branches: 438 -> 435 (-0.68%). Functions: 34 -> 34.
- Only invoke-system-run-plan.ts drops coverage: 3 statements, 3 branches.
  These are removed preparation-only probes; no source file drops >5 statements.
- Both Vitest coverage groups succeeded; no groups were excluded.
- Blacksmith Testbox: both shard files pass (47 infra + 1 unit tests).
  Final file durations: 2.80s infra; 1.95s unit.
- check:changed passes, including core-test types, boundary checks, all
  dead-export scans, and type-aware lint (0 warnings, 0 errors).
- Local oxfmt, plain oxlint, and git diff --check pass.
- No production or shared test-support changes; no bugs or flakes found.
- Proof workflow: https://github.com/openclaw/openclaw/actions/runs/36738903333

* test(ai): deslop t0414 tests

* test(commands): deslop t0423 tests

* test(google): deslop t0416 tests

* test(board): deslop t0409 tests

* test(infra): deslop t0410 tests

* test(gateway): deslop t0429 tests

* test(doctor): deslop t0422 tests

* test(signal): deslop t0426 tests

* test: retain distinct regressions in batch d118
2026-09-30 21:36:05 +00:00
Peter Steinberger
a948e56619
fix(test): adapt Tool Search recipe to published baseline
Use structured Tool Search config from 2026.9.7 while retaining the legacy code-mode migration specimen for older baselines. Read the existing coverage receipt version for baseline assertions; keep post-upgrade cleanup checks strict.
2026-09-30 13:23:51 -07:00
Peter Steinberger
441e664f94
refactor(memory): move Forget planning reads to the worker (#162045)
Use the existing memory retrieval pool for noncreating index selection, counts, and vector inspection. Capture caller placement and matching facts before dispatch, share the existing scrub rules, and return only selected row identities and counts. Keep final lineage checks and write settlement with their existing transaction owner.

Validation: causal host count regression, 150 owning tests, native fetched-row and byte bounds, captured-input and read-error controls, full changed-file gate, and independent review.
2026-09-30 13:11:03 -07:00
Peter Steinberger
404e47a13c
test(core,plugins,ui): remove low-value tests (batch d116) (#162019)
* test(custodian): deslop t0394 tests

* test(media): deslop t0397 tests

* test(imessage): deslop t0396 tests

* test(agents): deslop t0392 tests

* test(plugins): deslop t0383 tests

* test(cli-runner): deslop t0388 tests

* test(config): deslop t0405 tests

* test(ui): deslop t0403 tests

* test(channels): deslop t0400 tests

* test(voice-call): deslop t0398 tests

* test: preserve independent contracts in d116 cleanup
2026-09-30 11:50:25 -07:00
Peter Steinberger
8e047c3ce4
test(core,plugins,ui): remove low-value tests (batch d114) (#161993)
* test(copilot): deslop t0366 tests

* test(config): deslop t0377 tests

* test(plugins): deslop t0368 tests

* test(auth): deslop t0381 tests

* test(ui): deslop t0382 tests

* test(release): deslop t0384 tests

Port the shard while preserving newer SDK acknowledgement coverage and all six strengthened trusted-tooling cleanup scenarios.

Validation: 109 tests passed with no skips on Blacksmith Testbox from a fresh-main proof checkout; oxlint, oxfmt, diff checks, and independent review passed.

* test(agents): deslop t0369 tests

Port 2533d418dcab93da1aec73ee224072e29402ba08 onto the campaign lane, preserving all 25 test blocks changed on the lane and its four test removals. Retain the existing max-lines baseline because preserving that coverage keeps the planner suite above the limit.

Validation: all 60 tests passed on Blacksmith Testbox from a fresh-main proof checkout, with no skips. Scoped oxlint, oxfmt, diff checks, and independent review passed.

* test(pdf): deslop t0374 tests

* test(auto-reply): deslop t0378 tests

* test(gateway): deslop t0371 tests

* test: preserve persistence and release workflow coverage

Retain cold logout reload, persisted plugin-state reopening, and real workflow-input compatibility while keeping the batch test reductions.
2026-09-30 10:34:56 -07:00
Peter Steinberger
96d66fed3e
refactor(memory): move Forget transactions off the host thread (#161959) 2026-09-30 10:31:20 -07:00
Peter Steinberger
cfd219c519
fix(release): skip pending ClawHub publications and surface recovery for failed ones (#161985)
The ClawHub release planner and prepared-artifact resolver read the new public publication-state endpoint (/api/v1/packages/{name}/versions/{version}/publication). Only absent versions are republished; pending ones are skipped, and failed ones are excluded, with the recover command printed to the step summary. A 404, or a 200 without a state field, falls back to the legacy version probe.
2026-09-30 09:35:54 -07:00
Peter Steinberger
dc49c824d3
fix(release): verify beta floor for core npm packages (#161968)
Share filesystem core package discovery with npm bundle preparation. Include every selected core package in beta-floor diagnostics and block postpublish on any core floor failure, while preserving reused and superseded core selectors.
2026-09-30 16:21:44 +00:00
Peter Steinberger
30e80e790d
fix(ci): bound isolated Gateway fixture workers
Keep the Gateway isolated/database-worker cohort at its existing two-worker budget so cold startup has room within unchanged test deadlines. Preserve former eight-worker complete generations as conservative timing floors without relabeling them as current measurements.

Validation: 325 owning tests, selected changed checks, and fresh managed review. Generated job counts and coverage are unchanged. Exact changed CI remains required.
2026-09-30 09:20:07 -07:00
Peter Steinberger
e15fa0b879
test(core,browser): remove low-value tests (batch d105) (#161658)
* test(daemon): deslop t0250 tests

* test(net): deslop t0248 tests

* test(auto-reply): deslop t0231 tests

* test(browser): deslop t0260 tests

* test(mcp): deslop t0255 tests

* test(commands): deslop t0245 tests

* test(auth-profiles): deslop t0253 tests

* test(agents): deslop t0244 tests

* test(doctor): deslop t0268 tests

* test(test): deslop t0252 tests

* test: retain boundary coverage and repair d105 checks

* test: preserve distinct d105 lifecycle and boundary contracts
2026-09-30 08:54:27 -07:00
Peter Steinberger
0b3de8ae20
fix(update): settle service receipts and interrupted rollback (#161851)
* fix(update): settle service receipts and interrupted rollback

Await the existing bound worker phase receipt before native service stop,
then revalidate the original executor/requester. Retain already admitted
compensation through signal settlement, including its later phase writes,
without retaining the unbounded forward operation.

Keep current-core and rollback authority policies distinct. Preserve the
existing ledger kernel, records, schema and recovery semantics.

Related: #161385
Canonical stale-baseline context: #161766

Co-authored-by: Vincent Koc <vincentkoc@ieee.org>

* fix(test): register post-update worker fixture once

Keep its existing worker registration instead of adding a duplicate that
prevents CI compact timing generation. Preserve the strict uniqueness
guard and the three newly required worker routes.

The actual CI merge failed manifest generation and its existing registry
regression; the baseline and one-entry correction passed the same selector.

* test(update): retain real owners in module failure fixtures

Load the mutable-signal and execution-guard owners in the existing VM
fixture so post-update failures reach their original backup-retention
assertions. Keep all19 inner cases and the synthetic service boundary.

The original three-file CI group failed on the unexpected dependency call;
the two real-module additions restore17 outer passes and19 inner passes.

---------

Co-authored-by: Vincent Koc <vincentkoc@ieee.org>
2026-09-30 08:49:22 -07:00
Peter Steinberger
192b90aaae
ci: plan extension types without serial core discovery
Retain complete core graph validation in the existing required parallel boundary row for canonical extension-only changes. Discover all four noncore compiler consumers while preserving full fallback for aliases, missing inputs, mixed changes and incomplete inventories.
2026-09-30 06:45:23 -07:00
Peter Steinberger
cbb17e87e9
ci: stabilize native declaration emission
Keep native declaration visits deterministic so regenerated SDK types preserve valid extension package receipts. Type-check-only workers retain their existing concurrency and all diagnostics and boundary seals remain enforced.
2026-09-30 06:08:43 -07:00
Peter Steinberger
72cca0df98
test(agents, memory, ui, tooling): remove low-value tests (batch d112) (#161835)
* test(transcripts): deslop t0348 tests

* test(memory-core): deslop t0340 tests

* test(agents): deslop t0334 tests

* test(line): deslop t0361 tests

* test(docker-e2e): deslop t0362 tests

* test(scripts): deslop t0338 tests

* test(release): deslop t0359 tests

* test(markdown): deslop t0357 tests

* test(ui): deslop t0360 tests

* test(vitest): deslop t0363 tests

* test(ci): align config owner expectation after test cleanup
2026-09-30 11:59:03 +00:00
Peter Steinberger
b06666e3ad
test(runtime): remove low-value tests (batch d110) (#161714)
* test(auto-reply): deslop t0258 tests

* test(config): deslop t0331 tests

* test(memory-core): deslop t0285 tests

* test(updater): deslop t0320 tests

* test(telegram): deslop t0333 tests

* test(gateway): deslop t0341 tests

* test(agents): deslop t0318 tests

* test(gateway): deslop t0345 tests

* test(gateway): deslop t0328 tests

* test(gateway): deslop t0337 tests

* test: retain unique safety contracts in reduction batch

* fix(ci): keep native compiler import proof on Node
2026-09-30 11:58:18 +00:00
Peter Steinberger
4d63cbedc6
test(core,plugins,ui): remove low-value tests (batch d111) (#161796)
* test(agents): deslop t0342 tests

* test(config): deslop t0347 tests

* test(ui): deslop t0336 tests

* test(gateway): deslop t0335 tests

* test(gateway): deslop t0349 tests

* test(logging): deslop t0315 tests

* test(copilot): deslop t0353 tests

* test(system-agent): deslop t0339 tests

* test(plugin-state): deslop t0344 tests

Consolidate nine files into five while retaining cursor release, schema ownership, rollback, expiry bounds, and runtime authority checks. Testbox: 46 tests pass; covered statements remain 746 and covered branches drop from 395 to 392. Existing script guard and lint failures in ci-build-manifest.mjs remain outside this shard.

* test(plugins): deslop t0346 tests

* test: retain distinct security and preference coverage

* test(plugins): match archive security errors correctly

* test: retain lifecycle and import boundary coverage

* test(gateway): type browser projection cases as named records
2026-09-30 11:44:03 +00:00
Peter Steinberger
bf4b16a81d
refactor(meeting): declare browser meeting plugins through one factory (#161628)
Zoom meetings, Teams meetings and Slack huddles each carried the same
thirteen-file glue layer (entry, CLI, CLI metadata, config, errors, node
host, node invoke policy, runtime facade, probes, setup, Chrome transport,
types), feeding one platform adapter into a dozen separate
MeetingPlatformAdapter helpers. A new MeetingPlatformAdapter member,
defineBrowserMeetingPlugin, composes those helpers from one declaration,
so each plugin now declares only what is genuinely platform-specific. A
fourth browser meeting platform costs one declaration.

The helpers the factory replaces stay exported and are marked deprecated,
since published plugin versions call them. Tests inject Chrome bindings
through the factory's public spec instead of mocking core internals. The
three plugins drop their unused typebox dependency.

Zoom and Teams now require OpenClaw 2026.9.8 (install and plugin API
floors, catalog, docs), like Slack huddles, because they call a member
2026.9.7 lacks. Released plugin packages already require a matching host.

All plugin registrations and 42 generated page-script hashes are
identical to main.

Release note: Zoom meetings and Microsoft Teams meetings plugins require
OpenClaw 2026.9.8 or newer.
2026-09-30 04:36:48 -07:00
Peter Steinberger
0c0faf026c
test(core,plugins,scripts): remove low-value tests (batch d106) (#161661)
* test(scripts): deslop t0269 tests

* test(gateway): deslop t0267 tests

* test(doctor): deslop t0262 tests

* test(auto-reply): deslop t0259 tests

* test(scripts): deslop t0263 tests

* test(skills): deslop t0261 tests

* test(plugins): deslop t0256 tests

* test(cli): deslop t0266 tests

* test(config): deslop t0265 tests

* test(codex): deslop t0277 tests

* test(cron): fix reduced fixture lint and prune stale baseline

* test: remove dangling references to retired fixtures

* test(gateway): preserve distinct media retention coverage
2026-09-30 11:20:15 +00:00
Peter Steinberger
509cb47840
refactor: move workspace journal storage off the Gateway thread (#161539)
* refactor: move workspace journal storage off the Gateway thread

* fix: include workspace journal contract in PR wrappers

* fix: preserve repository authority through journal commits

* test: capture first-signin startup and RPC timing on timeout
2026-09-30 04:16:02 -07:00
Peter Steinberger
c2ac326cb0
build(macos): pin the app runtime to OpenClaw Bun 57fadf566d (#161825)
Moves the bundled runtime from fork release ee83b78b18 to
openclaw-v1.4.3-20260930-57fadf566d-webkit-f20ce77445: the upstream Bun
sync through ba3f27d1d1, the full CI-line compatibility port, the macOS
/dev/fd copy fix, child_process stdio handles, bounded compile-cache exit
persistence, TLS over paused CONNECT transports, and the net/fs/
structuredClone parity fixes.
2026-09-30 04:08:34 -07:00
Peter Steinberger
f6bba09343
ci: pack regular CLI stripes within the existing budget
Include generated regular CLI children in both serial packing paths while preserving the 150-second child and 250-second job limits. Keep runner, worker, preparation and process policies intact.

The focused planner regression and neighboring budget case pass in 33.87 seconds on macOS. A local manifest probe emits 88 instead of 89 rows with the identical 3,884-target inventory; Linux integration proof remains required before landing.
2026-09-30 02:57:50 -07:00
Peter Steinberger
94c99f73e0
ci: defer published upgrade survivor to hourly verification
The intact published-upgrade-survivor runs in every admitted hourly main CI run
and Full Release Verification through its normal_ci child. Preserve the exact
legacy-operator-state scenario with auto-auth when the target declares it.

No PR owner changes select this survivor, including updater, Doctor,
state-migration, or direct survivor changes. The other five Docker seed lanes
keep their existing PR owner selection.

Frozen historical targets retain their supported base fallback; invalid
catalogs fail. No survivor scenario case or assertion is removed.
2026-09-30 02:57:50 -07:00
Peter Steinberger
95ed4ee478
fix(ci): run first-hop self-upgrades in waves and restore them to stable validation (#161774)
Raise first-hop lane weight to two under the existing npm limit of five, admitting at most two concurrent lanes while permitting one to overlap the weight-three survivor. Budget three 3500-second waves plus 20 minutes for the survivor and 10 minutes for setup/artifacts: 205 minutes, rounded to a 210-minute job timeout.

With the six-way contention bounded, return package-update-self-upgrade to the stable profile it was dropped from for 2026.9.7 (#161291, #161257).
2026-09-30 09:37:11 +00:00
Peter Steinberger
ae9ac5f3e8
perf(ci): reuse prepared test routing plans
Carry the exact per-target routing results through the same planner invocation instead of resolving them twice more. Preserve canonical ownership checks, plugin opt-ins, fallback behavior, and selected rows.
2026-09-30 02:15:06 -07:00
Peter Steinberger
093f16d0b5
fix(tooling): avoid SDK surface compiler cleanup failures (#161724) 2026-09-30 02:06:24 -07:00
Peter Steinberger
6c5feec599
fix(macos): compile every Swift module for the app's deployment target (#161754)
SwiftPM builds each dependency for its own floored platform, so
KeyboardShortcuts (macOS 10.15, built for 12.0) emits its weak
Clock.sleep(for:) ContinuousClock specialization with a 112-byte async
frame while modules built for macOS 13+ emit a 128-byte one. The linker
coalesces the body and the frame descriptor independently, and the
2026.9.7 arm64 build paired a body that writes byte 119 with the 112-byte
descriptor: the 2026.9.6 launch-abort shape that the release async-frame
audit now rejects.

Pass -target <arch>-apple-macosx<LSMinimumSystemVersion> to every Swift
module in both release build commands so all coalesced copies share one
frame layout, independent of dependency pins or link order.
2026-09-30 01:54:48 -07:00
Peter Steinberger
16be7f0f27
feat(macos): run the private app runtime on the OpenClaw Bun fork (#161603)
* feat(macos): run the private app runtime on the OpenClaw Bun fork

* fix(macos): let the bundled runtime load unsigned plugin addons

* test(macos): remove obsolete worker pruning closure tests

* test(macos): drop stale packaging test imports

* test(macos): migrate CI scope cases to runtime paths
2026-09-30 01:27:03 -07:00
Peter Steinberger
373701372f
fix(ci): declare the native write callback receiver-free 2026-09-30 01:09:48 -07:00
Peter Steinberger
f7468c5161
ci: reuse boundary compiler receipts across unrelated file additions 2026-09-30 01:09:48 -07:00
Peter Steinberger
68ef53a4a6
fix(release): publish parent fails after dispatch on a gate-blocked stale child, and on a slow beta-floor sync (#161632)
* fix(release): sweep stale publish children before dispatch and wait for the beta-floor sync

* test(release): drive stale-child guards before dispatch in workflow fixtures
2026-09-30 00:04:54 -07:00
Peter Steinberger
9dc08b1fac
fix(release): publish preflight stalls silently, frv continue is silent, and the SDK acknowledgement is reported late (#161633)
* fix(release): bound publish preflight observation, narrate frv continue, and surface SDK acknowledgement early

* fix(release): replay failed preflight reads with their exact filter

* test(release): model the exact-tag release lookup in preflight inventory fixtures
2026-09-29 23:53:05 -07:00
Josh Avant
bb812693f3
fix(ios): qualify releases against the stable Gateway (#161682) 2026-09-30 01:50:38 -05:00
Josh Avant
63144fe0c4
fix(qa): Inspector collection fails when a pairing link is reused (#161540)
* fix(qa): pair Inspector pages with fresh dashboard links

Acquire a fresh single-use dashboard handoff for every Inspector page.
Cover real pairing, replay rejection, identity and receipt selection,
reload, and Chat draft preservation in a release-only browser fixture.

Real Gateway proof: 81.29s wall including prerequisites, 17.84s test body.

* fix(ci): complete Inspector pairing fixture inventories

Align private-server discovery and non-release counts with the Inspector
fixture, and include it in the frozen-target fallback command. Shorten
nearby comments to preserve the existing workflow size limit.

Workflow guards: 137 passed, 8 skipped; 85.65s wall.
Focused routing: 6 passed; 11.77s wall.

* fix(ci): include preflight manifest in trusted checkout
2026-09-30 01:48:56 -05:00
Peter Steinberger
982f9f8bcf
refactor(scripts): reuse report grouping owner (#161672)
Reuse the existing scripts groupBy owner in the shared test-report renderer. Preserve first-seen order, finding order, complete counts, output caps and all CLI/guard contracts. Remove six production lines without changing tests or dependencies.

Validation: 19 existing focused cases, scoped lint, both zero-cycle checks and independent review. Release notes: internal tooling cleanup only.
2026-09-30 06:07:52 +00:00
Peter Steinberger
4b9e0b3866
refactor: consolidate cross-directory duplicate code (#161352)
* refactor: consolidate cross-directory duplicate code

Reuse canonical parsing, bounded-map, session-scope and service-binding owners. Share Markdown breakpoint scanning and Responses web-search detection; route lazy WebSocket consumers through the existing portable transport and remove their obsolete build rewrites. Preserve SDK, wire and persisted-state contracts.

* refactor: reuse allow-from projections in channel setup

* refactor: share literal schema merging through an internal owner

Keep schema guards and provider-specific flattening intact. Isolate the shared web-search predicate from exported projection types so its private declaration does not enter SDK type closures.

* refactor: reuse shared listeners and remove forwarding helpers

Reuse existing listener, normalization, setup-visibility and cache-control owners. Preserve callback ordering, error propagation, string filtering and public SDK exports. Remove obsolete internal aliases and their stale test mocks.

* refactor: retain SDK-reachable setup visibility wrapper

Drop the setup-visibility cleanup group after the strict SDK comparison identified its namespace declaration in 123 reachable export closures. Preserve the existing API without acknowledgement or baseline changes.

* refactor: reuse canonical helpers across runtime consumers

Reuse result, numeric, string, filesystem, blocked-key, and map-eviction owners while preserving caller policy and output. Remove redundant private branches and repeated comments. Doctor inputs, transformations, warnings, and persisted state remain unchanged.

* refactor: consolidate shared runtime and tooling operations

Preserve per-caller validation, ordering, diagnostics and public SDK contracts while reusing canonical owners. Correct the Discord plural reply-count diagnostic with an existing handler regression.

* refactor: narrow consolidation to retained runtime owners

Drop the new SQLite scope and cache invalidation extractions after their local fixture preparation exceeded the maintainer quick-validation budget.

* refactor: keep Discord cleanup behavior neutral

Drop the optional diagnostic spelling change and its unvalidated new test under the maintainer quick-validation limit. Runtime import edges are unchanged; the reverted test only removes an import edge from the verified acyclic graph.

* fix: repair consolidation checks and restart fixture

Remove two obsolete imports, tighten assertion counts after cast removal, and load the real null-writer in the restart outcome VM fixture. The focused restart fixture passes all 19 existing native cases.

* refactor: drop unproven session-runtime clone groups

Keep the original Codex environment reader and session-runtime helpers after an unattributed worker teardown failure. Preserve the test assertions and leave worker lifecycle repairs to their owner.

* refactor: preserve main sandbox string filtering
2026-09-29 23:05:36 -07:00
Peter Steinberger
5d0afd0673
fix(release): skip non-shard timing jobs and accept the historical config-runtime alias (#161635) 2026-09-29 22:42:38 -07:00
Peter Steinberger
14233c0e6d
chore(release): close out 2026.9.7 on main (#161587)
* test(gateway): await yielded orchestrator cleanup

Wait for the existing registry cleanup publication instead of racing real SQLite worker completion against a two-second polling window. Reuse the plugin-subagent observer and bind the new waits to the test abort signal so timeouts release their subscriptions.

Keep yielded follow-up dispatch on real timers so the acknowledgement helper cannot prematurely run the registry sweeper against mocked run liveness. Preserve all requester lineage, announcement, delivery, cleanup, and predecessor lifecycle assertions.

(cherry picked from commit 0591abe71e703b7cf7c69a42cd8b93287d9fe8d9)

* fix(test): stop resume registry sweeper before gateway teardown

The resume Gateway fixture left its subagent registry sweeper scheduled
across files. A later tick used the retired module generation to open a
shared-state worker against the next fixture's state directory. The worker
registered with the old database lifecycle, escaped current teardown, and
rejected the recreated SQLite pathname.

Reset the fixture-owned registry without persistence before closing the
Gateway so its timers stop before database retirement.

Validation: the six-file same-worker driver reproduced the pathname error
before the fix (270.19s instrumented), then passed all 49 tests (215.72s).
The delete-state-lifecycle file passed all 10 tests alone (102.65s wall).
Targeted oxfmt, oxlint, git diff --check, and independent review passed.
All diagnostic instrumentation was removed.

(cherry picked from commit 613b3f8818d21d0c0d65333b29aabce11dd5afe9)

* test: settle reply recovery before deleting fixtures

The aborted-restart fixture could delete its store while asynchronous
main-session recovery-owner release was still preparing a write. Tracing
reproduced that release recreating/registering the prior store during the
following onAdopted test, invalidating the shared registry generation.

Join reply successor-admission barriers and close each fixture database root
before deleting its directory or resetting the reply registry. Extract the
fixture owner to keep the existing large test file from growing. Re-enable
the adopted-claim cancellation test. Production registry guards and retry
contracts are unchanged.

Validation with OPENCLAW_E2E_SKIP_BUILD=1 pnpm test:e2e:gateway against
src/auto-reply/reply/agent-runner.runreplyagent.e2e.test.ts:
- Original file: 220 passed, 1 registry failure, 1 skipped; 358.38s Vitest.
- Original isolated adoption case repeated 20 times: passed; 170.95s Vitest.
- Repaired aborted-restart/cancellation/adoption sequence: 3 clean runs,
  3 tests each; CLI wall 225.68s, 190.48s, 99.70s.
- Repaired full file: 221 passed, 1 different steering timeout; 528.47s Vitest.
  Remaining failure: keeps the replacement source when retired admission
  completes, first source admission did not persist (existing 5s guard).
  Focused replay passed in 151.18s Vitest; this does not establish a fix for
  that timeout. No timeout was changed.
- Source test types, scoped oxlint, formatting, and independent review passed.
  Broad check:changed guards passed through the core graph boundary; stopped
  before its duplicate messaging type pass. pnpm tsgo:test:src passed that
  graph plus the remaining source test graphs.

The full original CI shard and Telegram case were not replayed within the
requested 35-minute investigation limit.

(cherry picked from commit 0b679e62f6)

* fix(release): retire the 2026.9.7 live-shard waiver before main moves to 2026.9.7

RELEASE_WAIVED_LIVE_FILES (28597852) keyed on package version 2026.9.7 dropped three live files whose only case was skipped on the release branch. Main keeps those cases enabled, so once the closeout sets main to 2026.9.7 the adapter would silently stop running them. Refs #161083 #161084.

* chore(release): record the shipped 2026.9.7 changelog on main

CHANGELOG/2026.9.7.md, its contribution record, the index entry, the finalized Unreleased section, and the Matrix plugin changelog, byte-identical to the 2026.9.7 release SHA c074824a.

* chore(release): set main to the shipped 2026.9.7 version and record its update compatibility

Root version 2026.9.7 with pnpm release:prep version alignment, the macOS Info.plist version, and the verified npm tarball (sha512-/8N2Ln…RQwRWA==) recorded in the update compatibility inventory.

* test(gateway): move yielded orchestrator follow-up cases into their own module

The 0591abe7 forward-port grew agent.sessions-and-models.test-utils.ts past its line-cap ratchet. The yielded-orchestrator follow-up matrix is a self-contained case set, so it moves unchanged into agent.yielded-orchestrator.test-utils.ts, still loaded by agent.test.ts in the same module graph.
2026-09-29 22:39:14 -07:00
Peter Steinberger
f34c2a21db
feat(release): let a release lead record an exact-job flake instead of blocking publication (#161515)
* feat(release): accept exact-job recorded flakes in release validation

A release lead can classify one failed Normal CI job of a Full Release
Validation run as a flake through the trusted classification workflow. The
receipt binds the exact job id and attempt, CI child run, FRV parent run and
attempt, and Release SHA, and carries a tracking issue or PR plus a reason.
Release Decision, the manifest, the publisher's live re-derivation, the step
summary, and the GitHub release notes tail treat it as a visible advisory.
Required classes and every other child stay blocking.

* docs(release): fold recorded flakes into the shared release boundaries

* fix(release): scope flake receipt discovery to the CI child run

* fix(release): require main lineage for flake receipt producers

* test(release): copy the flake classification module into tooling fixtures

* fix(release): keep advisory-only release notes verifiable
2026-09-30 04:31:07 +00:00
Dallin Romney
7a438dc93e
fix: validate extended-stable package upgrades (#161450)
* fix(release): validate extended-stable package upgrades

* fix(release): route upgrades by baseline channel

* fix(release): pin extended-stable registry upgrades

* fix(release): bind extended-stable registry candidate
2026-09-29 20:49:42 -07:00
Peter Steinberger
5367527c27
fix(codex): preserve native assignments across parent rotation (#160952)
* test(upgrade): isolate native assignment fixture lifecycle

* test(upgrade): use explicit fixture control flow

* test(upgrade): decode native fixture websocket frames

* fix(codex): preserve native assignments across parent rotation

* test(codex): cancel native rotation at the commit boundary

* test(codex): use fresh MCP fixture threads

* test(codex): remove redundant attestation call assertion

* test(codex): tighten rotation fixture types

* test(codex): return a valid workspace resume response

* test(codex): isolate schema lifecycle bindings

* test(codex): route assignment lifecycle through the database broker

* ci: refresh native assignment checks after main repairs

* ci: refresh native validation after parent-audit fixture repair

* ci: refresh native checks after duplicate route repair
2026-09-29 20:43:38 -07:00
Peter Steinberger
4157ed60ff
fix(ui): stop serving a failed Control UI build on the next Gateway start (#161438)
* fix(ui): stop serving a failed Control UI build on the next Gateway start

Rolldown writes the complete bundle, runs the writeBundle hook, and only
then reports aggregated resolve errors and exits non-zero. scripts/ui.mts
built straight into the served dist/control-ui and reused the runtime build
identity, so a failed build left a correctly stamped tree that the Gateway's
asset health check accepted as ready on the next start (observed: a bare
markdown-it-emoji import broke the Control UI after a gateway-profile worktree
install skipped ui/ dependencies).

scripts/ui.mts now builds into a same-depth dist/control-ui.build-<pid>-*
sibling, runs both validators against it, and renames it into place only on
success, restoring the previous output if publication fails. Dead-process
leftovers are reclaimed on the next build and tsdown's dist cleaning skips
in-flight staging trees. The gateway worktree-setup workload now installs
./ui... because the Gateway builds the Control UI from source on first start.

* fix(ui): restore the previous Control UI bundle after an interrupted swap

A builder killed between moving dist/control-ui aside and renaming its staged build into place left the previous complete bundle only in a dead-process .retired sibling, which the next build's leftover cleanup deleted. The next build now restores the newest dead-process retired bundle when the served path is missing, before reclaiming leftovers, so a failed retry still serves it.

* fix(ui): keep interrupted Control UI builds out of packages

A killed builder cannot reach its cleanup, so a dist/control-ui.build-* staging or retired sibling can survive until the next UI build. Exclude those siblings from the package files list, which also drives the dist inventory, and type the new UI test fixtures for the scripts test lane.

* fix(ui): retry transient Windows denials when publishing the Control UI build

Windows scanners and indexers can briefly deny renaming a freshly written or served directory with EPERM, which failed an otherwise valid rebuild on its first attempt. Every publication rename now retries EPERM, EACCES, and EBUSY on a bounded 100-1600 ms schedule (about 3 s) and keeps restoring the previous bundle when the denial is permanent.
2026-09-30 03:28:17 +00:00
Peter Steinberger
9d2ef5e1da
feat: add process census evidence for retained artifacts (#160292)
Process-census capability needed to distinguish a dead managed-service handoff owner from a surviving descendant: a Windows process census (PID, start identity, command line, cwd, owner SID with the foreign-owner rule from the Unix contract), verified Unix UID provenance for incomplete observations, and retained-artifact reference matching that reports matching versus unverified PIDs. Existing callers keep their classification when the new evidence is absent.

Refs #159897 (the reclaim itself follows in #160488).

Landed under the pre-existing-red rule: the remaining CI failures were current main reds in the merge window (fast-lane config expectation fixed by e0ec544eb1; cron service tests fixed by 5f76cc437d; update-candidate-canary from b36eb3e7b1).
2026-09-29 19:59:42 -07:00
Peter Steinberger
53ccbd1b7b
feat(release): watch FRV children and rerun one child with a bounded budget (#161516)
Adds `pnpm frv watch --run <parent>`: it resolves Full Release Validation children from the parent's dispatch-job logs and reports each attempt transition and failed job once, with runner labels. It tolerates transient GitHub failures and resumes from a small state file.

Adds `pnpm frv rerun --run <parent> --child <key|run-id> [--max-attempts N]`: a bounded, audited single rerun-failed-jobs request. It reruns the green producer when a consumer binds its run attempt (#161317) and checks the new attempt for duplicate or missing jobs.

Retires `pnpm frv prioritize --run`, since the priority variable no longer controls admission. `--restore` stays.
2026-09-30 02:47:22 +00:00
Peter Steinberger
306e8a892e
fix(release): make FRV child evidence reuse match real receipts (#161520)
GitHub omits empty-string workflow_dispatch inputs from github.event.inputs,
so every sealed child receipt lacked the "" defaults that the parent-side
request filled in, and receipt validation rejected every candidate. The
per-candidate errors were swallowed, so dispatch logs never said why.

Treat empty and absent inputs as identical on both sides, only spend the
five full validations on receipts for this target and tooling, scan 100
runs, and log why each candidate was reused or rejected. Reuse now requires
the child's Tooling SHA to equal the current parent's, enforced at dispatch,
plan sealing, and final verification (previously any main-ancestor tooling
was accepted, which would have become live with this fix).
2026-09-30 02:43:26 +00:00
RoboClaw
65b0412dd6
fix(test): unblock cold isolated runtime tests (#161497)
* fix(test): unblock cold isolated runtime tests

Prepare selected runtime prerequisites inside the admitted container using the existing test-selection and build owners. Reserve a bounded 16 GiB envelope for artifact builds after reproducing an 8 GiB cgroup OOM; source-only tests retain 8 GiB and other isolation controls are unchanged.

Validation: 33 focused selection/admission/lifecycle tests; real cold isolated runtime build and 4 test cases passed with confirmed container cleanup; scoped changed checks with final test-lint repair passed; independent P0/P1 review scoped-clean.

Work session: https://team.openclaw.ai/chat/roboclaw/dashboard/bca94781-5883-4623-b4db-b67759e38d78

Co-authored-by: steipete <58493+steipete@users.noreply.github.com>

* fix(test): unblock cold isolated runtime tests

Worked on by:
- @steipete

Co-authored-by: steipete <58493+steipete@users.noreply.github.com>
OpenClaw-Publication: 55da5e5d-251f-45b6-b94d-c77b3e792326

---------

Co-authored-by: steipete <58493+steipete@users.noreply.github.com>
2026-09-29 19:22:31 -07:00