Commit graph

10748 commits

Author SHA1 Message Date
Jason (Json)
28e16eedc3
fix: dashboards require approval after native app authentication (#159860)
* fix: dashboards require approval after native app authentication

* fix: register native auth test entries and remove the client type cycle

* fix: recover native dashboard auth on older WebViews and reconnects

* fix: preserve native dashboard auth across current app paths

* fix(ci): use idle hosted tooling workers during overflow

* fix(ci): preserve oversized tooling refusal

* test: load first-hop helpers in the survivor fixture

* test(macos): wire conversation auth fixture and tokenless route

---------

Co-authored-by: Peter Steinberger <steipete@gmail.com>
2026-09-28 16:16:53 -06:00
Peter Steinberger
2e7f4f8fcd
fix(ci): pack hybrid hourly tooling within the row cap
Use the existing hosted retry file packing for hybrid tooling so an
indivisible file can share its spare worker without increasing its price.
The exact hourly hybrid plan drops from 80 total rows to 78, within the
unchanged 77 Node plus two dist limit. GitHub remains at 76 total rows.

Add the hybrid hourly cap regression and retain the GitHub ownership and
cap assertions. Preserve the complete file/config/environment inventory,
release descriptors, timing weights, runner policy, and worker limits.

Validation: 26,658 tooling tests and all test types passed on Linux
Testbox; script types and lint passed there. The connection ended after
those gates, so final root-test lint and format checks passed locally.
The integration file passed all 10 cases locally. P2 autoreview is clean.
2026-09-28 15:04:46 -07:00
Peter Steinberger
0a5ef7d9e0
fix(node-host): updates never install on Bun-only macOS and Linux hosts (#160575)
* fix(node-host): prepare updates on Bun-only POSIX hosts

Read exact registry manifests and stream SHA-512-verified archives in process, then use shared Bun staging inside the private node runtime generation. Preserve npm preparation on Node and Windows, including existing schema and candidate checks.

Cover Bun preparation, private bins, retained generations, invalid registry archives, loopback discovery, and Windows npm selection. Remove the Linux smoke blocker and document the remaining Windows limitation.

* refactor(update): share registry package document reads

Consolidate registry URL construction, HTTP fetching, JSON consumption, deadlines, and response cleanup for update discovery and node runtime preparation. Preserve caller metadata mapping, HTTP error strings, diagnostic labels, and body timeout policies.

Remove the redundant registry status wrapper and share the existing discovery error return. This removes 25 production lines without changing preparation or its regression tests.

* fix(node-host): skip npm freshness probing for Bun preparation

* fix(node-host): create the private Bun node_modules root before staging

* fix(node-host): seed the private Bun global project manifest

* fix(node-host): skip Node engine checks for candidates run on Bun

* docs(bun): link node update preparation history row

* refactor(node-host): move registry archive reads out of install-source-utils

install-source-utils sits under plugin install owners; importing the shared
registry document reader from it closed an import cycle through the provider
and plugin metadata graph. The in-process manifest and archive reads now live
in a leaf module that only the node host imports.
2026-09-28 14:59:44 -07:00
Peter Steinberger
87702484d1
fix: worker turns fail on older session-host nodes after the Gateway adds a worker tool (#160147)
* fix(gateway): keep worker turns working on session hosts that predate new worker tools

A Gateway built from main added the `presence` worker session tool, and
every worker-turn launch to a published openclaw@2026.9.6 session-host
node failed with `INVALID_REQUEST: invalid worker launch descriptor`
followed by `node worker cancellation timed out`. The installed node
supervisor validates `toolAuthority.allowedToolNames` against a closed
vocabulary, and bundle refresh does not replace the supervisor.

The Gateway now negotiates the launch vocabulary per node, following the
existing capability pattern: it advertises
`node-worker-launch-tool-names-v1`, updated nodes declare
`workerHost.launchToolNames` only to Gateways that advertise it, and the
Gateway treats an absent declaration as the frozen 2026.9.6 vocabulary.
Tool authority filters the projected tool set by the destination
vocabulary, so turn authorization and the launch descriptor stay
identical. Unknown declared names are ignored, so future worker tools
need no further capability.

Update behavior: Gateway-first updates keep older nodes hosting turns
without newer tools (one info log names them); node-first updates keep
the declaration unchanged for older Gateways; updated pairs get presence.

* fix(scripts): add worker tool authority to the PR wrapper inventory

src/infra/node-runner-inventory.ts is in the trusted-anchor PR wrapper's runtime import closure and now imports src/worker/tool-authority.ts for launch tool-name negotiation; the extracted wrapper could not resolve it.

* test(gateway): await the shared async node tunnel manager fixture

Main moved createManager into node-worker-tunnel.test-support.ts as an async helper (#160415); the launch-vocabulary lifecycle test still called it synchronously.
2026-09-28 21:58:24 +00:00
Peter Steinberger
0b44647a90
fix(scripts): accept the CI workflow's fail-fast expression in admin landing (#160721) 2026-09-28 14:54:22 -07:00
Josh Avant
a19ae3eda4
fix(qa): unblock isolated harness tool and evidence checks (#160527)
* fix(qa): recognize settled subagent batches

Share requester-settle wake recognition between mock input classification and completion handling. Accept the current batch wording while retaining installed-candidate session wording and the existing provenance check.

The real catalog-only handoff spawned and completed a child on the baseline, but QA ignored its completion and timed out. Two focused regressions failed before this fix. The real handoff now passes, all 70 owner/sibling tests pass, and check-changed passes. Standalone changed-test wall times: input 3.21s; handoff 33.76s.

* fix(qa): admit canonical repository checkpoint commands

* fix(ui): stabilize Run Inspector evidence collection

Bind rendered evidence to the public run, execution, and selected decision
receipt identities. Add a shared collector that uses the component-owned
route model and a separate page, preserving the caller's Chat surface and
unsent draft through collection and later reload.

Cover exact run/execution selection, receipt cursor reload, back navigation,
wrong requested identity, missing receipt, and page cleanup in the existing
mock-Gateway Chromium harness. Document the collector and per-tab auth setup.

The regression failed on baseline before product edits because the rendered
run identity attribute was absent. Terminal R's optional assistant transcript
idempotency key is a distinct producer gap; this does not manufacture that
key, relabel historical cells, or change product appearance.

* fix(telegram): confine QA runtime and preserve readiness evidence

Run standard-library drivers through existing Python without UV inline-script virtual environments. Confine private runtime state, drain readiness pipes, retain structural diagnostics, and require verified teardown receipts before releasing recovery state. Preserve doctor, recovery, group, and published-upgrade callers.

Validation: native sandbox regression failed before the fix; 140 harness tests pass, and 72 final focused owner/sibling tests pass in 5.73s. Build exits 0. Broad changed checks stop on two unchanged TS2459 package-update test errors. Focused lint matches all 135 baseline findings with zero additions. No live credentials or Telegram sends.

* fix(telegram): avoid native scenario barrier watchers

* fix(qa): share production publication guard contract

* fix(qa): require named message availability in mock provider

A catalog dispatcher does not advertise every delivery tool. Require the
exact message declaration and a usable invocation surface, including trusted
system/developer instruction carriers. Finish with the recovered child result
when message is absent.

Prove the absent-message regression through the mock HTTP provider and retain
catalog delivery, namespace, text-only fanout, and subagent handoff coverage.

* fix(qa): pass checkout roots to script scenarios

Expand the selected repoRoot separately from each scenario outputDir in the
maintained test-file runner. Prove the argument and CWD contract through a
real subprocess with external artifacts and fresh passing producer evidence.

* fix(qa): distinguish tool declarations from instruction prose

* fix(qa): bind forked-context evidence to native receipts

* fix(qa): bind repeated-ingress MCP scheduler

* fix(qa): add owned provider continuation checkpoints

Let maintained mock scenarios hold one session continuation before response
bytes and observe its request cursor and tool-call identity. Reuse the
provider request log and scenario signal/shutdown lifecycle; replacement
requests proceed normally without touching Gateway decision state.

* fix(qa): keep harness proofs within their owners

* test(qa): exclude all registered runtime consumers

* fix(qa): preserve runtime inputs and enforce checkpoint launches

* test(qa): skip runtime proofs when tools are absent

* fix(qa): complete checkpoint launcher test admission

* test(qa): compile repeated ingress child before execution
2026-09-28 16:48:47 -05:00
Kimi Yu
9190ad7c12
fix(codex): retain completed command output with Codex 0.158.0 (#160487) 2026-09-28 14:38:08 -07:00
Peter Steinberger
b779775249
fix(release): admit 2026.9.7 recovery helper and plugin scan inventory (#160717)
fix(release): admit 2026.9.7 recovery helper and plugin scan inventory

Full Release Validation for 2026.9.7 (parent 36466521174) failed two
release-tooling checks that main never exercises:

- The installed-package verifier applied the 6 MiB root-dist cap to
  dist/package-update-activation-recovery.mjs, the ~66 MB sealed
  self-contained recovery helper added by #158491. Treat it like the other
  self-contained bundles (80 MiB cap, no reparse), as #155877 did for the
  sqlite-store worker. Every main-based FRV since #158491 failed here.
- plugin-npm-security-scan had no reviewed inventory for release/2026.9.7.
  Freeze the 9.6 counts, record four post-9.6 test-file-only findings
  (codex auth-refresh harness spawn, feishu proxy env test, imessage
  client test #156334, signal socket probe count after #159648) in the
  current inventory, and map release/2026.9.7 to it, as #153372 did for 9.6.
2026-09-28 14:36:10 -07:00
Peter Steinberger
8857192115
fix(test): include all runners for package directories (#160653)
* fix(test): run package directory tests in their owning projects

* ci: refresh package directory checks after planner repair
2026-09-28 14:36:06 -07:00
Irish Joseph
68a9ce0324
fix(browser): read Windows browser version via environment data so spaced paths stop breaking the PowerShell probe (#160432)
* fix(browser): read Windows browser version without breaking spaced paths in PowerShell probe

Windows PowerShell appends any extra -Command argument to the script text,
so passing the executable path as an argument made the version probe fail
with a ParserError for any path containing a space (including the default
Chrome install path), and the child stderr leaked into doctor --json.
Pass the path via environment data and capture probe stdout only.

* test(browser): prove Windows version probe on spaced paths natively

* test(browser): keep failed Windows metadata probes quiet

---------

Co-authored-by: Ayaan Zaidi <hi@obviy.us>
2026-09-29 02:51:17 +05:30
Peter Steinberger
de73318f5d
test(agents,plugins,ui): remove low-value tests (batch d087) (#160605)
* test(discord): deslop t0071 tests

* test(plugins): deslop t0093 tests

* test(ui): deslop t0089 tests

* test(cron): deslop t0087 tests

* test(cli): deslop t0104 tests

* test(codex): deslop t0096 tests

* test(codex): deslop t0105 tests

* test(agents): deslop t0097 tests

* test(agents): deslop t0098 tests

* test: preserve independent contracts in batch d087

* test(agents): preserve independent tool-wrapper fixtures

* test: fix typed lint in batch d087 fixtures
2026-09-28 21:15:35 +00:00
Dallin Romney
cb8e8c7324
fix(release): scan trusted release-tool dependencies (#160694)
* fix(release): scan trusted release-tool locks

* test(release): refresh Vercel lock integrity

* test(e2e): include first-hop timing fixture
2026-09-28 14:09:00 -07:00
Peter Steinberger
5c3d2ee897
fix(plugins): npm-sourced plugins fail to install on Bun-only hosts (#160224)
* fix(plugins): install npm-sourced plugins on Bun-only hosts

Bun-only installs have no Node, so every npm-sourced plugin operation that
spawned `npm` failed. OpenClaw's managed plugin roots rely on npm's lockfile,
peer planner, and lock-based rollback, so keep npm's semantics and run a
pinned npm CLI under Bun instead of switching package managers.

- Add npm 11.20.0 as an exact dependency (npm 12.1.0 fails the managed peer
  planner under both Bun and Node).
- One owner, resolveNpmCommand(), serves all plugin install, update,
  uninstall, peer-planning, prune, and npm-config call sites: Node keeps
  `npm` unchanged; Bun runs `<bun> <bundled npm-cli.js>`.
- npm's bundled minimatch/brace-expansion/ip-address sit below the workspace
  security floors and cannot be overridden; a maintainer-approved exception is
  bound to npm@11.20.0's exact bundle, and the npm lock mirror verifies those
  members against the pnpm-locked npm tarball.

* refactor(plugins): keep npm fund suppression out of the shared spawn helper

Keep managed Bun npm installs quiet through their existing --no-fund arguments and safe install environment. Leave the shared process helper unchanged from main to reduce PR #160224 CI fanout.

* fix(plugins): keep the bundled npm lookup error private
2026-09-28 14:01:29 -07:00
Peter Steinberger
f0280ee9c3 fix(e2e): size first-hop budgets from measured update times
Allow 1800s per first-hop container and 2100s per lane using measured
published/candidate hop durations with a slow-host margin. Give the
six-source release self-upgrade job 80 minutes for two resource-limited
waves and setup. Preserve phase timing and concise update-step evidence.

Keep updater behavior, live authority checks, source coverage, and runner
concurrency unchanged. OverlayFS runtime-copy and journal-read overhead
remain a separate product performance follow-up.

Validation: 211 focused tests passed locally with one worker (55.36s wall);
inner/outer budget assertions failed against the prior harness. Bash syntax,
formatting, lint, timing formatter smoke, and diff checks passed. Independent
Codex review found no actionable P0-P2 findings.
2026-09-28 13:22:12 -07:00
Peter Steinberger
548a1b899f
fix(pr): qualify attributed CI job deadlines
GitHub Actions can report an exhausted Node job deadline as cancelled.
Recognize only an explicitly attributed root with a matching current Actions
check-run, complete deadline annotations, consistent timing and unchanged
workflow. Retain its cancelled status separately from fail-fast collateral.

Keep source attribution, reviews, security, authority and exact-head checks.
Native regression and refusal controls pass; isolated Linux qualification
preserves the existing command deadlines and retained incomplete CI coverage.
2026-09-28 13:14:13 -07:00
Peter Steinberger
3ad12d9287
perf(nodes): finish remote session turns without status polling delay (#160162)
* perf(nodes): finish remote session turns without status polling delay

Negotiate bounded status waits with paired nodes so completed turns return
as soon as their durable receipt settles. Retain polling for older nodes
and preserve current authority, exact receipt identity, cancellation, and
physical cleanup barriers.

Overlap independent tool cleanup and transcript settlement before the
terminal ACK. Reuse main's capacity parser, the closed hosting schema,
and canonical command/deferred owners; remove duplicated inventory
comparison and receipt validation paths.

Validated on the assigned Linux lease with 388 focused tests and
pnpm tsgo:core. Production LOC versus the rebased merge-base is net -3.

* perf(cli): keep node inventory schemas out of root help

* test(nodes): use the shared worker capacity constant

* refactor(nodes): simplify concurrent transcript settlement

Preserve concurrent cleanup and transcript barriers in the inlined worker turn owner after the prompt-caching integration. Keep main failure precedence and block terminal acknowledgment when transcript settlement fails.
2026-09-28 13:05:40 -07:00
Peter Steinberger
28340c41f8
fix(pr): read complete REST facts for prior-CI admission
When GraphQL keeps a valid merge projection UNKNOWN, explicitly approved prior-CI admission may select the existing complete writer-bound REST observation. Preserve known facts and REST provenance, and revalidate live authority after the final complete snapshot. Ordinary merge and security requirements remain enforced.

Validation: 138 composed native cases, causal original refusal and late-authority mutant, selected static/type/lint checks, and independent review. No partial snapshot or synthetic passing CI state is introduced.
2026-09-28 12:02:14 -07:00
Peter Steinberger
df582a4010 fix(ci): pack partitioned planner proof within the main-tier row cap
Use the spare worker beside an indivisible hosted tooling file only when a
single companion leaves the existing cost unchanged and fits the existing
whole-job budget. Files above that budget remain separate.

The hourly plan drops from 80 Node rows plus two dist rows to 74 plus two.
Preserve all file estimates, timing constants, caps, canonical ownership,
and test assertions. Full release descriptors remain identical.

Validation: 1,168 tests in all 11 requested planner/routing files passed;
targeted hourly and overflow-refusal regressions, oxlint, oxfmt, script
types, line-cap and erasability checks, diff check, and P2 review passed.
2026-09-28 11:46:42 -07:00
Peter Steinberger
eadee73a32
chore(crabbox): require Crabbox 0.67.0 (#160560)
* chore(crabbox): require Crabbox 0.67.0

* chore(crabbox): finalize the 0.67.0 minimum upgrade

* test(crabbox): refresh version fixtures across consumers
2026-09-28 11:29:35 -07:00
Vincent Koc
4373f42f5f
fix(pr): keep correction publication valid after first push (#160587) 2026-09-29 00:58:00 +07:00
Peter Steinberger
3e4e9cd4ae
fix(pr): recover retired auto requests with replacement admin proof
Allow an explicitly selected different head after confirmed auto cancellation to use current prior-CI admin admission. Preserve the accepted intent, cancellation, captures, and CAS ancestry; rerun replacement-head review, preparation, security, authority, and CI attribution checks. Same-head provider-rejection recovery remains separate.

Validation: causal original refusal; 54 Linux native recovery cases; CLI forwarding; selected types, guards, lint and independent review. No failed CI or cancelled coverage is relabeled as passing.
2026-09-28 10:49:40 -07:00
Peter Steinberger
cb3755a343
perf(gateway): retain prepared state across channel config reloads (#160274)
* perf(gateway): retain prepared state across channel config reloads

Preserve session rows and model catalog readiness when channel transport settings change. Keep model/auth and roster invalidation under their existing owners and await atomic catalog replacement. A 7,908-row file-reload fixture drops from 1,253 ms to 46 ms with no rematerialization.

* test(gateway): align reload fixtures with runtime snapshots

Publish prior configuration through the runtime snapshot owner used by hot
reload. Keep the agent-local and neutral compaction assertions, share auth
fixture construction with the config test-support module, and use Vitest's
call-order matcher for credential publication ordering.

* test(ui): wait for lazy tooltip readiness before measuring motion

Wait for the open, positioned popup after the hover timer starts its lazy
module import. Preserve the existing motion and placement assertions instead
of reading shadow parts before custom-element registration completes.

* test(ui): share canonical liveness in progress widget fixtures

Seed the canonical session with the same running facts as its main-scoped
list response. Fresh descriptor snapshots can then reconcile without clearing
fixture-only liveness fields. Preserve the writer-scope negative control and
all existing progress and activity assertions.

* test: stabilize planner and heartbeat CI proof

Skip import parsing for files outside a targeted source scan while keeping full-graph reads complete. Synchronize heartbeat delivery fixtures on their existing lifecycle events instead of a shorter wall-clock deadline.

* test: stabilize CI fixture readiness and preparation

Keep independent planner inputs in separate cases, model completed AI declarations in the build stub, and wait for roster or composer readiness before UI interactions. Serve sharing evidence through the canonical mock session owner so owner-role fixtures render without a header error.
2026-09-28 17:47:34 +00:00
Peter Steinberger
63a7b9f74c
fix(packaging): restore cloud-session bootstrap for CommonJS bundles (#160583)
The bundleDependencies change in #160015 exposed exact-path closure checks rejecting valid undici CommonJS requires. Resolve require edges with Node file and directory rules while keeping ESM imports strict, and carry package metadata through bootstrap inspection.

Restore cloud-session dispatch on the next Gateway update. Add checker and transferred-artifact regressions; verify a real undici 8.10.2 bootstrap artifact without starting or stopping a Gateway.
2026-09-28 17:44:53 +00:00
Vincent Koc
567132c06b
fix(lint): plan against effective memory capacity (#160581)
Automatic semantic-check planning now uses effective memory capacity, so a physically large host with a small container ceiling selects the constrained graph and worker policy. Explicit settings and complete selected-file coverage are preserved.

Validated with 156 focused tests, native type/lint checks, original-code failure controls and a real 7 GiB cgroup on a 31 GiB host. Separate one-worker test costs and the limits of hosted per-file attribution are recorded in the PR.

Related: https://github.com/openclaw/openclaw/issues/160010
2026-09-29 00:44:42 +07:00
Peter Steinberger
a57426fa03
test(core,qa-lab,tooling): remove low-value tests (batch d086) (#160480)
* test(heartbeat): deslop t0066 tests

* test(tooling): deslop t0084 tests

Port shard 044e24ef93954c7b340e1f6901f47bfb4a718386 onto the campaign lane. Preserve the lane-added opaque retention child-owner regression and all three assertions verbatim.

Validation: Testbox passed all 222 tests in one file with zero skips (73.73s shard wall time). Targeted oxlint, oxfmt --check, git diff --check, and independent review passed.

* test(agents): deslop t0076 tests

* test(cli): deslop t0083 tests

* test(agents): deslop t0075 tests

* test(cron): deslop t0086 tests

* test(config): deslop t0064 tests

* test(qa-lab): deslop t0090 tests

* test(agents): deslop t0078 tests

* test(gateway): deslop t0085 tests

* test: retain boundary coverage and stabilize heartbeat proof

Retain distinct wire, routing, authentication, ownership, and recovery cases
identified during review of batch d086. Preserve the surrounding reductions.

Capture heartbeat awareness observations inside the transport-completion
callback and assert them after the run, preserving ordering without racing
startup against a separate five-second deadline.

Validation: focused Testbox suites, changed-file gates, targeted formatting
and plain lint, and independent review passed.
2026-09-28 17:43:50 +00:00
Peter Steinberger
c7ff0090ea
fix(e2e): release lanes fail on a missing scheduler entry and stale service state (#160306)
* fix(e2e): emit scheduler entry for Docker clients

Emit the scheduler imported by mounted release Docker clients and guard all Docker client runtime dist references against the configured build outputs.

Validation: generic guard fails on the original config; all 33 config tests pass on Blacksmith Testbox. pnpm test wall: 7.73s warm, 31.90s cold worker preparation; test time 192ms. Scoped lint, formatting, scripts/root-test types, and independent review passed.

* fix(e2e): reload first-hop service fixture between lanes

Reload the lane systemd shim after deleting its unit so a published updater does not see a stale loaded definition. Print both captured service-install streams when setup fails, preserving effective-service admission checks.

Validation: reset regression fails on original harness with the real fixture. All 17 systemd fixture tests pass on Blacksmith Testbox; pnpm test wall 13.62s, new case 556ms. Scoped formatting, lint, shell syntax, scripts/root-test types, and independent review passed.

* fix(e2e): stop gateway before restart recovery preparation

Standalone Doctor can start a stopped service. Use the existing confirmed-shutdown helper before preparing the separate recovery update, preserving restart, auth, and inactive-only start assertions.

Validation: all four new recovery cases fail on original harness and pass with the fix, including failed stop, active service, and open listener refusal. All 14 phase tests pass on Blacksmith Testbox; pnpm test wall 1.85s, new cases 15ms. Scoped checks and independent review passed.

* test(e2e): stop an inactive fixture service idempotently

* test(e2e): align survivor fixtures with recovery shutdown

Keep the verified shutdown phase before inference preparation. Model idempotent systemd stop in the prepared-service fixture and include the phase in companion and frozen-target harnesses without changing their auth, membership, restart, or survival contracts.

Validation: origin/main passes all 55 targeted tests; original branch reproduces 17 failures. Blacksmith Testbox passes all 898 survivor tests across 42 files, including recovery-phase and cron-seed. Changed-file test walls: mobile 2.33s, parking 7.36s, membership 3.41s. Formatting, shell syntax, focused lint, and independent review pass.

* test(e2e): include script clients in dist-entry guard

Scan both Docker client roots and resolve each runtime import against its own client URL. This closes the release guard gap for scripts/e2e without adding runtime entries or duplicate tests.

Validation: 33 config tests pass on Blacksmith Testbox; pnpm test wall 42.30s with cold worker compilation. Removing the scripts-only runtime-context entry passes the original guard and fails the extended guard. Formatting, scoped oxlint, and independent P2 review pass.
2026-09-28 10:41:34 -07:00
RoboClaw
8ff319f71c
fix(test): avoid service-permission failures before tests start (#160392)
* fix(test): avoid service-permission failures before tests start

Worked on by:
- @VACInc

Co-authored-by: VACInc <3279061+VACInc@users.noreply.github.com>
OpenClaw-Publication: 3dfcb5be-749a-48f1-968a-6b25efab0758

* fix(test): avoid service-permission failures before tests start

Worked on by:
- @VACInc

Co-authored-by: VACInc <3279061+VACInc@users.noreply.github.com>
OpenClaw-Publication: 996f3d1d-862c-4055-9e99-f85e2ab8dcf3

* fix(test): avoid service-permission failures before tests start

Worked on by:
- @VACInc

Co-authored-by: VACInc <3279061+VACInc@users.noreply.github.com>
OpenClaw-Publication: 799d41fc-e479-496d-a0d9-bc77d3c7ad73

* fix(daemon): include generated systemd units in inventory

* fix(daemon): inspect loaded systemd Gateways before builds

* fix: ignore stopped unrelated Windows tasks in live dist fence

* fix: break daemon inventory type import cycle

* fix: ignore unrelated unreadable launchd plists in test fence

* test: copy errno owner into declaration fixture

* fix: fence all test runtime build outputs

* test: isolate E2E setup from AI declaration repair

---------

Co-authored-by: roboclaw-bot <309084314+roboclaw-bot@users.noreply.github.com>
Co-authored-by: VACInc <3279061+VACInc@users.noreply.github.com>
2026-09-28 13:38:50 -04:00
Peter Steinberger
54d19ff2ca
ci: partition changed-node planner proof by ownership
Preserve all existing planner assertions across seven tooling files, with
independent dependency-hub scenarios separated from leaf-config coverage.
Retain every file in protected discovery, fast CI routing and qualification.
Apply the native planner-family budget during initial and measured packing,
and retain the artifact fixture's required native capacity.

Linux owner, type, boundary, lint and deadcode checks passed, along with
seven constrained row replays and the final focused routing regressions.
P2 review is clean. Broad hybrid class-vCPU-minute growth models at 1.64%
with observed setup overhead; final PR matrix limits remain unchanged.

The recent native CI monolith took 532 seconds. The split Testbox rows
peaked at 199 seconds, so the 150-second and five-minute goals remain open.
2026-09-28 10:34:50 -07:00
Peter Steinberger
14cc3a4da8
chore(release): retire the internal Tideclaw alpha release track (#160261)
* chore(release): retire the Tideclaw alpha release track

Tideclaw alpha/nightly publication is retired. Alpha remains readable as
history (existing v*-alpha tags, published versions, changelog and
upgrade-survivor baselines, product version ordering), but it can no longer
authorize a release.

Every active release boundary now rejects an alpha version or -alpha.N tag,
the alpha npm dist-tag, and tideclaw/alpha/* workflow or tooling routes:
preparation, Full Release Validation publication selection, npm preflight,
approval receipts, core/plugin npm and ClawHub publication, native handoffs,
and finalization. Channel mappings throw for alpha instead of falling through
to latest. The Tideclaw branch routes, alpha dist-tag options, the alpha FRV
publication route, and the alpha-only Docker runtime-assets job are removed.
Beta, stable, extended-stable, and correction releases are unchanged.

The release-openclaw-nightly skill is deleted and the release skills and docs
no longer describe the alpha track.

* test(release): drop retired alpha preparation and finalization expectations

* test(release): drop remaining retired alpha references
2026-09-28 10:01:49 -07:00
Vincent Koc
f010fc0b57
fix(tooling): stop reporting lazy imports as duplicate exports (#160558)
Recognize transparent literal dynamic-import forwarders in the shared export-collision guard while preserving member, argument and body restrictions. Both variable and inline forwarding paths use the same predicate; runtime behavior is unchanged.

Validation: original guard reproduces nine repository false positives and fails all three new positive fixtures. Candidate passes the repository guard and all 54 fixtures, including six new rejection cases, under a 7 GiB process-tree limit with swap disabled. Script types and formatting pass. Independent and exact-head ClawSweeper reviews identify no actionable findings.
2026-09-28 23:59:48 +07:00
Vincent Koc
f5302f0b24
fix(tooling): retain compiler artifacts when cleanup recording fails (#160176)
Latch uncertain compiler cleanup at the shared artifact owner before marker I/O. Retain the existing owner/child claim when recording fails, preserve both errors and refuse a later inherited generation until verified recovery.

Validation: 49 focused ownership/preflight cases, formatting, script types and typed lint passed under a 7 GiB process-tree limit without swap. Direct and nested ENOSPC cases deny later writes before mutation; the published v2026.9.6 updater environment also preserves serving artifacts through refusal and recovery. Claim formats and readers are unchanged.
2026-09-28 23:59:28 +07:00
Peter Steinberger
abb37032e7
fix(ci): skip unmatched source parsing in test selection (#160517)
Targeted import-graph scans now parse only files containing a requested search term. Omitted rows never enter the edge cache, so complete graph reads retain full import coverage. Preserve every PR-exempt hourly/release ownership assertion and the existing timeout.
2026-09-28 09:23:13 -07:00
Vincent Koc
517cbd9a80
fix(ci): bound Testbox lease ownership and SSH lifetime (#160052)
Select Crabbox 0.67.0 for native Testbox SSH teardown while preserving supported offline binaries and managed caches for other providers. Keep doctor and execution on one cache selector. Record retained allocation provenance and revalidate the original caller, checkout and receipt immediately before delegation.

Retire schema1 ephemeral lease receipts through explicit stop and reallocation. Update local and remote proof guidance. Release-owned changelog content is untouched.

Real wrapper/native 0.67 proof at 09994f preserved allowed-reuse receipts, refused a different supported session and post-admission receipt withdrawal before any native Testbox call, and exercised the exact v2026.9.6 schema1 writer followed by refusal, stop, fresh schema2 allocation, reuse and terminal SSH cleanup. The four production and four Testbox test blobs remain identical after integrating main741c9cb. Freshness36/36 and the earlier scoped hosted checks passed; final exact-head proof and review remain required.

The prior CI blocker repeated an inherited one-second placement polling failure and left reconciliation unjoined during teardown. Main-owned e4ea649 observes the real admission boundary and always joins recovery; c136ec6 and 958de2c repair plugin retention lifecycle proof and disposal capture. Integrate those upstream repairs instead of adding unrelated runtime changes. Main also owns the equivalent planner fixture filter, so it is absent from this PR.

Production net +247 owns allocation provenance, final admission and provider-specific binary selection; tests net +456, docs net +32. Follow-ups: existing claim-restoration fencing, direct-caller rollout, remote-proof docs planner ownership and SQLite fixture maintenance isolation. No fleet upgrade.

Co-authored-by: Vincent Koc <vincentkoc@ieee.org>
2026-09-28 23:22:36 +07:00
Peter Steinberger
714c2c6baa
perf(ci): avoid redundant planner scans (#160531)
Prefilter wide reverse-import queries before parsing source facts, preserving
unmatched files for later full-graph discovery. Reuse the canonical commands
inventory and prepare scoped target patterns once per planning call.

Warm single-case wall times improved from 68.30s to 31.41s for PR-exempt
ownership and from 118.08s to 50.82s for leaf-config/dependency ownership.
Final alternating orders: integration -> leaf 36.72s -> 52.04s;
leaf -> integration 50.10s -> 30.17s. Keep all assertions, coverage,
PR-tier selection, and 120s timeouts unchanged.

Validation: 340 tests across the three planner owner files, plus 1779
source-selector tests; script typecheck, scoped lint/format, diff check,
and independent review. The three-file single-worker run took 777.53s;
individual test-file times were 625.774s, 132.342s, and 11.642s.

Related: #160224, #159431.
2026-09-28 09:18:44 -07:00
Peter Steinberger
f6c237fd0d fix(test): keep prebuilt E2E readers from rebuilding AI
9c69495d6f added missing AI declaration repair to generic Vitest admission.
That admitted another shared-artifact writer after direct E2E setup marked
its generation prebuilt, and during explicitly prebuilt local runs. The
unchanged bounded runner reproduced two tsdown invocations for ordinary
preparation and one for prebuilt preparation when declarations were absent.

Keep that repair in its existing CI shard owner, which prepares qaRuntime's
missing declarations before concurrent readers. Remove the duplicate generic
admission calls; ordinary E2E retains its one AI build and local prebuilt
consumers require caller-prepared artifacts. Retarget declaration tests to
that CI helper and document the boundary. Preserve @vincentkoc's #159624
matching Control UI assets, stale-asset rebuilding, read-only prebuilt
identity rejection, and cancellation guarantees unchanged.

Timing: controlled direct-E2E tsdown invocations fall 2 -> 1 (ordinary) and
1 -> 0 (prebuilt). The prebuilt admission fixture, including four stub
readers, took 787ms; this is not a real compiler benchmark. Full unchanged
run-vitest-bounded file: 39 tests passed, 49.29s pnpm wall (baseline: two
expected failures, 44.56s). Five sibling files: 400 passed, 83.99s wall.
Modified vitest-prebuilt-ai file: 8 passed, 2.64s pnpm wall, 3ms test time.
No new tests or increased test inventory; existing boundary regressions
fail on the original code without weakening their expectations.

Validation: oxfmt, scoped scripts/root-test oxlint, tsgo:scripts,
tsgo:test:root, and git diff --check passed. Siblings cover UI admission,
prebuilt identity rejection, E2E setup, CI shard ordering, and AI repair.
2026-09-28 08:51:11 -07:00
Peter Steinberger
9c69495d6f fix(ci): prepare packed AI declarations before prebuilt tests
qaRuntime intentionally omits declarations, but private-QA CI marks its output
prebuilt and skips the normal typed E2E setup. The packed AI tarball then
contains no declarations and fails its external TypeScript compilation.

Check the selected consumer's manifest type entries and run the existing typed
AI build only when an entry is absent. Share this preparation between local
Vitest admission and the CI shard owner, before concurrent readers start.
Keep explicit skip-build, complete typed output, unrelated shards, and the
sparse historical workflow adapter unchanged.

Proof: cold qaRuntime reproduced TS7016 locally and on Node 24.19.0 Testbox;
the repaired concurrent shard passes with all 17 type entries. Linux typed
AI compilation took 4.7s; the unrelated shard passed in 2.88s with no typed
build and no declarations. Median unrelated preparation over 1,000 local
calls changed from 0.001ms to 0.009ms, with zero added build commands.

Validation: check:changed passed; 162 focused tests and all 203 shard-runner
tests passed. New regression file: 8 tests, 2.19s pnpm wall; complete modified
shard suite: 6.98s pnpm wall. Independent review's historical-adapter concern
was rejected after executing the sparse adapter with the prebuilt flag and
without the current prerequisite module.

Canonical full private-QA pnpm build passed (665.84s), followed by the ordinary
non-prebuilt packed-package E2E consumer (58.34s wall, 16.49s test duration).
2026-09-28 07:40:40 -07:00
Peter Steinberger
670415932e
fix(update): updates and repair fail on Bun-only installs without Node (#159431)
* fix(update): keep Bun-only installs on Bun through update, repair, and install

Bun-only installs (no Node anywhere, CLI launched as `<bun> openclaw.mjs`,
Gateway service pinned to that Bun) could not update: maintenance children
were spawned with a bare `node`, and the package preinstall rejected every
Bun lifecycle without a persistent Node.

- resolveNodeRunner() now returns the running Bun executable, making it the
  single owner for OpenClaw CLI children (finalize, repair maintenance, and
  legacy compat chunks now use it).
- Service-runtime decisions stay Node-owned: under Bun there is no
  Bun-as-Node fallback, and the running Bun is not judged by Node engines.
- The preinstall skips Bun's temporary lifecycle node shim and, only when no
  persistent Node exists, accepts an explicit OPENCLAW_PACKAGE_BUN_LAUNCHER
  (Bun 1.4+). The updater sets it when it runs under Bun.

* fix(update): stage Bun and pnpm updates before reading admission-gated options

Candidate-owned admission stages the package with admission-gated
accessors. The Bun/pnpm staging preflight spread those params into the
package-manager probe, which evaluated requirePackageReplacement before
admission and failed every candidate-driven Bun- or pnpm-managed update
with "Staged update has not been admitted for activation." Pass only the
probe's fields instead.

Also align the Bun-only docs with the refusals observed from published
2026.9.6 updaters.

* fix(update): keep split-root Node service preflight under Bun

Retain the recorded Node runner for an owned split-root Gateway when the
updater runs under Bun, so target engine compatibility is checked before
service refresh. Node-driven updates keep their current runner; Bun
services keep the Bun runtime rules.

The caller regression failed before the fix and passed afterward. After
merging main, all 155 focused tests and both requested typechecks passed.
Caller test cost: 4 tests in 90.68s wall, including worker preparation;
the new regression itself took 16ms. Autoreview: scoped-clean at P2.

* docs(install): run the Bun-only service install through Bun

* fix(install): exempt only Bun's temporary node shim in preinstall

Restrict the exemption to bun-node-<hex> directories targeting the running Bun. Persistent node aliases to that Bun now fail closed before later valid Node, with or without the launcher marker. Addresses the P1 finding on #159431.

Validation: the two new launcher-contract cases failed before the fix (69 passed, 2 failed; 16.22s wall). After the fix, all 71 passed (5.32s wall; 306ms test time). Scripts typecheck, changed-file formatting, diff whitespace check, and scoped P0-P2 autoreview passed.

* fix(install): recognize the Bun fork's per-user node shim directories

Address the ClawSweeper P1 on #159431 by recognizing the stock and pinned fork lifecycle shim formats, including debug names and 16-hex private fallback suffixes. Keep the running-Bun realpath check and persistent-alias rejection unchanged.

Validation: the focused preinstall suite failed before the fix on all four fork acceptance cases (76 passed, 4 failed; 3.45s wall), then passed all 80 tests after the fix (5.02s wall; 237ms test time). pnpm tsgo:scripts, targeted oxfmt check, and git diff --check passed. Independent autoreview was scoped-clean through P2.

* test(security): keep the no-output policy timeout test off the real deadline

The test fires its mocked 1 s no-output timer by hand once the forked
process tree has written its pid file, but the 10 s overall deadline
still ran on a real timer. On a loaded runner, starting the process tree
can take the whole deadline, so the overall timeout won (CI run
36401631363 failed at exactly 10 s with "policy command timed out").

Intercept the overall timer with a non-firing placeholder so only the
path under test can fire. A local probe with a forced 10.5 s startup
stall fails with the CI signature before and passes after.

* test(process): keep forked no-output fixtures silent after publishing their pid

The forked descendant in writeForkingNoOutputScript printed "ready" to
stderr after writing its pid file. Tests fire the mocked no-output
deadline as soon as that file exists; the runner defers the decision by
one 0 ms timer, and output arriving in that window refreshes the idle
timer, which clears the pending decision (the timer has been reused
since 912b21b9). The run then settled only on the 10 s overall timeout,
failing "kills forked policy command children on no-output timeout" in
CI run 36401631363.

Drop the post-pid output; the tests fire the deadline by hand, so nothing
needs it. Revert 80fda958, which masked the overall timeout on the wrong
diagnosis and turned the race into a 120 s hang (CI run 36407894516).

A forced-ordering probe (output 100 ms after the pid, decision deferred
300 ms) fails with the CI signature on the old fixture and passes on the
new one. install-policy.test.ts and secrets/resolve.test.ts both pass.
2026-09-28 07:15:04 -07:00
Peter Steinberger
3e0def1451 fix(ci): drop references to the removed worktree run-end cleanup test
fb55609468 (#160308) consolidated run-end cleanup coverage into the
surviving service tests but left the deleted file in CI policy watches,
the proof inventory, timing metadata, and a QA scenario code reference.

The scenario failure was reported in CI run 36427486045, job 108949566158
on #160224:
https://github.com/openclaw/openclaw/actions/runs/36427486045/job/108949566158

Before: the catalog test failed with the missing worktree cleanup codeRef
(1 failed, 44 passed) in a worktree outside the older parent checkout.
A nested worktree falsely passed because the resolver found the parent's
old test file. Concurrent commit 3a300c650a repaired the scenario reference;
retain that fix and remove the three remaining obsolete script entries.

After: catalog 45/45, inventory/timing/node-plan 225/225, and 29 targeted
changed-plan cases pass. The complete four-file tooling baseline also
passed 550 tests. Scoped oxfmt, oxlint, and git diff --check pass.
Independent review reports no actionable P0-P2 findings.
2026-09-28 06:59:54 -07:00
Dallin Romney
baf2edda0d
fix(release): compact npm preflight manifests (#160464) 2026-09-28 06:29:00 -07:00
Peter Steinberger
b27c79ba15
fix(ci): prevent leaf-config ownership check timeouts (#160458)
Resolve target shape once before scanning database-worker owners, avoiding repeated directory probes for every inventory entry. Keep literal-file, directory, glob, and watch selection unchanged.
2026-09-28 06:18:32 -07:00
Peter Steinberger
860a5b863b
refactor(workspace): consolidate transfer lifecycle and fixtures (#160415)
* refactor(workspace): consolidate transfer lifecycle and fixtures

* refactor(scripts): measure workspace manifests through their capture owner

The reconciliation core no longer re-exports a naming-only manifest wrapper; the workspace computation benchmark now imports captureWorkspaceManifest from workspace-manifest-worker directly.
2026-09-28 06:13:08 -07:00
Peter Steinberger
b95394a915
perf(state): stop re-reading the ownership lock on every shared-state read (#160139)
* perf(state): monitor ownership for explicit shared-state reads

Keep process/projection exclusion, schema leases, native-handle custody,
physical database identity, maintenance admission, and SQLite transaction
and commit grants. The hot cost was repeated observation of the same
process owner at independent read lifetime boundaries.

Monitor owner/projection sidecars every second and cache canonical owner
paths for one second, invalidating them on ownership and schema lifecycle
transitions. Only the dedicated shared-state read transport opts in.
Reads tolerate that bounded window after out-of-band lock replacement or
alias retargeting; writes and generic broker jobs keep fresh checks.
Foreign and maintenance ownership remain strict. No schema or dependency
changes; update behavior is unchanged.

One warm Control UI reload on synthetic state (46 RPCs), base f5866a5d0bcc
versus candidate: realpathSync.native 1478 -> 913 (-38.2%), openSync
536 -> 275 (-48.7%). Counter windows were 8.087s and 9.810s. The top-40
stack buckets show explicit-read sidecar opens 256 -> 0; candidate
monitoring adds 20 opens. This is syscall evidence, not a latency claim.

Validation: fake-timer ownership/projection theft, alias expiry, lifecycle
and failed-cleanup invalidation, and no warm read ownership syscalls.
The transport regression fails on the base with an owner-sidecar open.
Both core typecheck lanes, touched-file CI oxlint, line-cap ratchet,
diff checks, and independent P2 review pass. Runtime/UI builds pass
with declaration emission disabled; runtime postbuild/stamps are included.

The broader run initially timed out the unchanged synchronous snapshot
test after 162.790s against its 120s deadline under heavy host load.
Its same-file-order replay with the original concurrent lifecycle suite
passed in 5.002s without a fix, so the original cause remains unresolved.
The selected 57 files finish with 852 passing tests and two existing skips
across the successful runs. No timeout or assertion was weakened.

New ownership test cost: 8 cases, 7.879s summed test time and 80.49s shard
wall under host load. The two changed read/borrow files shared a 417.61s
unit shard; initial worker transformation dominated that cold run.
Benchmark and broad-suite evidence use the original f5866a5d0bcc base;
rebase integration preserves the unchanged ownership implementation.

Benchmark limitations: an initial cold-start warmup timeout was discarded.
Both base and candidate reproduce existing Teams/Zoom plugin shutdown
cleanup failures, but both isolated Gateways stopped and state was removed.
Both UI builds retain the existing 363.4/362.9 KiB startup-JS advisory.

* fix(scripts): add gateway state owner directory to the PR wrapper inventory
2026-09-28 12:27:53 +00:00
Peter Steinberger
8aec7806a9
fix(ci): keep PR-exempt ownership proof within its timeout (#160408) 2026-09-28 05:24:09 -07:00
Peter Steinberger
dfd4cc8a5d
fix(e2e): recognize delegated post-core update workers (#160384) 2026-09-28 05:06:29 -07:00
Peter Steinberger
733ceac1a7
fix(windows): discover Gateway Startup-folder launchers (#151674)
* fix(build): refuse dist rebuild under a live managed Gateway

Stop pnpm build and run-node auto-build from deleting hashed dist modules
while a managed Gateway ExecStart still points at this checkout.

* chore(build): drop unused fence message field

* fix(build): also fence direct tsdown and run-node rebuilds

Cover the cleanTsdownOutputRoots path and refuse early in run-node so
live managed Gateway dist cannot be wiped outside build-all.

* fix(build): keep live Gateway stop off the rebuild path

Dispatch gateway stop/restart from existing dist, apply --profile
before service inspection, and fence only physically overlapping
checkouts.

* fix(build): keep live dist fence off tsdown declaration graph

Load daemon inspection lazily so tsdown fixtures and plugin-sdk dts
generation do not import service-layout. Move run-node recovery tests
to a sibling file so the line-cap ratchet does not grow.

* fix(build): use import type for SpawnOptions in live-dist tests

* fix(build): dispatch source-only QA reports before the live dist fence

qa parity-report and qa coverage already run from source without
rebuilding private QA dist. Check that path before refusing a live
managed Gateway rebuild.

* fix(daemon): discover managed Gateway bindings across profiles

The live-dist fence needs every installed managed selector, not only the
current OPENCLAW_PROFILE. Reuse includeManagedOpenClaw scans and leave
findExtraGatewayServices semantics unchanged.

* fix(build): refuse live dist rebuild for every overlapping Gateway profile

A default-env build could still replace dist under a sibling profile
Gateway. Inspect all managed bindings, name offenders, and point operators
at stop or openclaw update.

* fix(daemon): keep managed Gateway binding discovery under lint limits

Move profile binding mapping out of inspect.ts, list managed services via
listManagedOpenClawGatewayServices, and use toSorted for profile naming.

* fix(build): keep live-dist refusal out of updater restore and system census

Propagate admissionRefused so update-gateway-build does not rm live output
roots after a pre-mutation fence refuse. Enumerate system-scope managed
units with explicit systemdReadTarget so user and system siblings both reach
the fence.

* test(build): pass a compare function to toSorted in the fence fixture

* fix(build): satisfy live-dist fence CI type, knip, and assertion gates

Treat toSorted comparators as possibly undefined, stop re-exporting the
unused binding type, and read launchd profile env through the record
owner instead of a type assertion.

* fix(daemon): inspect systemd template instances in the live-dist census

Discovered `openclaw@.service` files were forwarded as unit names, so the
fence queried a non-runnable template and fail-opened past a live instance
when a separate user Gateway was installed. Reuse the existing instance
resolution owner and cover user+system template inspection through the
service reader.

* fix(tooling): pin foreign-cwd tsconfig and clear CI gates

- pin TSX_TSCONFIG_PATH in the CLI shim only when the cwd lacks one, so
  fixture-spawned run-node.mts resolves workspace profile imports
- pass --import scripts/tsx.mjs to live-updater spawns and pin the
  tsconfig in run-node-lifecycle fixture envs (production loader parity)
- drop the unused systemd re-export for knip and move the template-
  instance test to service.systemd-scope.test.ts for the line cap

* fix(daemon): inspect registered Gateway services accurately

Read strict Windows service commands from registered actions and preserve
supported launcher encodings without treating dynamic CMD expansion as a
verified command. Retain exact task identity and bound systemd selection.

Use one platform inventory collector for the existing full and diagnostic
projections. Report incomplete inspection through Doctor while preserving
status JSON and the existing managed-service filters. Keep load-state
inspection behind a leaf capability instead of reverse facade imports.

Refs #151608, #151463.

* fix(daemon): classify service commands by position

Share runtime and root-option parsing so Node display names and profile values
cannot classify a Node service as a Gateway. Preserve literal shell, env and
generated launchd wrappers, and honor the executable selected by launchd.

Read systemd's single-quoted arguments through the existing parser and preserve
literal apostrophes when rendering. Keep strict Windows command rejection
separate from lenient diagnostics, and align the native-boundary fixtures with
that contract without weakening ownership or source-preservation assertions.

* fix(daemon): inspect direct registered task actions

Read literal executable actions from their registered Scheduled Task and
revalidate the action before returning command facts. Strict runtime inspection
uses the same registered command instead of a default launcher.

Keep executable paths out of managed launcher provenance. Direct actions remain
outside automatic update service management when no restorable CMD/VBS launcher
exists, preserving the prior stop and definition ownership boundary.

* test(windows): prove installed service upgrade paths

Extend the existing native Scheduled Task proof with verified immutable package
handoff and fixed fresh, published 2026.9.3, and published 2026.9.4 CLI cells.
Require live version/build identity, selected PID replacement, peer continuity,
strict registered-action inspection, and settled cleanup before evidence
publication and disposable installation retirement.

Keep source-only and repair modes, permissions, deadlines, and native lifecycle
owners intact. Direct executable fixtures remain disabled and preserve the
updater's unsupported-mutation boundary. Native execution remains pending.

* test(doctor): retain managed Windows launcher provenance

Keep the generated CMD path on the existing managed-service fixture before
and after reinstall so Doctor admission sees the installation it models.
Preserve all stop, install, restart, rollback, and direct-action refusal
expectations without changing production behavior.

* test(windows): exercise native autostart ownership boundaries

Extend the existing published-updater cell with its stopped, task-owned
Scheduled Task. Capture real admission, exercise native enable/disable,
and preserve the definition and files after foreign-owner refusal.

Test retained admission separately from restoration so legitimate same-root
refresh remains supported. Restore the original XML through the existing
fixture lifetime; do not broaden product control or workflow permissions.

Modeled owner tests, selected source checks, and independent review pass.
Actual Windows execution remains required before a native proof claim.

* test(windows): retain sanitized installed command failures

Keep unexpected command stdout and stderr in bounded failure diagnostics so
JSON-mode CLI errors survive native proof failures. Reuse the existing
terminal and support redaction owners before clipping; withhold incomplete
captures while preserving the existing truncation failure.

Move the existing command runner into its own test-support module and update
both consumers without changing timeout, exit, signal, or cleanup semantics.
Real-child regressions reproduce both privacy defects and pass with the
correction. The original installed Gateway failure remains undiagnosed.

* fix(daemon): preserve Windows probe budgets and Doctor cleanup eligibility

Use the existing cold PowerShell startup budget for Task Scheduler queries
without an explicit deadline. Keep periodic activation checks bounded and
preserve unknown results rather than treating timeouts as task absence.
The native probe test now exercises the production default.

Share Doctor's existing legacy cleanup classification with its registered
preview. Keep unsupported platforms, scopes and unrecognized Linux unit
names as findings without advertising removal or invoking unrelated cleanup.
Extract the classifier and complete cleanup test group without dropping
assertions or changing native mutation ownership.

Retain the failed installed Windows runs and their source identities;
new package and native upgrade qualification remain separate requirements.

* chore(daemon): retire stale Doctor size baseline

The split Doctor service module no longer needs a max-lines suppression.
Remove only its stale baseline entry so the shrink-only ratchet matches
actual source. Product, dependency, workflow and fixture bytes are unchanged.

* fix(build): keep native fixture import closures complete

Keep the legacy source-update transaction at its existing direct CLI call
site so native declaration-library consumers do not load its CLI-only source
resolver closure. Copy the exact PID helper inputs into the runtime fixture.

Use the real legacy-loader file URL in its existing subprocess invocation,
so execution and unused-file analysis share the same dependency reference.
Preserve all assertions, lifecycle guards and package runtime behavior.

Co-authored-by: Donnie Fiander <44792682+DonnieFi@users.noreply.github.com>

* test(windows): retain sanitized service proof observations

Record bounded install and status facts before semantic assertions so a
successful CLI exit cannot hide the native inspection reason. Exclude
free-form stderr and private response fields, and preserve existing
execution, deadline, and cleanup assertions.

* fix(windows): discover exact Startup Gateway launchers

Carry the focused Startup discovery changes from #151674 onto the current
service inventory owner. Preserve exact file identities, selected fallback
classification, Scheduler inspection failures, and published updater
fingerprints. Extend the existing installed Windows cells with read-only
Startup diagnostics and owned-file cleanup.

Focused local regressions pass, including Doctor failure on the original
collector and success with Startup discovery. Native Windows qualification
and the managed binding/fence integration remain pending.

* test(windows): retain safe installed-service observations

Retain allowlisted install/status JSON before strict semantic assertions,
without copying private response fields or opaque successful stderr.
Compose service observation and exact sibling-refusal checks through one
internal options record and migrate every test/support caller together.

Preserve native process cleanup, readiness assertions and command budgets.

Co-authored-by: Donnie Fiander <44792682+DonnieFi@users.noreply.github.com>

* test(update): isolate source updater fixture inputs

Give the synthetic compiler declared memory capacity through the existing
memory owner instead of inheriting competing CI workers. Keep real build
heap admission and all lifecycle assertions unchanged.

Use the established TypeScript loader for the Linux-only live-updater CLI
case so it reaches the platform refusal rather than failing during import.

Refs #152875. Retains the exact failing Linux CI evidence and contributor work.

Co-authored-by: Donnie Fiander <44792682+DonnieFi@users.noreply.github.com>

* fix(windows): preserve native status inspection defaults

Keep the CLI RPC default distinct from an explicit timeout so Windows
service inspection can use its existing cold-start budget. Forward
explicit load-query deadlines through the Task Scheduler owner while
preserving lifecycle defaults and timeout diagnostics.

* fix(windows): preserve native status inspection defaults

Keep the CLI RPC default distinct from an explicit timeout so Windows
service inspection can use its existing cold-start budget. Forward
explicit load-query deadlines through the Task Scheduler owner while
preserving lifecycle defaults and timeout diagnostics.

(cherry picked from commit 9c3bb9c32a)

* fix(daemon): remove inventory file-helper type cycle

* fix(models): retain discovered models after refresh failures

Record successful legacy catalog results at the producer boundary so unavailable refreshes retain the accepted inventory. Preserve explicit outcomes, advisory SDK fallback behavior, and first-discovery starter policy.

* fix(models): preserve skipped catalog outcome semantics

Mark bundled static, configured, and advisory catalog projections with
explicit empty outcomes so legacy success inference cannot promote them
to observed account inventory. Preserve live outcomes and helper types.

Keep exact auth provenance histories and move existing fixture/policy
code into focused owners where required by the line-cap ratchet.

Validation: 447 producer and sibling cases, 56 shared self-hosted cases,
95 auth/policy cases, causal missing-outcome failures, maintained checks,
and independent review.

* test(plugin-sdk): keep discovery loader types acyclic

Move the shared loader type into a leaf consumed by both discovery
contract helpers. Preserve its public provider-test-contracts export
without a child-to-parent type import cycle.

Validation: maintained Madge check reports zero cycles; core, all core
test graphs, extension test types, lint, formatting and independent
review pass. Runtime behavior and previous catalog proof are unchanged.

* fix(plugin-sdk): mark generated static catalogs explicitly

Keep the generated non-live, non-strict catalog adapter from claiming
successful acquisition for manifest or configured rows. Preserve null,
errors, strict and custom callbacks, static catalogs, and public types.

Validation: three existing controls fail before the correction; all49
owner and sibling cases pass afterward, with types, lint, line caps and
fresh independent review clean.

* fix(models): remove unlanded catalog outcome inference

* test(windows): handle omitted scheduled task settings

* test(windows): handle exported task defaults and reuse pinned packages

Preserve disabled-task and cleanup assertions when Task Scheduler omits default settings. Bind the package to its source commit and permit only the reviewed fixture paths to differ in tooling. Includes the canonical sharing-fixture retirement correction from cd2f12e (#158205).

* test(windows): retain installed proof before native cleanup

Persist settled commands and the original inspection failure before awaited cleanup. Prepare the existing compiled worker cache before each installed campaign and give installed Actions steps room for the unchanged native test and teardown bounds.

* test(windows): diagnose native probe context after cleanup

Compare read-only PowerShell probes under native and isolated OS contexts only after failed cold acceptance and verified cleanup. Preserve probe deadlines, process ownership, original results and package identity. Carry the canonical failure-only Gateway hook phase observer for the separate CI recurrence without claiming a causal repair.

* test(windows): record native module cache context

Retain only projected environment key names and module-cache path kind/hash so the diagnostic can distinguish an actual caller-selected cache from an absent one without recording values or changing launch behavior.

* test(windows): preserve native context in installed fixtures

Replace the duplicate OS environment filter with the existing native service projection and case-aware merge. Preserve caller-selected PowerShell module cache routing while retaining private application, temporary and npm paths and excluding application credentials and Node options. Two old-code regression failures and 23 fixed owner/parity cases pass; native timing remains to be verified without changing deadlines.

* test(windows): preserve native context and failure evidence

* test(windows): await installed readiness and guarded cleanup

Installation acknowledges service activation without certifying runtime readiness. Await the existing HTTP readiness owner after install and update before the unchanged Task/PID/RPC/build assertions. Stop installed supervisors through the guarded service owner before generic task deletion; the probe PID-file cleanup does not own those processes. Preserve original failures, cleanup authority and all existing native deadlines.

* test(windows): distinguish absent tasks during fixture cleanup

Use authoritative Scheduler inspection before guarded stop. Confirmed absence after uninstall or failed registration skips only stop; unknown inspection still fails, and full process/port cleanup remains required. Preserve task authority on a genuine stop failure.

* test(daemon): wait for installed gateway readiness before inspection

* test(daemon): stop installed fixtures through their profiled CLI

* fix(daemon): give actionable Startup refusal guidance

* fix(update): retain native custody during partial-stop recovery

Compensate a failed native stop through the existing guarded service restart,
revalidating the original binding inside its operation lock. Require completed
recovery and preserve the original failure without starting a build or custom
shell command. Successful and failed-build custom restart contracts stay intact.

Four published-shell regressions fail on the former path and pass after this
repair; 21 unchanged controls, selected checks, and independent review pass.

Co-authored-by: Donnie Fiander <44792682+DonnieFi@users.noreply.github.com>

* fix(daemon): inspect extra Windows services before removal

Keep verified Node and legacy diagnostics visible while offering read-only
Scheduled Task inspection in Doctor and deep status. Preserve the existing
service cleanup owners and diagnostic JSON shape.

Carry the canonical SQLite fixture host-context and dependency-selection
repairs from #158209 and #158409 for the inherited CI collection failure.

* test(windows): expect informational Startup fallback findings

* fix: restore asynchronous harness task completion

Complete the shared task-content projection cutover for asynchronous finalization and delivery. Preserve the captured Incognito policy and exact task-assignment fences.

The unchanged worker suite reproduced six ReferenceError failures before the fix and passes all seven cases afterward; the sibling SDK runtime suite passes all 30 cases. Independent review is clean through P2.

* fix(daemon): compose read-only Windows inspection hints

Preserve exact Startup file inspection while adopting read-only Scheduled Task advice and Inspection headings. Keep Node diagnostics visible without suggesting deletion. Combine both owner behaviors and reuse the qualified lower correction without unrelated fixture changes.

* test(windows): accept omitted PID for stopped Startup runtime

* test(windows): budget installed lifecycle phases separately

Give the fixed installed cells room for their serial setup, unchanged updater command window, native assertions and cleanup. Keep source-native and command deadlines, the quiet-run detector and the hard job limit unchanged. Record real completed phases after durable evidence writes, and verify workflow envelopes against the actual body and teardown configuration.

* fix(qa): retain the current Telegram renderer reference

Carry the canonical one-line metadata correction from #158313. The retired formatter was moved into format.ts, already listed by this scenario; execution and all assertions are unchanged. The owning catalog suite passes all 56 cases and fresh review is clean.

* test(windows): budget installed service phases and retain progress

* test(daemon): retain installed update failure progress

* test(windows): retain failed installed update progress

Read failed published-updater progress through the existing asynchronous SQLite owner before native cleanup. Retain only safe phase, status and step timings; preserve all command, body and teardown deadlines.

Validation: five installed-fixture tests, services types, targeted changed checks and independent review through P2 passed. Shared diagnostic implementation qualified in the Windows fixture owner.

* test(daemon): budget the full published Windows update

* chore(ci): carry canonical line-cap repairs

Carry the necessary kernel environment and publication alias hunks from #158625 (8f22acdd), and the interrupted replay fixture extraction from #158538 (25f4dfa5). Preserve their assertions, state contracts, and Peter Steinberger attribution. Update installed workflow comments to match the already-qualified body budgets.

* test(windows): retain final updater verification diagnostics

* test(windows): verify stale admission refuses foreign task resume

* fix(daemon): bound aggregate Windows service inventory

Carry one monotonic deadline through Scheduler discovery, launcher reads, registration revalidation, and missing-launcher metadata. Preserve completed findings and report incomplete inventory when the shared budget is exhausted.

* fix(windows): share inspection deadlines with Startup launchers

Compose the shared Windows inventory deadline through Startup directory inspection, launcher reads, and definition revalidation. Preserve completed discoveries and report every exhausted scan, including arbitrary filenames and the final completed Task.

Use the same caller deadline for exact Startup state capture and reread. Stalled file operations remain unknown diagnostics without extending the native inspection allowance. Retain immediate Task-query failure fallback, profile bindings, and process cleanup propagation.

* test(openrouter): split Fusion prompt coverage below line limit

* test(daemon): complete system template runtime fixture

Provide the loaded-runtime reader's required ControlGroup metadata. The missing
native response correctly made inspection unavailable and broke the template
regression after main integration; production behavior is unchanged.

Validation: the original fixture fails locally for the CI assertion, then all
76 service-scope and loaded-runtime cases pass after correction. Independent
P2 review is clean. No assertion, timeout, or package source changed.

* test(openrouter): align the canonical Fusion suite label

* fix(daemon): classify registered helpers without profile admission

Let read-only registered inventory classify faithfully revalidated commands
without requiring an OpenClaw profile. Keep selected-service profile
admission strict and preserve launcher, Task, script and deadline checks.

Native disabled-discovery exposed a false warning for a static node --version
helper. The actual collector regression fails before this fix and passes
with all 206 owner and sibling cases afterward.

* fix(daemon): separate task discovery from profile admission

Registered inventory must classify fully inspected unrelated commands without requiring an OpenClaw profile. Preserve resolved matching profiles for selected-service reads and all native definition revalidation. Reproduced through the real inventory collector after the installed Windows discovery fixture failed; retain the selected-service refusal control.

* fix(build): retain Startup identity in external recovery guidance

Carry the shared stop, successful rebuild, and start guidance while keeping Startup-only siblings tied to their exact entry paths. An already-current update does not repair stale output. Fence tests pass 44 cases with one existing skip; script types, lint, formatting and independent P2 review pass. This copy-only successor is separate from the frozen f919 Windows package source.

* fix(daemon): recognize released waiting task launchers

Recognize the exact waiting VBS wrapper shipped by 2026.9.3 during
owned service reconciliation. Keep custom launcher behavior unknown and
preserve all command, root, Task and authority checks.

The real audit entry point rejected this released form as TaskLauncher
unknown-edit before repair. The causal regression and edited-launcher
control pass with 55 audit and 140 backup/rewrite sibling cases.

* test(windows): preserve primary installed update failure facts

Retain bounded sanitized fields from the original published-update JSON before unchanged failure assertions. This preserves primary reasons and failed checks that the output tail can omit, without changing capture ceilings, command deadlines, process ownership, or cleanup.

* fix(test): read Windows proof archive members portably

* test(windows): normalize native task export line endings

* fix(windows): enable verified disabled tasks on explicit start

Preserve captured autostart policy during update recovery; an explicit Gateway or Node start revalidates the selected task and its launcher/package or recorded wrapper before enabling and running it. Keep stronger local start fingerprints out of published update drivers serialized command shape.

* fix(windows): retain starts with unavailable enable metadata

Only an explicit disabled policy admits enabling. Preserve native Run for readable Tasks whose optional Enabled field is absent, while retaining typed failures for failed inspections. Keep the finalization fixture original gateway-entrypoint exports when replacing its install resolver.

* docs(windows): explain partial effects of an explicit start

A successful enable can remain after a later launch failure. Preserve the existing independently audited enable contract and require current authority for any further control; do not imply automatic rollback of a failed start.

* fix(build): bound exact Startup service inspections

Pass the existing Windows inspection allowance when the build fence reads an exact Startup entry. A stalled launcher read now aborts within that budget so the fence can still inspect and refuse a later live sibling. Preserve ordinary reader defaults and the cold native-Node import boundary.

The real fence/state/file-reader regression fails before the fix and passes afterward. All 122 fence and Startup/deadline cases, selected type/lint/export checks, native Node import, and independent scoped review pass.

* docs(update): clarify existing-build recovery commands

Source-runner gateway stop and restart use the existing CLI by default. Explain that applying source changes requires the documented stop, build, and start sequence.

---------

Co-authored-by: Donnie Fiander <44792682+DonnieFi@users.noreply.github.com>
2026-09-28 04:58:29 -07:00
Peter Steinberger
bff74e5f0b
fix(pr): recover explicit admin base-change rejections
Allow explicit same-head recovery when the retained prior-CI admin REST
request received the exact GitHub 405 base-branch-modified response.
Preserve each original intent and response in the outcome ancestry, bind
recovery to the current outcome with compare-and-swap, and repeat all live
review, security, authority, CI-evidence and head checks before dispatch.
Unknown, mixed, truncated or other provider responses remain ineligible.

Qualify the sent-request path through the native wrapper, including repeated
rejections, final-await tampering, stale outcomes and revoked authority.
Validation: 448 distinct recovery/CLI/sibling cases pass across retained
runs; fresh independent reviews are clean. The initial positive CLI fixture
needed an explicit main target; its causal original-code failure is retained.
All root type shards, coercion and dead-export checks, and script/test lint pass.
The pinned baseline's unrelated UI shard budget is already repaired on main
by 52ae7e314d.
2026-09-28 04:46:34 -07:00
Peter Steinberger
d36da73091
fix(models): prepare catalog pairs before replacing outputs (#160157)
* fix(models): prepare catalog pairs before replacing outputs

Forward-port the paired-output correction from 45f97d88dd. Preserve recoverable old and new artifacts after partial publication without overwriting foreign replacements.

Co-authored-by: Vincent Koc <vincentkoc@ieee.org>

* fix(models): reject unknown recovery cleanup identities

* fix(models): keep symlinked catalog outputs and exact recovery identities

Resolve paired output symlinks to their final targets before preparation, preserving links and creating dangling leaf targets when their parents exist. Reject aliases to the same target. Keep the v1-only writer unchanged.

Use BigInt identities for recovery cleanup and output snapshots; retain unknown or changed recovery entries. Consolidate cleanup without changing interrupted-publication retention.

Validation: 27 focused publisher tests pass with --maxWorkers=1 (8.15s wall); scripts and test-scripts typechecks, targeted lint, formatting, and diff checks pass. Real publisher CLI proof preserves symlinked outputs, including symlink/.. traversal and a dangling target. Codex autoreview found no actionable P0-P2 findings.

---------

Co-authored-by: Vincent Koc <vincentkoc@ieee.org>
2026-09-28 04:46:19 -07:00
Peter Steinberger
316f2ea6d5
fix(node-host): update checks fail on Bun-only installs without Node (#160154)
* fix(node-host): discover updates on Bun-only installs without spawning npm

* docs(bun): note headless node update preparation still needs npm

* docs(bun): record node update discovery change

* test(node-host): reject aborted registry reads with an Error
2026-09-28 04:43:50 -07:00
Peter Steinberger
ed911c3e46
perf(node-host): start worker launch helpers faster on session hosts (#160163) 2026-09-28 04:43:19 -07:00
Peter Steinberger
e3bee38f02
fix(test): prevent TUI identity PTY crashes in clean checkouts (#159832)
* fix(test): prevent TUI identity PTY crashes in clean checkouts

* test(ci): include TUI preparation in planner expectations

* test(ci): select eligible planner control fixtures

Select 96 PR-eligible tooling files before testing overflow, and pin the synthetic timing source. Require the same inherited-hosted group eligibility when selecting a recipient and exercising its negative controls.

Co-authored-by: Peter Steinberger <steipete@gmail.com>
2026-09-28 04:08:23 -07:00