Commit graph

2 commits

Author SHA1 Message Date
Peter Steinberger
fcdd870217
fix(watch): complete cache migration and ignore unchanged invalidations (#161543)
Complete the fs-safe watcher migration for catalog and usage-template caches, with non-persistent subscriptions so one-shot commands exit naturally. Preserve atomic-save and missing-root recovery behavior.

Re-read selected sources after undetailed invalidations and restart the dev child only when content changes. Keep unchanged Skills coverage available, and honor CHOKIDAR_INTERVAL for forced polling and automatic fallback. Remove redundant native transports and transport-specific fixtures.

Completes #159226 and supersedes #158182. Dependency pins remain owned by main.

Co-authored-by: Vincent Koc <vincentkoc@ieee.org>
2026-10-02 09:14:04 -07:00
Peter Steinberger
52992fb086
refactor: use fs-safe for file watching and preserve Doctor plugin captures (#159226)
* refactor(infra): share filesystem observation policy

Centralize polling overrides, guarded source admission and metadata sampling. Bound remote file notifications and join accepted output before shutdown.

* refactor: migrate config, skills, memory and dev watchers

Use fs-safe invalidations and scope replacement while keeping source selection, settling, reload and indexing policy with each owner. Remove first-party Chokidar, bespoke native transports and duplicate observation tests. Preserve joined application work and watch-limit degradation; isolate manual Gateway writer fixtures from filesystem readiness.

Supersedes #158182. Thanks @vincentkoc.

Co-authored-by: Vincent Koc <vincentkoc@ieee.org>

* fix(memory): use host filesystem observation for captured plugins

Pair watch with SDK Root admission so captured plugin dependencies cannot split
fs-safe's module-local Root registry. Preserve the strong Root check and route
existing watcher tests through the host SDK boundary.

Keep the eager Skills subscriber snapshot with Array.from and remove the new
lint suppression without changing the production suppression allowlist.

* fix: preserve fs-safe fallback for legacy polling overrides

Treat legacy false, zero and empty polling overrides as a native preference
so Bun and installations without the native addon retain working observation.
Capture Config polling-recovery eligibility per observer to preserve the
explicit override retry limit.

Prepare the Skills recovery test through its existing worker owner and load
the real snapshot dependency before cases begin. Keep recovery assertions,
timeouts and artifact verification unchanged.

* test(memory): check resolved observation backend health

* fix(plugins): preserve native capture authority and lifetime

* test(ci): scope native Doctor proof to its runtime owners

* fix(plugins): break capture storage type import cycle

* fix(memory): import observation types through host SDK

* test(gateway): isolate operation journal fixtures

* fix(qa): point worktree lifecycle scenario at surviving run-end cleanup tests

#160308 deleted src/agents/worktrees/service.run-end-cleanup.test.ts, which the
managed-worktrees-workboard-lifecycle scenario still listed as a codeRef, so
extensions/qa-lab/src/scenario-catalog.test.ts failed on main. The surviving
run-end cleanup outcome coverage lives in service.test.ts (late claims, stale
lifecycle writes) and service.removal-safety.test.ts (dirty retention).

(cherry picked from commit 3a300c650a)

* test: isolate planner contracts and share installer shell

* refactor(gateway): schedule remote skill refresh (#160318)

## What Problem This Solves

Remote-node skill refresh still kept a private debounce timer and threaded its handle through Gateway startup and shutdown.

## User Impact

Skill changes keep the existing 30-second debounce, now owned by the kernel scheduler. Shutdown joins an active refresh and suppresses late broadcasts. Updating needs no operator action or config, storage, or public SDK migration.

## Why This Change Was Made

Schedule refresh directly with the existing scheduler. Remove the timer getter/setter, delay option, runtime handle, and redundant close hook.

Production **+18/-36/net -18**; tests **+100/-38/net +62**; docs **0**. Production counts use src/** and extensions/** with the work-order test exclusions.

## Evidence

Deslop and Codex autoreview completed with no actionable findings through P2. Focused proof passed on blacksmith-testbox lease `tbx_01m3k4y56rf9r63tzs4hck3pek`, [run 36378613468](https://github.com/openclaw/openclaw/actions/runs/36378613468); unchanged source/test hashes were verified in the later proof. Each command used `pnpm test <file> --maxWorkers=1`:

| File | Tests | Command wall |
|---|---:|---:|
| src/gateway/server-startup-early.test.ts | 14 | 14.763 s |
| src/gateway/server-close.test.ts | 80 | 23.385 s |
| src/gateway/server-startup-lifetime.test.ts | 14 | 37.299 s |

The closeout docs PR will carry the complete census and merged-main sleep/shutdown proof; that live proof is still pending.

Exact-head `node scripts/check-changed.mjs --base 5362ba0ca3 -- <all 6 changed paths>` passed on `4fafc2add7`: production typecheck, 25 dependent test type graphs, lint (0 warnings/errors), dead-export scans and boundary guards. Check wall **1610.89 s**. Provider blacksmith-testbox, lease `tbx_01m3khvbn4fa9jb0bhe5qqp3we`, [run 36397188865](https://github.com/openclaw/openclaw/actions/runs/36397188865). The candidate was materialized in a clean detached worktree because native sync retained its hydration HEAD.

## Inherited CI failures and landing evidence

Completed exact-head CI [36402003337](https://github.com/openclaw/openclaw/actions/runs/36402003337) has two underlying failures plus its aggregate gate. Both match independent PRs:

- `gateway-agent-skill-refresh.e2e.test.ts:304` (called from line192): expected lifecycle count3, actual2. Identical without this cutover on approval [run36401825740/job108861794753](https://github.com/openclaw/openclaw/actions/runs/36401825740/job/108861794753). The unchanged synchronous `src/skills/runtime/refresh-state.ts` producer invokes the test's earlier registered listener directly. This PR changes a separate downstream remote-bin refresh consumer; the failing lifecycle count precedes the debounce/broadcast assertion.
- `subagent-completion-blocked.e2e.test.ts:78`, ordinary delivery exhaustion: expected suspended, actual pending. Identical on questions [run36403551682/job108867383832](https://github.com/openclaw/openclaw/actions/runs/36403551682/job/108867383832). Neither PR touches that completion owner.

Focused tests and the complete exact-head Testbox gate passed. The maintainer's standing instruction authorizes pinned admin squash over these inherited failures. No CI rerun or weakened assertion was used. The installed wrapper rejects the CLI admin argument shape; the native workflow's protected GraphQL merge route preserves the same squash, message, and expected-head payload used by that CLI operation.

(cherry picked from commit df785c0971)

* test(gateway): adapt skills proof to fs-safe lifecycle

* test(skills): normalize Windows observation lookup

* test(memory): isolate guarded reconciliation from native hints

(cherry picked from commit f07aececc0b5641a00adfd0271c89e7c7162d38d)

* test(channels): update Synology context builder inventory

Match the selected builder spelling after the Synology inbound route was inlined in #160639. Preserve the existing caller inventory assertion and production behavior.

* test(skills): assert subscription renewal during recovery

* fix(ci): exchange hybrid parallel groups within the job cap

Reuse the existing bounded exchange optimizer for final hybrid consolidation. Preserve group ownership, child worker limits and admission budgets while reducing the fs-safe composition from 80 to 78 rows. Scope exchange by backend so valid Blacksmith layouts and cap-independent plan identity stay stable.

* fix(test): retire subagent sweepers before SQLite owners

Stop the original subagent registry before non-isolated file cleanup retires
its SQLite owners. Join accepted sweeps and cleanup tails while preserving
synchronous fixture resets, successor scheduling, and the actor path guard.
Await the new reset contract in the concurrency benchmark too.

Baseline controls demonstrate an old registry tick recreating a retired
shared-state owner and premature settlement of three in-flight retirements.
The receiving candidate passes 102 owning/sibling cases, including the
17-case nested fixture (16 passed, one expected skip). The benchmark with
eight child sessions also passes. The owning run took 156.42 seconds.
Types, lint, formatting, export scans, and repository guards were qualified.
The exact historical hosted timer schedule remains unobserved.

* test(ci): isolate worker boundary fixtures from prebuilt mode

Keep synthetic historical runner fixtures in their explicit no-dist mode so
an inherited CI prebuilt flag cannot enter package preparation before the
worker-owner boundary. Preserve all six fixture modes and their assertions.

* test(ci): align runner fixture with canonical main

---------

Co-authored-by: Vincent Koc <vincentkoc@ieee.org>
2026-09-29 01:05:11 -07:00