The Telegram userbot skill's node:test (*.test.mjs) and Python (*.test.py) files are not Vitest-routable, so aggressive PR selection drops them as direct targets. Their only CI owner is the Vitest wrapper test/scripts/telegram-e2e-userbot-skill.test.ts, which the policy watch selected for 20 named non-test scripts only. Test-only skill PRs (e.g. #162968) and later scripts (reply-policy-*, telegram-runtime) planned nothing in either selection mode. Watch the skill's whole scripts/ directory, the agents/openai.yaml the wrapper parses, and the fixture it loads by URL. The wrapper stays in the release-only core-tooling family; it now also runs on PRs that touch the skill, the only PRs that can break these self-contained scripts.
51 KiB
| summary | read_when | title | sidebarTitle | ||
|---|---|---|---|---|---|
| How the slowest Node test families are split, balanced, packed, and cached |
|
Node test lanes | Node test lanes |
How the slowest Node test families are split, balanced, packed, and cached. Part of the CI scope and routing index.
The slowest Node test families are split or balanced so each job stays small without over-reserving runners:
- Hourly main and ordinary manual/release validation retain complete plugin and channel contract jobs with their existing weighted process selections and runner fallbacks. PRs select exact contract tests through the changed-owner Node plan; frozen historical manual targets keep their separate jobs.
- Additional checks: Current targets combine export-collision, session accessor/transcript reader, SQLite transaction, and SQLite schema checks into one serial source-contract row, preserving each command and collecting every failure. The complete plan has five rows on main and six on current-target dispatch; PRs select the applicable static rows; the SDK API report is allocated only for dispatch. Frozen targets retain their eight original rows and historical command fallbacks.
- Core unit fast/support lanes run separately; unit-src, Control UI, and gateway-core each use three deterministic file-weighted stripes, while the security and media/UI companion configs retain their scoped whole-config support groups; core runtime infra splits into process, shared, hooks, secrets, and three cron domain shards.
- Auto-reply runs as balanced workers, with the reply subtree split into agent-runner, commands, dispatch, session, and state-routing shards; dispatch further isolates core, delivery, and lifecycle entrypoints.
- Agentic gateway/server (control-plane) configs split across chat, auth, model, HTTP/plugin, runtime, and startup lanes instead of waiting on built artifacts.
- Normal CI packs only isolated infra include-pattern shards into deterministic bundles of at most 64 test files, reducing the Node matrix without merging non-isolated command/cron, stateful agents-core, or gateway/server suites. Heavy fixed suites stay on 8 vCPU while most bundled and lower-weight lanes use 4 vCPU. Previously promoted compact workloads retain 8-vCPU capacity through their semantic owners and selected files before packing. Timing or build-ownership changes cannot transfer that capacity to an unrelated row. Hosted stripes inherit their parent owner's capacity; whole named manual and release plans retain their existing routing. Full named bundles include their runtime-build mode when present, so
checks-node-bundle-infra-small-runtime-1andchecks-node-bundle-infra-small-1identify different prerequisite groups. Their test inventories and runner limits are unchanged; unique names let release collectors retain both jobs without ambiguous attempt evidence. - Hourly main on GitHub-hosted runners bounds agent-chat groups to 30 files and Gateway-methods groups to 96. Chat groups and Gateway-methods groups above 32 files run in separate rows; small Gateway tails keep ordinary packing. The measured updater test runs alone instead of extending a storage row. Storage retains its 64-file bound. These main-only splits retain every selected test, config, worker limit, build prerequisite, and timing generation; PR, full-release, and paid-runner plans keep their existing packing. The complete hourly main matrix remains within its 77-row budget.
- Canonical pull requests resolve their runtime tests from bounded changed owners, transitive import consumers, workspace/public SDK aliases, protected regressions, and explicit policy watches. The Node constructor consumes that selected set. Windows and UI jobs retain complete ordinary inventories when their platform/area owners change; unrelated UI families can still admit precise protected or affected files. Global inputs never request the full compact runtime repository. Missing selector capability fails preflight. Main retains its complete automatic tier, and ordinary manual/release validation adds the existing release-only inventories.
- Static correctness is independent of runtime targeting. Production and test types, semantic lint, formatting, SDK boundaries, schema drift, and Knip retain their required checks; ambiguous static ownership keeps full compiler/lint coverage. Runtime targeting does not need a full-suite fallback to preserve these guards. See the scope rules.
- A changed tooling test or owner selects its exact owner tests, affected transitive test importers, protected regressions, and explicit watches, subject to the named process-proof exclusions below. Tooling configs retain their file parallelism, worker limits, runtime prerequisites, backend routing, and capacity promotion. They never request a full-family PR fallback.
- Precisely resolved metadata-bearing files retain their canonical config, runner, worker/admission policy, and preparation requirements. Canonical allocation prices known subsets using existing file weights, retaining a cold-process allowance capped at the complete group price and charging preparation separately; opaque inventories retain their full price. Compatible whole execution envelopes may share an existing row without merging child processes or increasing their workers; serial siblings may share that row while parallel siblings retain their stripe-family separation. The PR admission pass partitions selected files into rows with a target of at most 150 estimated test seconds, using existing file/group timing evidence and conservative estimates for missing observations. It preserves complete files and every assertion; an indivisible file or canonical group can exceed that target. These are admission estimates, not measured job walls: setup, builds, queueing, and runner contention remain separate costs. Reduced groups use distinct timing identities so partial observations cannot overwrite full-parent timings.
- Control UI changes retain complete ordinary unit, mocked E2E, real-Gateway, and bootstrap families through their independent existing owner gates. Protected and affected files opt into other families precisely. Named release-only UI files retain the existing direct/owner opt-in mechanism. Empty real-Gateway phases are omitted while desktop-only carriers keep their required desktop proof. Configs, project isolation, worker policy, and runtime prerequisites remain unchanged.
- Directly changed plugin tests and their selected owner tests use exact-file selections in canonical configs. Unit-fast, contract, bundled, and E2E files retain their own suite boundaries. There is no whole-plugin PR fallback; hourly Plugin Prerelease and Full Release Validation retain the full extension inventory.
- The planner's
RELEASE_ONLY_TOOLING_SHARDSset and matching maintainer leaves in mixed fast configs defer the complete maintainer-tooling family on unrelated PRs. The tooling Vitest configs own the ordinary inventory and isolated/Docker catalogs. Maintainer leaves selected by fast configs retain their ordinary, isolated, or fake-timer owner and process pins; filtering mixed groups preserves product neighbors and gives the subsets separate timing identities. Dedicated product E2E and live tests, including the fivetest/scripts/*.e2e.test.tsgates, stay outside this tier. PRs touching a tooling test or owner select its owner files, transitive import consumers, and protected regressions, subject to the explicit release-proof exclusions below:scripts/**,src/scripts/**,test/**,.github/**,config/**, root package/pnpm inputs, tooling configs, and other inputs classified as tooling by the shared changed-path owner inscripts/test-projects.test-support.mts. That owner also covers Docker, agent/Crabbox tooling, app scripts/Fastlane, and extension scripts/package inputs. Directly changed tooling tests retain coverage on their own PRs except for the explicit process proofs. Hourly main and ordinary CIworkflow_dispatchinclude the family, including Full Release Validation'snormal_cichild against the frozen candidate. ItsRun Node test shardstep runs these unchanged tests before regular publication admission; an independent duplicate release test is not required. The existing approved preflight-only beta publication exception remains unchanged. Plugin Prerelease separately ownsagentic-plugins; CI's plugin exclusion does not control the tooling tier. Fork repositories keep full tooling because they do not use canonical PR targeting. Product tests retain their existing tiers except for the explicit runtime proof inventory below. - The explicit
CI_PROOF_TEST_FILESinventory keepstest/scripts/frv.release.test.tsandtest/scripts/install-ps1.release.test.tsout of PR plans, including directly edited tests. Hourly main, manual CI, and Full Release Validation retain both complete process tours in the canonical tooling config. FRV keeps its in-process continuation contracts infrv.test.ts; the installer keeps its source guards ininstall-ps1.test.ts. No cases or assertions are removed. - The explicit
RELEASE_ONLY_RUNTIME_TEST_FILESinventory defers measured runtime integration tours from canonical automatic PR and push plans while faster owner tests retain their primary contracts. It covers Doctor configuration and migration compositions, CLI startup, Gateway model/plugin/session and worker-environment tours, and updater, packaging, and storage recovery proofs. The existing release inventory remains:src/flows/doctor-health.test.ts,src/infra/update-managed-service-handoff-foreground.test.ts,src/node-host/node-worker-supervisor.recovery.test.ts,src/state/openclaw-database-preflight.lifecycle.test.ts, and all eightsrc/config/state-startup-corpus*.test.tswrappers. It also retains the real Git archive/deletion matrices insrc/gateway/server.sessions.archive-worktree-lifecycle.test.tsandsrc/gateway/server.sessions.delete-worktree-lifecycle.test.ts, plus the native/broker child-process retirement matrix insrc/process/supervisor/adapters/child.service-lifecycle.test.ts. Faster session-lifecycle, identity, worktree-owner, relay-state, and readiness tests remain in automatic CI. The startup-corpus paths owner supplies the complete family in its existing partition order. The files keep their canonical configs, process isolation, runtime prerequisites, worker limits, cases, and assertions. Whole-config CLI filtering shares the canonical include and exclusion owner without introducing file stripes or changing its worker and serial-job policy. Parallel groups preserve the reduced coverage timing identity when applying worker policy, so automatic CI samples cannot overwrite full release estimates. Selected owner tests and direct test edits run in canonical PRs and exact-head release-gate substitutes; unrelated paths do not admit the remaining release matrix. Manual CI, Full Release Validation'snormal_cichild, local full plans, and noncanonical repositories retain the complete inventory. The release tier also keeps the separatepublished-upgrade-survivorDocker cell with the published baseline,legacy-operator-statescenario, andauto-authrestart mode; the synthetic foreground-handoff cases do not substitute for that published-driver × candidate proof. - The keyed-publication scale benchmark (
src/gateway/session-row-projection.keyed-marks.benchmark.test.ts) also uses the runtime release inventory. Full manual CI and Full Release Validation retain its 4,428-row SQLite fixture and 70,000 publications; ordinary PR and hourly main plans defer it. Direct edits still select it on their PRs. The smaller session-row projection tests keep exact-key invalidation, cross-agent isolation, row lookup, and refresh behavior in regular CI. The benchmark stays in the canonical Gateway core config with every assertion intact. - Ten native maintainer-tooling and release matrices use
RELEASE_ONLY_RUNTIME_TEST_FILES:test/scripts/ci-linux-git.test.ts,test/scripts/pr-worktree-provision.test.ts,test/scripts/pr-worktree-interruption.test.ts,test/scripts/pr-merge-outcome.test.ts,test/scripts/pr-merge-admission.test.ts,test/scripts/pr-merge-rest.test.ts,test/scripts/pr-merge-receipt.test.ts,test/scripts/pr-merge-recovery.test.ts,test/scripts/full-release-validation-at-sha.test.ts, andtest/scripts/package-acceptance-workflow.test.ts. Tooling-owner PRs defer these matrices unless their test files change directly; ordinary main and product-only PRs already omit the tooling family. Faster Git-owner, import-closure, hosted merge, completion, state-isolation, and release-contract siblings remain in regular tooling coverage. - The same explicit runtime inventory also defers native runner state-cleanup, empty-discovery, and entrypoint matrices, correction-merge integration, and the bundled browser MCP package-manager matrix from tooling-owner PRs. Faster state-lifetime, home-isolation, runner admission, native argument parsing, correction preparation, and bundled MCP artifact tests retain their primary contracts. The complete compositions keep every case and assertion in their canonical configs. Direct test edits still select them; full manual CI, including hourly main and Full Release Validation, retains them.
- The media-persistence large-corpus proof and SQLite reliability process tour use the same release inventory for heap, sparse-session, large-payload, and composed crash/restore bounds. Ordinary media-migration, local snapshot repository, SQLite snapshot, and Doctor compaction suites keep their primary contracts in regular coverage; the complete resource-scale and process compositions remain in manual and release validation.
- The release inventory also retains the complete CLI help, Gateway-backed exit, web-output backpressure, plugin-authoring, and packaged-skill tours, the real npm dependency-repair and installed-provider authentication compositions, and the 256 MiB Git backup heap proof. Faster output finalization, Gateway RPC, command policy, skill readiness, plugin repair/discovery, and backup round-trip tests remain in PR and hourly main coverage. The moved files keep their canonical configs and all assertions, run when edited directly, and remain selected by full-tier CI and Full Release Validation.
- The MCP release proofs (
src/agents/agent-bundle-mcp-retention.test.tsandsrc/agents/mcp-stdio-client.cleanup.real.test.ts) retain the nested shared-worker file-order check and all six real POSIX relay/anchor shutdown cases, including their production cancellation deadlines. Ordinary CI keeps the private-transcript and requester-lifecycle tests, MCP cleanup-error propagation, and supervisor cancellation/exit tests with injected clocks. The composed process proofs run through the same agents-core config during manual and release validation. - Startup corpus tests use changed-owner selection on PRs; the six-file fixed smoke set below is the unconditional runtime floor. Canonical main retains its regular startup corpus, and full manual/release validation retains the complete family. A complete exact-tree Node coverage receipt removes the otherwise empty fast startup-corpus row. Historical manual targets preserve their supported split-file or legacy sharded corpus fallback. Current PRs do not add a separate full corpus job.
- Canonical
mainpushes use a Blacksmith integration compact with nondist Node jobs plus the dist boundary descriptor. Former multi-config walls (CLI plus CLI-process, isolated plus fake-timers unit fast, and the logging/process/runtime-config trio) are split into per-config shards so no single group floors a lane. Real Node+TSX command tests belong to the isolated CLI-process catalog, so ordinary CLI tests do not prepare a runtime. The process catalog splits by complete file costs, with split sizing bounded below by its complete file costs and runtime prerequisite so an older aggregate timing cannot hide newly owned work. File packing also includes the prerequisite; runtime-consuming CLI children may share one preparation in the same serial job when their complete combined estimate fits the existing 150-second budget. Each child retains its selected files, isolated process and two-worker limit; deliberately separated fixed stripe families remain apart. The gateway process file stays alone because its cold proof already takes 200 seconds. CLI children retain the 150-second sizing target and their two-worker limit. Ordinary hybrid bins containing only non-build CLI children may combine up to 250 predicted seconds, with each original child still admitted separately below 150 seconds; the hosted runtime prerequisite itself has a 160-second floor. They omit the low-signal-per-push tooling and TUI PTY groups while retaining product-runtime groups outside the explicit release tier, including three file-weighted stripes apiece for unit-src, Control UI, and gateway-core. Blacksmith serial admission stays at 200 seconds for the large class and 276 seconds for the small class. Ordinary groups that can share two process slots admit 360 predicted aggregate seconds; a group already above its serial cap stays alone. Manual dispatches and Full Release Validation retain the full named per-shard matrix. Hourly main runs the complete main-tier inventory; the pre-existing release-only compositions remain in ordinary manual and Full Release Validation plans. - Hourly main includes the complete tooling inventory through the existing compact planner within its 77-row Node cap. The explicit release-only runtime compositions stay in ordinary manual/release validation unless a PR selects their tests. A successful PR merge-ref result proves its tested tree, not later main revisions; the existing gate aggregates only the jobs actually selected for each revision.
- Doctor session and cron tests keep the two longest SQLite files in separate shards, with the remaining files together. Each shard runs complete files under the commands config's existing execution policy. The three owners cannot share a compact job; Blacksmith gives them 8-vCPU placement and its 200-second admission target. The two new owners charge the observed file body plus the full group overhead until the canonical timing refit has two successful main-run samples. Fixture sizes, assertions, and process isolation are unchanged.
- Direct measurements for generated hosted stripes own their packing weights just like native groups. GitHub runs use hosted measurements and hybrid attempt-one runs use Blacksmith measurements; only unmeasured children divide the parent estimate. This keeps an expensive measured child from being packed beside extra work under an artificially small estimate while preserving its test partition, worker pins and per-child file execution policy.
- Changed-plugin jobs run their independent config envelopes serially, including configs whose inner process chunks also run sequentially. The existing worker policy can then use the detected CPU budget without reserving capacity for overlapping plans; multi-target jobs retain their separate concurrency policy.
- The sixteen tooling stripes use measured file weights to keep the Git owner, managed-process, worker-artifact, and transform-cache proofs in separate jobs. Go setup follows explicit selected files first. For whole-config tooling plans, only the ordinary tooling config owns the docs i18n Go tests; isolated and Docker catalogs do not. Unknown historical configs retain their original Go setup. The shard containing the unified declaration compiler fixture retains the existing larger runner; hosted splits carry that placement through each child's selected files before packing. Its synthetic compiler graph uses the existing 1,024 MB heap override, while production full builds keep their resource guard. Git owner cases each own an isolated checkout and process tree, so their real timeout and cleanup checks overlap at most two cases per runner. Worker transform proofs use a separate suite with its own fixture lifetime; late child cleanup cannot remove another suite's inputs. These changes preserve the test cases, deadlines, and process isolation; measured admission determines the job count within the existing hybrid row cap.
- Blacksmith numbered tooling bins request the 32-vCPU class after packing while retaining their logical runner classes, names, file inventory, serial project execution and two-worker pins. Tooling files use the shared worker scheduler; Docker helper fixtures retain their separate serial config. In run 33689551111, the three slowest tooling bins received two CPUs and 8 GB of memory; their 314–342-second bodies set an eight-minute non-Windows wall. The larger request addresses that host-capacity mismatch without adding shards. Hosted and hybrid tooling placement is unchanged; the request needs native timing proof before claiming the eight-minute target.
- For ordinary UI jobs,
RELEASE_ONLY_UI_TEST_FILESdefers nine exhaustive Control UI matrices and tours from ordinary canonical PR and main plans:app-sidebar.stress.browser.test.ts,board-fixture.e2e.test.ts,chat-attachment-menu.e2e.test.ts,chat-mobile-bubble-margin.e2e.test.ts,chat-session-entry.e2e.test.ts,github-link-hovercard.e2e.test.ts,native-embed-settings.e2e.test.ts,settings-layout.e2e.test.ts, andtheme-muted-contrast.e2e.test.ts. A directly changed file stays selected on its own PR or push; source-owner and directory changes do not widen this tier. Full manual CI and Full Release Validation keep the complete canonical UI configs. Forks, unknown changed-path inventories, and historical compatibility targets also retain complete coverage. Focused board, composer, media-spacing, hovercard, settings, and contrast siblings stay in ordinary CI. The planner narrows discovery before the existing three UI shards and eight weighted Control UI E2E shards. Ordinary E2E rows use three parallel workers while private-server and runtime-budget projects keep one; complete manual inventories retain twelve rows and two parallel workers. The separate real-Gateway and browser-extension jobs keep their own coverage and scheduling. - The same
RELEASE_ONLY_UI_TEST_FILESselection also feeds the separate real-Gateway job:cron-duration-save,desktop-resize,control-ui-automation-management(QA Lab),quota-reset-status,session-pr-reader-lifetime,chat-collaborator-scroll,mcp-app-conformance,usage-sessions-owner-attribution, and QA Lab'scontrol-ui-openclaw-delegation,control-ui-media-transcript, andsession-host-command-staterun in full manual CI and Full Release Validation'snormal_cichild. Their complete browser/Gateway compositions remain in the canonical prebuilt config, with unchanged assertions and per-project worker policies; faster mocked and owner-boundary siblings stay in ordinary CI. The shared path inventory preserves all other real-Gateway and QA Lab files. The desktop transport bootstrap retains its separate proof; historical targets without the prebuilt config keep their original explicit file list. Proof-job admission is unchanged: ordinary PRs already omit the real-Gateway job, and main-shaped plans retain directly edited test files. Browser-extension routing is unchanged. Real-Gateway preparation uses the existing runtime-only full build (OPENCLAW_BUILD_PRIVATE_QA=1 OPENCLAW_RUN_NODE_SKIP_DTS_BUILD=1 pnpm build);build-artifactsretains SDK declaration generation and validation. Keeping preparation independent avoids waiting for that job's later artifact checks. The current planner emits two real-Gateway rows from that same selected inventory: existing serial fixtures plus selected standalone companions, and the remaining audited parallel fixtures. The companions share the existing two-worker phase after serial execution, without another bundled preview. Prebuilt invocation previews and the platform-family fixture reuse the validated canonical Control UI bytes. Default mocked Gateway hellos receive that artifact identity, preserving explicit overrides and build-skew validation; ordinary local previews still build private assets. The real node/SSH desktop resize bootstrap is release-only and follows the same selected inventory; full manual/release runs and direct desktop-spec edits execute both carriers once, while frequent resizing and security-owner tests stay in ordinary CI. The planner owns placement and consumes the same parallel-eligibility allowlist as Vitest scheduling. Frozen targets and older planners retain one complete row; ordinary UI shard counts and runner labels stay unchanged. - Control UI browser projects greedily pack discovered test files using committed per-file timings, then cold-start basename hints, then source byte size. Bundled files default to at most two workers, bounded by the shared worker limit and local throttling; explicit Vitest CLI worker overrides remain available. Private source servers, real Gateways, and runtime-budget tests stay in a single-worker project; local whole-suite runs finish bundled work before starting that project. Per-file overhead is refitted from serial invocations; parallel invocations still update their file weights. Timing keys supply weights only: discovery still determines the complete test inventory, including new files and files without measurements.
- Broad browser, QA, media, and miscellaneous plugin tests use their dedicated Vitest configs instead of the shared plugin catch-all. Include-pattern shards record timing entries using the CI shard name, so
.artifacts/vitest-shard-timings.jsoncan distinguish a whole config from a filtered shard. - The browser-extension Chromium bootstrap command prepares
qaRuntime. It builds the native-host and relay JavaScript and runtime assets; the separate artifact job owns Control UI and plugin SDK declaration validation. The real Chromium flow and its assertions are unchanged. - Node E2E shards reuse
qaRuntimethroughOPENCLAW_E2E_USE_PREBUILT_DIST. When the packed@openclaw/aiconsumer test is selected, the CI shard owner checks every declaration entry in its package manifest and runs the typed AI build only if an entry is missing, before admitting workers. Generic Vitest admission never repeats this repair: local prebuilt E2E runs require caller-prepared artifacts, and ordinary E2E setup already builds the typed AI package. Unrelated tests and complete typed packages add no build;OPENCLAW_E2E_SKIP_BUILDstill leaves preparation to the caller. - The browser native-host launch test is a separate POSIX E2E case. Linux
build-artifactsruns it explicitly after building or restoring dist, usingOPENCLAW_E2E_USE_PREBUILT_DIST=1so the test cannot start another build. Its JSON report must contain exactly the named passing assertion in the expected file, with one passed test and zero failures, pending tests, or todos; missing artifacts, skipped tests, and absent results fail. The workflow step skips only when a frozen historical checkout lacks the test file: that is unavailable historical proof, not coverage. Current checkouts with a missing file still fail. Changes to the case, its installation fixture, or its relay-key fixture select the artifact job even on test-only diffs; unrelated browser unit tests stay build-free. Manual CI uses the same artifact step, independently of the release-only plugin sweep. - Linux Node shard jobs persist Vitest's filesystem module cache through the upstream Actions cache API. On Blacksmith runners, official cache actions use Blacksmith's colocated cache backend instead of GitHub's, so cache entries are backend-local even when their keys match. Blacksmith CI shards are restore-only and unpack the protected Blacksmith seed into isolated runner-local roots. While the GitHub-hosted outage backend is active, every
checks-node-*test shard,checks-ui, the ordinary shardedchecks-ui-e2ejob, both fast contract matrices, and the Vitest-runningchecks-fast-coretasks restore a separately published immutable transform seed from GitHub's backend. The composite action's single default-offrestore-test-cachesinput keeps the expansion easy to disable without changing cache keys or writer policy; mixed fast-core rows enable it only for tasks that invoke Vitest. The real-Gateway UI job does not restore these test caches; native and Control UI i18n lanes do not invoke Vitest. The hybrid planner profile uses matching key contracts across two backend-local archives: attempt-1 Blacksmith rows read the Blacksmith seed, while hosted retries can read only a separately published GitHub seed. Ordinary CI jobs remain restore-only; the separate trusted warmer owns protected backend-local cache publication. The non-cancelling warmer follows main, skips docs-only pushes, runs daily, and accepts manual or repository dispatch on main. Dependency publication and code warming serialize independently per backend, platform, and ref, so a new dependency seed can publish while a previous build is still running. Compiled Vitest workers remain in the full Linux code job, prepared and saved before native SDK compilation changes resolution topology; dependency jobs do no compiler work. The ordinary Linux code row followsOPENCLAW_CI_RUNNER_BACKEND. Hybrid mode also publishes a bounded hosted seed using seventeen files across the real CI-routing, plugin/channel contract, and UI package configs. It runs the same direct pnpm entrypoints and project concurrency as hosted consumers; the Node UI keeps the shard runner. Both Linux code rows install the checksum-pinned Bun fork and collect the same seven UI seed files on Bun through the normal Vitest launcher. The shared shard environment owner places Bun transforms in its separate runtime cache leaf. Node and Bun UI collection use the ordinary UI heap environment, while heavier Node groups retain the 8 GiB ceiling. The project, shard, extension-batch, and direct test wrappers resolve a shared cache root into the same config-owned leaf, independent of config order or single versus mixed-config invocation. Concurrent runs of one config borrow separate writable slots until their process groups join. An explicit cache path remains a caller-owned leaf. The shard runner keeps separate Node and Bun roots and clones both restored seeds before concurrent plans start. It omits the full build and broad agent/runtime collection; imports reached only by other files or test bodies can remain cold. A successful Blacksmith publication does not populate GitHub's cache backend. The Linux row launches each selected shard/config envelope through the normal runner in a fresh child process with concurrency one and--testNamePattern=(?!). Non-UI collection keeps the complete Node seed and also collects the existing Bun-compatible inventory through the same runtime partition owner as ordinary CI. Hosted warming adds only the compatible files already present in its bounded tooling seed. The non-UI runtime selector admits the exact collection filter; other test-name filters and unsupported flags retain the conservative Node route. Collection preserves the include patterns, environment, compiled imports and per-file cleanup while reusing stable config-owned cache leaves. Test bodies run in ordinary CI; imports reached only inside those bodies can remain cold until that run. The warmer finishes every selected envelope and saves the content-keyed transform and compile caches even when collection fails, then reports the failure after the cache saves; ordinary CI shard execution remains fail-fast. This prevents config-global state from leaking, avoids expanding filtered shards into whole configs, and retains transforms produced by the previous child. Setup computes the transform-input fingerprint once when transform caching is enabled and passes it to restore and generation validation; disabled caches do not scan the checkout. The fingerprint clears incompatible lockfile, package, tsconfig, and Vitest-config generations. After both runtime producers join, the final Node UI collection owns the single pruning pass. Before publishing, the trusted warmer scans and prunes the combined transform cache to 75% after it exceeds 2 GiB, and the Node compile cache to 75% after it exceeds 1 GiB. Consumer jobs never prune the restored seed. Vitest hashes module id, source content, environment, and resolved transform config, so ordinary partial source changes keep unchanged entries warm while changed modules miss safely. Coarse restore prefixes bridge workflow runs; normal Actions cache LRU and inactivity eviction bound old immutable archives. - Node shard jobs record resource snapshots before execution and after worker cleanup: CPU model, host load, free memory, Node’s available/constrained-memory estimates, and Linux CPU/memory/I/O pressure when available. Node reports a zero constrained-memory value when a constraint is unknown or absent; it does not mean zero RAM. Host-wide pressure can explain capacity differences between CI and a replay, but does not identify the cause of an individual test failure.
- Trusted Blacksmith Linux Node jobs restore root
node_modules, workspace importer trees (including plugin-local versions and links), and the workspace-local pnpm store from one immutable upstream Actions cache, which Blacksmith transparently serves from its colocated backend. Pnpm imports with hard links where the filesystem permits, and keeping the complete installed tree and store in one archive preserves those links. Pnpm's metadata cache lives beneath that same archived store root, so restored installs can verify supply-chain policy without depending on the producer's home directory. Source postinstall and build preparation leave pnpm-owned dependency trees intact. The key includes an explicit archive format, runner OS and architecture, the exact Node patch, and the semantic install-input fingerprint; there are no stale-prefix fallbacks. Frozen installs select tracked manifests from the dependency lockfile's importer and local-link graph, so unrelated release tools and test fixtures do not invalidate the archive. Mutable installs, custom pnpm hooks, local-file dependencies, and unsupported lockfile shapes retain the conservative tracked-manifest inventory; frozen pnpm reconciliation still validates every restore. Manifests are canonicalized before hashing. The repository-ownedopenclawmetadata block and non-install scripts are excluded because pnpm and the audited direct root hooks do not read them, so runtime schema, publication metadata, formatting, and ordinary test/build script edits keep the dependency tree warm; unaudited lifecycle-hook drift fails closed until its source inputs join the fingerprint contract. Dependency, package-manager, hook-source, and lockfile changes always select a new immutable archive. Every exact restore runs frozen offline pnpm reconciliation, so an unchanged archive validates without registry access or importer relinking. If reconciliation fails, setup first clears every importer tree and rebuilds it offline from the restored store, then clears both modules and store and retries from the network rather than serving a partial tree. Setup then disables pnpm's redundant pre-run dependency check so install and frozen reconciliation remain the only dependency writers; shard commands must not launch concurrent implicit installs. The separate trusted warmer publishes the toolchain and exact dependency archives immediately after setup succeeds, before build and transform warming; preflight and downstream CI jobs are restore-only. Canonical pushes and same-repo pull requests opt into exact restores only on actual self-hosted runners, including hybrid attempt 1. An exact miss automatically falls back to the coarser pnpm store cache. Manual CI dispatches, fork pull requests, hosted lanes, and hosted retries use only that store cache. Cache restore/save failures are optimization misses rather than correctness failures, and normal branch scoping, LRU, and inactivity eviction bound obsolete archives. The former mutable dependency StickyDisk path was retired after repeated successful writers acknowledged commits that later runs still restored as empty filesystems. - Node shard and build-artifact jobs also restore Node's portable on-disk compile cache through immutable Actions caches. In GitHub-hosted outage mode, the hosted Vitest lane set above restores the separately published GitHub test-scope archive alongside its transform seed. Build, QA and test orchestration share the protected seed populated by the trusted warmer's full build and test collection. The existing
testcache namespace is retained for already-published archives; there is no separate build-only publisher. Ordinarybuild-artifacts, QA and test jobs only restore caches. PR and ordinary test jobs only read protected snapshots, so feature-branch bytecode never enters the shared seed and PR traffic creates no cache archives. This reuses V8 bytecode for Node-loaded orchestration, build tooling, and external dependencies when runtime identity and module paths relative to the cache match, including when only part of the source graph changes. Portable mode does not guarantee reuse across arbitrary checkout layouts; Node validates exact runtime partitions and source bytes internally. A maximum-size 2 GiB transform archive costs roughly 15–20 seconds to restore at about 125 MB/s; measured fast-contract transforms are roughly 21 seconds against an approximately 8-second restore, and broader cold imports reach roughly 100–143 seconds. The optimization should be reverted if measured savings fall below restore cost. Ordinary Vitest runs preserve the configured Node compile cache. Vitest disables bytecode caching in workers and their child processes for V8 and custom coverage providers; explicitNODE_DISABLE_COMPILE_CACHE=1disables caching for the entire invocation. - Both Linux code-warming rows publish the native SDK declaration archive to their own backend before test collection; the full Linux row also publishes it before its build. Hosted lint stripes and the dedicated package-boundary lane restore the matching archive from their runner's backend and validate native compiler inputs and complete outputs before reusing it. Hybrid mode maintains that hosted seed without duplicating the full Blacksmith build. After saving, the warmer removes its native SDK output so the subsequent packaged declaration cache describes the same tree as ordinary build consumers. Package-store contents and pnpm store-location bookkeeping are not compiler inputs; installed dependency bytes, resolution topology, explicit inputs, and compiler identity still invalidate stale declarations. The all-Blacksmith profile retains its read-only sticky-disk path. The Control UI and UI E2E jobs share a Linux Playwright Chromium archive keyed by the exact pinned Playwright version. The independent dependency publisher also runs one standard hosted
macos-15row that installs dependencies and publishes only the pnpm store after an exact miss. It uses the existing OS, architecture, Node-version, package, and lockfile key; macOS CI remains restore-only. This row leaves exact dependency, build, transform, and compile caching disabled and runs no build or test warming. Pnpm owns pruning, while the before/after disk usage and saved archive size expose retained content; pruning does not impose a content-store size bound. - The build-artifact and Docker seed jobs restore the protected full-build cache through the shared Node setup action. Full, package, and
ciArtifactsbuilds sharescripts/write-plugin-sdk-entry-dts.ts: it stages the canonical public/privatetsdownSDK declaration groups and caches each group independently. A hit restores into fresh staging; both cold and cached generations must contain every selected SDK entry and a complete relative declaration closure before publication todist/. The SDK cache never adopts declarations from livedist/. Local plugin lint and package-boundary compilation use independent native declaration trees, not packaged declarations; see declaration ownership. The built Doctor plugin-index proof reuses that exactdist/output instead of invoking the E2E harness's fallback TypeScript build a second time. - Full and package builds separately cache the AI, workspace-package, and remaining unified declarations. The unified runtime always rebuilds JavaScript before the shared declaration owner stages one base group and five plugin groups; the later SDK stage uses the same owner for its two groups. Both stages restore unchanged groups into private staging and compile only misses. When CPU and available memory cover the two largest compiler heaps plus per-child headroom, up to two nonempty, independent compiler stages can run together; other plans remain serial. They share one before/after input snapshot and validate every selected entry, successful compiler receipt, relative declaration edge, and shared-chunk owner before publication, including after package preparation clears
dist. Conflicting shared bytes or changed consumed inputs fail before live declarations are written; cache records refresh only after successful publication. Canonical DTS configuration enables TypeScript stable type ordering so an unrelated literal allocation does not reorder an unchanged exported type in a rebuilt group. Full publication prunes obsolete declarations while preserving signed app bundles and Control UI assets; SDK publication owns only its flat entries and preserves other groups’ shared chunks. The AI and workspace-package steps also rebuild JavaScript on cache hits. Implicit and explicit declaration-enabled builds share those seeds; runtime-only profiles use an uncached graph and cannot publish declaration generations. Each declaration group hashes ordinary source bytes from its successful compiler Program, so edits to existing unconsumed tests, UI sources, and workflows retain its cache hit. The exact checkout-local.cache/vitestscratch root is excluded from resolution discovery; explicitly consumed inputs, installed aliases, and adjacent or nested paths still invalidate normally. It still validates inherited configuration, generator and package/plugin metadata, compiler identity, and resolution topology; consumed declaration dependencies and new resolution candidates invalidate the generation. Workspace-package changes conservatively invalidate the AI and package declaration caches. The protected warmer publishes build archives immediately after a successful full build, before unrelated test warming and pnpm maintenance. Every warmer attempt gets a new immutable archive key; coarse restore prefixes supply prior groups, and per-step signatures remain the sole content-validity owner. Rebuilt groups replace their complete owned cache trees, including obsolete bytes whose previous record is missing or invalid. GitHub's cache quota and inactivity eviction bound old generations; identical warmer attempts can publish separate archives. The weekly Node 26 minimum lane instead publishes a 14-day artifact after successfulmainruns and restores only artifacts whose immutable producer identity resolves to that workflow onmain, avoiding quota churn without allowing PR code to write a shared cache. Private-QA declarations are never persisted in Actions caches because cache namespaces are not confidentiality boundaries. check-additional-boundariesruns the complete supplemental guard list (scripts/run-additional-boundary-checks.mts) with four concurrent child processes and per-check timings. Its 20 checks retain individual failures, deadlines and process cleanup. The shared four-rule focused scan runs once across all source roots; the narrower public lint commands remain available. Prompt snapshots run in their separate lane. Package-boundary compile/canary work stays together, and runtime topology architecture runs separately from the gateway watch coverage embedded inbuild-artifacts.- On the 32-vCPU self-hosted build runner, Gateway watch, channel tests, and the core support-boundary shard start together inside
build-artifactsafterdist/anddist-runtime/are already built. GitHub-hosted fallback runs keep Gateway watch serial so low-core contention cannot consume its readiness deadline. Full Node builds then verify Discord component attachment filenames through a serial public Gateway message action, checking the built revision and retaining the named-test JSON result; frozen targets that predate the case explicitly report unavailable proof. Both paths then run the two built TUI PTY artifact canaries alone.
The Discord proof uploader runs only after its producer step completes with success or failure. It preserves diagnostics for a real proof failure and still requires the declared files. When an earlier failure or cancellation skips the producer, or cancels it before completion, the uploader stays skipped; that coverage remains unrun.
PR owner plans first use the existing compact rows, then split work toward a 150-second test budget. If those splits exceed the final Node matrix cap, the planner retains the compact selected-owner rows, including plugin work. This keeps the selected files, configs, worker limits, and process owners intact; broad PRs may have longer rows instead of failing preflight solely because of splitting. Dist descriptors do not consume the Node row budget. If the retained plan still exceeds the cap, only changed-target chunks are partitioned by build and concurrency requirements, then balanced by predicted seconds into the remaining rows while preserving every selected target. Plans that cannot fit those separate policies or whose other owners already fill the cap still fail preflight.
Explicitly selected plugin tests retain their canonical config, native-loader isolation, worker policy, and group timing. The release-only switch controls the full plugin sweep, not the availability of its owner metadata; unrelated PRs still do not acquire that sweep.
The aggressive PR smoke selects two complete existing files:
src/config/io.load-async.test.ts for cold configuration loading and metadata
admission, and src/plugins/loader.runtime-registry.test.ts for plugin loading
and registry lifecycle. They retain their canonical configs and assertions and
occupy at most two Node rows. The full-selection kill switch retains the previous
six-file smoke inventory, including Gateway serving, config compatibility and
migration, and message-tool delivery. All six still run on hourly main.
The PR_EXEMPT_RUNTIME_TEST_FILES inventory in
scripts/lib/ci-proof-test-inventory.mts keeps measured slow integration tests
out of unrelated canonical PRs and exact-head PR fallback dispatches. Hourly
main and Full Release Validation select these complete files through their
canonical owners: normal CI owns core, UI, and tooling tests; Plugin Prerelease
owns the complete extension runtime inventory on its hourly schedule and in
the release campaign. Standalone manual CI retains its core inventory; run
Plugin Prerelease for extension coverage. Hourly main includes the complete
tooling inventory so both deferred tests and protected regressions retain their
canonical owners after PR selection narrows.
Direct test edits and selected subject owners opt files back into bounded PR
plans through the existing changed-target owner and transitive import graph. Missing/deleted tests and broad
unresolved inputs alone do not enable the inventory. Reduced groups retain
distinct timing identities and the original configs, prerequisites, process
isolation, worker limits, cases, and assertions. This tier is separate from the
older release-only runtime inventory; the tooling family now also runs hourly.
The inventory owner enumerates only audited paths that are files in the selected
checkout. Deleted paths need no runnable owner; every live entry still requires
exactly one hourly and release owner. New and renamed paths participate in changed-owner selection immediately;
separate auditing governs their inclusion in this explicit deferred inventory.
Choose complete integration, process, corpus, or migration suites from measured
PR runs, while retaining their focused owner tests on PRs. Check recent product
fixes, including regressions moved by test splits, before adding an exemption.
Hourly GitHub-hosted plans reuse the existing complete compact inventory and
serial tooling packer. Complete file estimates use the hosted cost scale; Blacksmith
process observations stay with Blacksmith. Packing preserves each child group,
its two-worker limit, and its timeout while keeping the complete Node matrix
within the 77-row main-tier cap. Tooling groups that execute on the same hosted
runner may share a row while retaining the strongest original capacity owner.
Explicit source watches live in
scripts/lib/ci-policy-test-watch.mts, including dynamically launched workers
and scripts that the import graph cannot discover.
Explicit policy watches retain their matching tests in PR CI even when the broader
tooling or runtime suite is deferred. This includes wrapper dependency checks,
Gateway client callsite scans, and upgrade-survivor package checks. Unrelated
deferred tests stay excluded. Changed test files that no Vitest config routes,
such as skill node:test and Python suites, are never direct targets; the watch
for their Vitest wrapper must cover those test files too, or test-only edits run
in no PR job.
Aggressive PR selection keeps changed tests and direct runtime import consumers, including package and SDK aliases. A changed module with fewer than 20 direct importers also selects tests at depth two; hubs stop at direct consumers. Changed packages also select direct package-specifier test consumers; that pass does not restart traversal through package readers and bypass the module hub cutoff. Erased type imports remain owned by typechecking.
Same-directory tests are selected only when the directory contains at most 30
test files. Larger flat directories use existing explicit owner mappings and
name-prefix siblings (foo.ts selects foo*.test.ts), without recursively
selecting child directories. Existing fixture, generated-input, and non-import
policy owners remain explicit. Config changes keep their planner/guard owners
instead of selecting the entire config inventory. Protected/deferred tests follow
these same rules, rather than opting in entire owner areas.
The preflight job summary lists every selected file and its selection rules.
Set repository variable OPENCLAW_CI_NODE_SELECTION=full to restore the previous
PR selection immediately. Hourly main and ordinary manual/release plans keep
their complete inventories regardless of this variable.
Source-module edits also select six explicit non-import guards: PR wrapper source closure, wrapper provisioning, eager import closure, updater swap-fixture dependencies, type-suppression inventory, and plugin SDK surface reporting. This conservative watch covers new, renamed, and deleted modules and new import edges, including dependencies missing from the inventory that should have named them. The existing required architecture group still checks source diffs for import cycles and topology changes.