qaRuntime intentionally omits declarations, but private-QA CI marks its output
prebuilt and skips the normal typed E2E setup. The packed AI tarball then
contains no declarations and fails its external TypeScript compilation.
Check the selected consumer's manifest type entries and run the existing typed
AI build only when an entry is absent. Share this preparation between local
Vitest admission and the CI shard owner, before concurrent readers start.
Keep explicit skip-build, complete typed output, unrelated shards, and the
sparse historical workflow adapter unchanged.
Proof: cold qaRuntime reproduced TS7016 locally and on Node 24.19.0 Testbox;
the repaired concurrent shard passes with all 17 type entries. Linux typed
AI compilation took 4.7s; the unrelated shard passed in 2.88s with no typed
build and no declarations. Median unrelated preparation over 1,000 local
calls changed from 0.001ms to 0.009ms, with zero added build commands.
Validation: check:changed passed; 162 focused tests and all 203 shard-runner
tests passed. New regression file: 8 tests, 2.19s pnpm wall; complete modified
shard suite: 6.98s pnpm wall. Independent review's historical-adapter concern
was rejected after executing the sparse adapter with the prebuilt flag and
without the current prerequisite module.
Canonical full private-QA pnpm build passed (665.84s), followed by the ordinary
non-prebuilt packed-package E2E consumer (58.34s wall, 16.49s test duration).
Admit two source-only singleton processes on measured Linux capacity while preserving file isolation, worker ceilings, cache ownership and focused failure draining. Keep serial fallback prices until exact parallel measurements qualify.
The constrained 199-case comparison improved from 184.40s to 147.31s; three qualified replays completed in 146.13-149.34s with at most 4.51 GiB aggregate memory. Runner classes and row counts remain unchanged. Linux owner, planner, changed-file and Knip proof plus focused fixture corrections and P2 review completed.
Finish ordinary failed packed configs when the existing continuation flag is
selected. Publish native completion counts only after selected plans, child
processes, and cleanup finish, so consumers can distinguish assertions from
partial or interrupted execution.
* fix(ci): overlap ordinary test plans around exclusive barriers
* test(ci): await scheduler admission signals
* fix(ci): keep aggregate scheduling out of metadata discovery
* fix(ci): size workers after final plan admission
Pin the reviewed Bun fork with the corrected WebKit string bounds checks, remove the CSS tokenizer workaround, and reduce native worker, fallback, browser and fixture overhead while preserving all assertions and Node GC coverage.
The complete matched Linux comparison preserves 1,494 files, 25,127 passing cases and four skips. Bun reduces summed shard command time by 3.60% and the slowest shard by 1.18%. The unsharded run passes within the unchanged 10 GiB sampled-PSS guard. Full release verification retains both runtimes.
The shard relay labeled every child stdout and stderr line, hiding Vitest workflow commands from the Actions parser. Forward annotation and group commands at column zero while preserving labels on ordinary output. Cover both streams, split commands, escaped payloads, and final partial lines through the shard runner.
* fix(ci): reuse JavaScript caches across test layouts and runtimes
* ci: report memory and pressure around test shards
Record Node memory estimates and Linux host pressure before shard execution and after worker cleanup so cache and replay timings retain resource context. Keep scheduling, cache policy, test inventory and deadlines unchanged.
The original CLI custody timeout remains unexplained after exact-source native replay and bounded phase traces. These diagnostics do not claim a fix. Existing shard/warmer suites pass 121 cases in 6.85 seconds; scripts types, typed lint, docs sanity and independent P2 review pass.
* test(codex): reuse provider schema runtime and complete interrupt fixtures
* ci: defer exhaustive UI matrices to release validation
* test(codex): keep prepared auth fixtures with shared test support
* ci: preserve UI release selection across test runtimes
* test(backup): isolate opaque SQLite warning fixtures
Keep the warning inventory independent of ambient backup scratch while retaining all six opaque warnings and byte verification and restore assertions. Move the unchanged summary-result fixture to existing test support to keep the oversized test file below its growth limit.
The original shard replay passed. Real legacy scratch reproduced a seventh warning; isolation preserves all 130 cases with that noise present. The original CI warning text was unavailable, so its exact source remains unconfirmed.
* ci: run compatible UI tests on Bun with isolated transform caches
* perf(ui): narrow Bun tokenizer workaround and skip zero-width styles
Keep all JIT tiers enabled while preventing the ordered CSS-tokenizer runaway at its end-of-file predicate. Avoid redundant style resolution for zero-width marquee labels while preserving resize recovery and reveal cleanup.
Keep PR runtime partitions, native shards, worker budgets, assertions, and release Node/Bun coverage intact. Repair the fixture dependency URL and make the marquee regression establish keyboard focus modality explicitly.
* perf(test): split sidebar cases across UI workers
Keep the existing 484 sidebar cases in four topic entrypoints so native UI sharding and worker scheduling can share their cost. Move mock and global cleanup into the common sidebar fixture, scope footer cleanup to its cases, and preserve the complete cache-warmer collection.
* perf(ci): warm Bun UI caches and defer costly FTL compilation
* ci: use measured Gateway test worker limits
Admit up to eight workers for the Gateway isolated/database-worker cohort only on roomy serial self-hosted jobs, with a 28 GiB memory floor and the existing two-worker fallback. Preserve sibling caps, timing history, coverage and matrix budgets.
* test(ci): preserve measured Gateway placement expectations
* test: identify compiler lifetimes across PID reuse
* test(ci): respect proof tiers in Gateway policy coverage
* test(ci): retain synthetic Gateway placement ownership
Use the upstream synthetic Gateway fixture for recipient and worker-cap assertions after rebase. Keep the fixture isolated from the changing project inventory and preserve every placement and execution-policy assertion.
* test(ci): preserve Gateway caps through dense packing
Keep the upstream settled serial row, prediction, inventory, and neighboring packing assertions while checking the measured Gateway and unmeasured sibling worker limits at their group owner.
* ci: keep session history benchmark on Node
Retain the observed-failing Node process and SQLite lifecycle benchmark in the existing Node-required runtime partition. Preserve the complete test, profiles, assertions, and deadline, with routing coverage for bun-compatible and dual policies.
Co-authored-by: Peter Steinberger <steipete@gmail.com>
Use measured CI-only six/eight-worker memory tiers for roomy serial self-hosted jobs. Retain local behavior, actual-host fallback limits, Gateway exclusivity, unproven group caps, and historical timing floors. The twelve predefined Blacksmith probe samples passed; native PR CI remains a separate validation step.
Fix Gateway integration CI timeouts caused by admitting another test plan beside cold in-process Gateway startup. The shared config policy gives Gateway core, database-workers, methods, methods-isolated, server, and server-isolated exclusive plan admission.
Keep the existing packed jobs, duration budgets, runner allocations, file partitions, timing identities, and worker ceilings. Preserve canonical Gateway metadata during precise test selection and resolve scoped Vitest arguments once after inherited global flags, preventing duplicate extension partition arguments. Ordinary jobs retain their concurrency. Installed OpenClaw behavior and test timeouts are unchanged; affected CI jobs may take longer because plans run serially.
Validation: exact-head CI run 35266135431 completed green with zero pending checks; recorded runner and changed-plan proof passed 277 tests, check-changed passed, and 36 Gateway argv comparisons matched. Recorded independent review and exact-head ClawSweeper review found no actionable defects.
* fix(ci): avoid overloaded jobs when runtime test groups grow
Retain measured workload descriptors across file additions and reuse existing ordinary capacity without adding jobs. Preserve timing identities and worker ceilings while moving complete runtime groups.
* test(ci): compare worker policies from original admission
* fix(ci): keep preflight job outputs under GitHub's 1 MiB cap
Pull requests that need the full canonical Node test plan failed preflight
with "Job outputs exceed 1,048,576 bytes", which made ci-gate report every
lane as selected=missing. The compact plan lists 7,749 striped test files
explicitly, and GitHub measures job outputs in UTF-16, so the effective cap
for all preflight outputs is 524,288 characters. The manifest reached ~522K
characters on the github profile as the tracked test inventory grew.
Emit each Node matrix row's groups as gzip+base64 JSON at the manifest
boundary and unpack them in the shard runner. On current main this takes the
github-profile manifest from 522,522 to 195,410 characters and the blacksmith
profile from 450,639 to 158,571, with identical jobs and group membership.
The manifest imports the codec through the existing target-plan seam, the
trusted runner checkout lists it, and plain OPENCLAW_NODE_TEST_GROUPS_JSON
remains for the Vitest cache warmer.
Related to #139929 and #136820.
* fix(ci): load the Node test groups codec only for grouped rows
The manifest imported scripts/lib/ci-node-test-groups-codec.mts
unconditionally through importTargetPlan. An ordinary manual
workflow_dispatch whose target_ref predates the codec is not a
compatibility target, so the import threw before planning even though
such targets emit an ungrouped plan and never call the encoder.
Import the codec only when the emitted Node rows carry groups, and keep
the clear failure when a grouped plan comes from a target without it.
The workflow-guard fixture can now omit the codec so a codec-free manual
target runs the real manifest step.
* fix(ci): preserve grouped historical test plans
Negotiate the packed group codec against the selected target and fall back to the same five-field legacy projection when the codec is unavailable.
Five-field projection adapted from #140036.
Co-authored-by: Peter Steinberger <steipete@gmail.com>
---------
Co-authored-by: Vincent Koc <vincentkoc@ieee.org>
Co-authored-by: Peter Steinberger <steipete@gmail.com>
* ci: let lint wrappers own Go resource limits
* test: allow unset Go variables in CI lint fixture
* ci: keep hybrid control jobs on hosted runners
* ci: reduce sharding for seven-minute target
* ci: bound packed test concurrency by runner capacity
Keep existing group boundaries and matrix rows while using two child slots
only on CI hosts with enough CPU and memory. Join admitted children before
surfacing unexpected failures.
Lower final push/PR Node caps to 64/120 and record declaration cache miss
reasons and whole-runner memory peaks. Retain runtime preparation ownership
and shared UI failure diagnostics.
* ci: give serial tooling jobs more runner capacity
Keep compact test groups, names, worker pins and row counts intact while
requesting larger existing runners for subprocess-heavy tooling and Docker.
Repair the CI fixtures for the current TypeScript library and named Docker
step, and clean up their focused lint findings.
* fix(release): preserve VCR mirror source digests
Transport only attestation-verified digests across secret-scanned job outputs, reconstruct immutable GHCR refs inside the VCR mirror, and add an approved mirror-only recovery path.\n\nCloses #129466
* fix(release): verify VCR recovery sources
Revalidate attestations and release-version labels before any VCR registry write so manual recovery preserves the immutable source boundary.
* test(release): keep VCR regression scoped
Leave global workflow-to-test routing cleanup for a follow-up; this PR directly changes and runs both VCR regression suites without forcing metadata-complete CI.
* fix(ci): preserve caches after warmer failures
Finish every selected cache-warm group, save content-keyed transform and compile caches, then fail visibly after the save steps. Ordinary CI remains fail-fast.