* fix(ci): cap Testbox concurrency and idle spend * test(ci): align workflow guards with Testbox admission * fix(ci): declare optional Testbox admission inputs * fix(ci): narrow Testbox CLI rejection errors * test(ci): verify admitted Testbox workflow behavior * fix(ci): reserve large Testboxes for memory-heavy proof * fix(ci): default routine Testboxes to one hour * docs(ci): clarify Testbox profiles and total job deadlines Co-authored-by: Vincent Koc <vincentkoc@ieee.org>
30 KiB
| summary | title | read_when | ||||
|---|---|---|---|---|---|---|
| CI job graph, scope gates, release umbrellas, and local command equivalents | CI pipeline |
|
CI continues during Full Release Validation; the legacy release-priority variable does not pause workflow admission. See deferred CI recovery for runs already deferred by older workflow revisions.
Native video smoke coverage uses four shards of four providers. Each provider has a ten-minute operation timeout plus 30 seconds of test overhead; each shard has a 50-minute job budget, leaving eight minutes for setup. These shards keep full-mode video testing disabled.
Broad PRs retain their compact selected-owner Node plan when time-based splitting would exceed the 130-row matrix cap. See Node test lanes.
This page is an index. CI is documented on nine pages, one per reader job. Open the page that matches your task.
Full hybrid extension lint packs the same canonical chunks into three existing rows, sharing setup and SDK preparation within each row. Targeted plans and frozen routes retain their existing layout; see runner profiles.
Automation admission filters known no-op events before runner allocation and concurrency, keeping automation on GitHub-hosted runners.
First-attempt PR Node matrices let the scoped monitor classify failures before cancelling eligible same-repository work. Fork monitoring is read-only. Exact known hourly-main test and supported static failures can remain advisory when the PR leaves their subjects unchanged and all remaining checks finish. Canonical PR reruns let every Node matrix leg finish so inherited failures do not cancel the remaining proof needed for an explicit admin landing. Native matrix fail-fast applies only to PRs in other repositories. Main and manual runs retain complete matrices. See failure cancellation.
First-hop compatibility uses a 3,200-second container budget and a 3,500-second lane budget, based on hosted 4-vCPU measurements with a slow-host margin. The release self-upgrade job gives first-hop lanes weight two at npm limit five, admitting at most two concurrently. It allows 210 minutes for three waves of six source versions, the survivor, and setup. Authenticated update restart uses a 2,280-second container budget, a 43-minute lane budget, and a lane-specific 1,500-second command timeout. Its OpenAI/recovery chunk allows 160 minutes for the npm-serialized lanes plus setup; see release-path chunks.
For the published-upgrade regression gate, see selection and routing, runner budgets, and Package Acceptance baselines. Weekly validation is listed under Update Migration.
Updater, state-lease, SQLite identity, native-plugin, startup-trace, and updater-tooling PRs also require the published-driver update cell. One GitHub-hosted Linux job runs the latest stable npm updater against the candidate package with two synthetic agents. Its ten-minute budget and result are included in openclaw/ci-gate; the broader main/release Docker survivor remains separate.
Full main CI and cache warming are hourly by default; OPENCLAW_CI_ON_PUSH=true restores their existing per-push admission. CodeQL, Workflow Sanity, and CI's security-fast keep their existing main-push scopes. Docs-only main pushes still skip the CI workflow and push-triggered cache warming. The cache warmer publishes dependencies independently of long builds and maintains a bounded hosted seed in hybrid mode. Every admitted canonical main run exercises one published-driver × candidate Docker upgrade; ordinary manual/release validation adds the other five Docker seed lanes. QA Smoke, real-Gateway browser checks, and named process proofs retain their selected main coverage and manual/release validation. Pull requests and exact-head PR fallback dispatches run static correctness gates, owner-bounded tests, transitive import consumers, protected regressions, and a six-file runtime smoke set. Node rows target at most 150 estimated test seconds. Single files and indivisible canonical groups can exceed that target; setup, builds, and queues are separate from test time. Node shards selecting sandbox container E2E cases prepare the Docker sandbox image when the runner does not already have it. Missing or unbounded runtime selection fails preflight instead of falling back to every test. Windows, browser, Docker, QA Smoke, packaging, contract, and extension families opt in through their existing owners; individual built-process proofs have independent owner flags. Full static fallback does not widen them. The PR-exempt integration tier retains measured slow tests in hourly main and Full Release Validation, with PR opt-in when their tests or subjects change. The existing Plugin Prerelease workflow owns complete extension runtime coverage hourly and in Full Release Validation; normal CI selects affected extension owners on PRs. Windows retains its complete inventory across five measured file shards on hourly main and ordinary manual/release validation; Windows-owner PRs retain that complete inventory.
Hourly iOS retains ios-build (tests) with Rust, voice, native Access, and focused lifecycle coverage. Debug builds select their simulator before compiling, overlap its boot and SimSlim preparation with compilation, and join preparation before either focused simulator test group. Preparation failures or a timed-out join fail the job with its log; test selection and build settings stay unchanged. Current smoke builds prepare testing products once with build-for-testing; the voice cleanup and lifecycle groups separately reuse those products with test-without-building. Historical targets and other phases retain their existing build actions. Managed attachment UI/export, Watch operation, and Watch delivery UI suites retain every assertion in full manual/release validation. Main-tier simulator builds use the native architecture without indexing or verbose test diagnostics; logs and xcresult bundles remain available. A coalesced scheduled iOS cancellation can leave openclaw/ci-gate green with a notice delegating iOS proof to a later scheduled job; it does not validate the canceled revision, and the workflow can still be canceled. Genuine failures remain red. Screenshot capture runs for its own changed inputs and full manual/release validation. See scope selection and capacity for the coverage trade-off.
iOS screenshot shards, release qualification, Store Release, and its screenshot-only operation use larger hosted capacity. Screenshot capture uses stock simulators and creates and cleans up one at a time; the screenshot-only operation can validate a selected branch without signing or uploading a release. The pairing, chat, and native Overview tests retain their existing assertions and deadlines.
Eligible core-source and core-test PRs use targeted type checks when every selected path exists in the checkout. GitHub and hybrid profiles distribute the selected consumers across their existing core stripes; the Blacksmith profile checks them in the central row. Ambiguous ownership and deleted core tests keep the full type-check coverage.
Preflight passes the complete changed-path manifest between steps as a local JSON file, so large PRs do not lose test-planning inputs to Actions output or environment size limits. Frozen targets that predate this transport retain their bounded JSON output contract. Missing or invalid inputs still reject current PR Node planning.
The Testbox check workflow requests the Blacksmith 16-class for routine dispatched proof, with a 60-minute total-job deadline including hydration. The explicit high-memory 32-class workflow retains 240 minutes for memory-heavy full-suite gates. The outer GitHub deadline can terminate active SSH commands; the separate 15-minute idle limit does not extend it. PR hydration checks stay on hosted Ubuntu; individual test deadlines remain unchanged.
Full GitHub and hybrid type checks run the five core stripes independently, retaining two compiler children per job. The last four rows then each run one root-test partition serially, leaving extension tests and scripts in the central row. Narrow plans reuse four already-selected rows when available; smaller selections retain central root checking. This adds no jobs or compiler overlap. Current hybrid full runs use three hosted extension-lint jobs; targeted layouts retain six stripe identities. Trusted hybrid first attempts place both packed core-lint rows on the Blacksmith 16-class and the final gate on the 4-class to avoid serial hosted assignment delays. Frozen targets keep their earlier layout; see static checks.
Additional checks and narrow-PR dependency scans start directly after preflight. Known full compiler selections skip discovery while retaining the core graph boundary in an existing required owner; see pipeline ordering.
The extension package boundary row has a 30-minute job budget for SDK preparation, all selected plugin compiles, input-receipt validation, the required negative canary, and cleanup. Hosted four-CPU runs spent about 19 minutes in the compile command alone; one completed compile and canary but exceeded the former 20-minute whole-job deadline. Other additional-check rows retain 20 minutes. Optional hybrid hosted overflow retains this row on Blacksmith: compiled receipt archives include checkout-specific paths and links, and hosted cold runs exceeded 22 minutes. Explicit hosted overrides, retry and trust fallbacks remain available. Other jobs keep their existing admission thresholds; the hosted row total counts only the rows actually offloaded. Compiler concurrency, coverage and cache guards are unchanged.
Core lint discovers separate source and UI TypeScript projects, retaining shared ambient declarations and imported dependencies. The source project also includes src/**/*.test-support.cjs; unrelated JavaScript files are not added as roots. See local checks.
Runtime topology checks inherit the existing Go memory defaults, with caller overrides and the full architecture check sequence retained.
Android native resource preparation uses the Mermaid renderer's filtered dependency install, including optional build tooling. Pnpm retains root dependencies but omits unrelated plugin packages; Gradle still builds the assets and runs the selected native tests and lint. Historical targets keep their compatibility path.
Android phone tests use up to four isolated JVMs on Blacksmith and retain Gradle-owned cache expiry. The same four normal rows split phone tests from app lint: Wear owns Wear tests and lint plus third-party app lint, and Kotlin lint owns Play/shared lint. Normal same-repository Blacksmith runs overlap all four rows; other routes retain two. All test and lint tasks remain selected.
macOS Swift CI runs the app and independent package suites in separate native phases, retaining every test and the existing concurrency and timeout limits.
Native test builds retain coverage and source-line backtraces while omitting IDE indexes and full debugger type metadata. Local development builds keep their normal debug settings.
Short hybrid jobs use a 40-row base threshold and 45-row hosted admission limit, with unchanged coverage and Blacksmith fallback when optional work does not fit.
Additional hybrid check offloads require fresh hosted assignment evidence. Eligible PRs can move dependency, core type and topology checks; main pushes can also move lint and central types within the same hosted row limit. Artifact builds retain Blacksmith because their measured hosted tail leaves no room for the 15-minute routing objective.
Windows keeps its complete explicit test inventory in five measured project-aligned shards for main and release validation. Windows-owner PRs retain the complete family; unrelated PRs omit it.
Real-Gateway browser checks use job budgets matched to their selected runner.
Control UI CI installs the Chromium revision pinned by Playwright even when the browser cache misses. Current targets use the installer's --require-playwright-chromium mode; historical targets retain their existing installer. Browser startup diagnostics include provider, page, WebSocket, and Chromium process events to diagnose a session-readiness timeout even when it is reported only after unrelated unit work finishes.
Browser extension CI launches the installed, patched Chrome MCP dependency directly, on Node and on the pinned Bun fork.
Build, QA and test orchestration restore the same protected Node compile cache. The trusted warmer populates build tools before collecting test imports, including the same seven Control UI seed files on Node and the pinned Bun fork in both Linux cache backends; ordinary CI remains restore-only.
In-process Gateway test configs use exclusive plan admission within existing packed jobs.
Changed-owner Node rows use existing file/group timing evidence with a 150-test-second admission target and the existing 130-row PR cap. Complete files, canonical configs, and worker policies remain intact; predicted test seconds are separate from measured CI wall time.
Plugin-sensitive PRs select their owner tests, transitive import consumers, protected regressions, and policy watches. The existing Plugin Prerelease workflow retains complete extension coverage hourly and in Full Release Validation; see Node test lanes.
Roomy serial Blacksmith Node jobs use measured Vitest worker sizing, with existing hosted, frozen-target, and overlapping-plan limits.
Source-only Linux Node shards can reuse content-validated compiled workers from the protected warmer; fixed preparation costs remain separate from test execution and runner capacity.
Changed-target shards containing canonical E2E tests prepare the private-QA runtime once before launching test children. Only a successful preparation step enables prebuilt consumption, so the children reuse its JavaScript, assets, and freshness stamps instead of starting another full build.
Vitest transform-cache fingerprints exclude the generated .ci-harness checkout so CI consumers and the protected warmer hash the same source inputs. Node bytecode caching remains enabled for ordinary Vitest runs; Vitest owns the worker-level coverage safeguard described in local testing.
Transform keys also include each project's dependency optimizer directory. This prevents cached UI imports from mixing separate projects' Lit instances when a focused run and a full run share the persistent cache.
Linux PR tests use Bun for the measured compatible unit lanes and Control UI Vitest job. Full Release Validation keeps their Node coverage and runs them on Bun too; see test runtime selection. Both runtimes group uncached, non-isolated UI files by environment in batches to reduce worker restarts while retaining native shard ownership and worker budgets.
Frozen-target CI loads its Node shard planner, planning helpers, and measured
costs from the pinned workflow_sha checkout. Test discovery and execution still
use the candidate source, so current shard budgets do not replace release bytes.
The npm/ClawHub release decision treats normal CI tests, plugin prerelease, cross-OS, performance, and QA lanes as advisory recorded evidence. Artifact, install-smoke, survivor, all first-hop compatibility, pack/npm qualification, package-integrity, and target-resolution proofs remain required. Aggregators follow required inputs; identity and provenance verification still apply.
For publication, the sealed manifest supplies the SDK evidence digest,
per-package npm decisions, and any approved OPENCLAW_RELEASE_STABLE_SOAK_WAIVER
text. Explicit publisher inputs override those defaults; historical manifests
retain their existing input contract. SDK API changes still need an
operator-supplied acknowledgement, and the sealed waiver applies only while the
repository variable still holds the same text. Publishers still validate live authority,
artifact bytes, and registry state at the mutation boundary.
Flaky tests never block npm/ClawHub publication: record advisory failures and investigate their owners without waiting for a green rerun. A passing replay alone does not prove a fix. Required artifact, install, update, target, and provenance proofs remain enforced. Native app publication is fully decoupled from npm/ClawHub, GitHub finalization, and main closeout; report each platform's readiness separately. The approximately 20-minute validation and one-hour publication targets require hosted timing evidence before being claimed.
Release closeout refreshes hosted full-release shard costs with
node --import ./scripts/tsx.mjs scripts/ci-shard-timings-refresh.mts --run <ci-child-run-id>.
The generator records successful hosted job walls, including setup, in the existing
config/ci-test-timings.json store. Release keys stay separate from compact CI
spans and survive the daily refit. The hosted full planner splits measured rows
above 12 minutes after file bundling, retaining exact coverage and worker settings.
Complete split generations keep subsequent plans from recombining expensive work.
An indivisible over-budget test fails planning with its owner named; unmeasured
rows still need native timing evidence before claiming the 20-minute objective.
Full Release Validation's exact-target UI job retains the current three native shards for both runtimes. Historical compatibility targets keep their original unsharded package command; see UI job budgets.
Set the repository variable OPENCLAW_RELEASE_RUNNER_GROUP to reserve a runner
group for Full Release Validation and its artifact, validation, and reusable
worker jobs. The Release Publish parent and its dispatched publish children read
the same variable. It selects the group; it does not provision runners or increase
concurrency limits. Missing group capacity queues jobs. Shared workflows receive
an optional runner_group from their release caller, including docker-release.yml
and vercel-container-registry-publish.yml from Release Publish; docker-image-refresh.yml,
ordinary CI, scheduled performance, and unrelated reusable callers retain their routing.
Approval and credentialed publish jobs (npm trusted publishing, ClawHub, Docker)
keep their default GitHub-hosted labels, and the hourly plugin npm preview routes
only when Release Publish dispatches it.
The runner count, matrix caps, and default labels do not change.
To reserve capacity outside ordinary PR/main pools:
- Create an org runner group with Linux runners labelled
ubuntu-latest/ubuntu-24.04, plus the Windows/macOS labels used by validation. - Grant
openclaw/openclawaccess to the group. - Set
OPENCLAW_RELEASE_RUNNER_GROUPto the group name; unset it to release the reservation and restore ordinary routing.
Linux runners for jobs that set semantic-checks: true also require:
- systemd as PID 1 and an active
systemd-logindservice. - Noninteractive sudo access to check logind, enable linger for the runner user, and start that user's systemd manager.
- cgroup v2 with memory and swap accounting, delegated to the user manager.
These requirements apply to custom release groups as well as ordinary runners. The shared setup action starts the user manager and verifies a real 64 MiB scope with swap disabled and group OOM termination before checks run. Unsupported runners fail setup; provision these capabilities before assigning the validation labels, or unset the release-group override to restore ordinary routing. Setup qualifies the backend; it does not itself limit later lint or compiler commands.
Full Release Validation starts source-only children alongside artifact producers after admission and reuse selection. Candidate consumers start as soon as the candidate is verified, while npm qualification and independent validation can continue; see the release procedure.
Release-dispatched validation children add one best-effort hosted receipt job
each, up to seven per full campaign and none for ordinary PR/main CI. It retains
job results independently of parent completion. With reuse_evidence=true, each
dispatch checks bounded prior receipts for its exact target and inputs, including
candidate bytes, and adopts only verified successful children. Other roles still
dispatch. Failed, cancelled, and active parents can supply green children; the
current parent seals and revalidates each immutable selection. Discovery adds no
jobs or Blacksmith registrations and falls back to fresh work on a miss.
Auto-reply reply tests run files in parallel with two workers per compact group. Their planner uses separate parallel timing identities; until those have measurements, serial group costs are divided by the effective worker count, with single-file groups retaining their full cost.
The Gateway isolated/database-worker cohort keeps its two-worker budget, including roomy serial Blacksmith and hybrid jobs, to leave cold startup headroom within existing test deadlines. Other packed groups retain their existing caps.
Commands tests share the existing worker budget across independent files. The Doctor session SQLite cases are split by operation while preserving the complete repair and recovery coverage; see shard weights.
The complete startup corpus uses eight state test files so existing workers can share its release/config matrix. Its explicit fallback prepares the runtime once and uses up to four workers, capped by available CPU parallelism; historical frozen targets retain their legacy process layout with CPU-bounded admission.
| Page | Read it when |
|---|---|
| CI pipeline jobs | The job table, the fail-fast order, and the Control UI size budgets. |
| Watch a CI run | Wait on one pull request head, recover a stuck run, and pass the evidence gate. |
| CI checkout ownership | Shared checkout anchors, fetch retry budgets, and trusted action policy. |
| CI scope and routing | Why a job did or did not run: changed-scope detection and manual dispatch. |
| CI runner classes | Trust-based runner routing, preflight and ratchet admission, Blacksmith classes, and runner backend modes. |
| CI capacity and shard weights | Runner registration, bounded PR concurrency, and measured shard packing. |
| Release validation workflows | Full Release Validation, live and E2E shards, Package Acceptance, install smoke, Docker E2E, and Plugin Prerelease. |
| Scheduled and maintenance workflows | OpenClaw Performance, QA Lab, CodeQL, the maintenance jobs, and ClawSweeper activity forwarding. |
| Local checks and Testbox | Reproduce a lane locally, keep the shrink-only ratchets, and run Crabbox or Testbox proof. |
Where each section moved
Every section heading from the previous single-page version keeps its anchor here, so an existing link such as /ci#pipeline-overview still resolves. Each entry points at the page that now holds the content.
- Pipeline overview
- Fail-fast order
- Control UI size budgets
- Watching pull request CI
- Recover an existing PR run first
- PR context and evidence
- Checkout ownership
- Scope and routing
- Measured shard weights
- ClawSweeper activity forwarding
- Manual dispatches
- Windows Testbox Probe
- Runners
- Blacksmith runner capacity
- Runner backend modes
- Runner registration budget
- Surface ratchets
- Local equivalents
- OpenClaw Performance
- Vitest paired benchmark
- Full Release Validation
- Live and E2E shards
- Package Acceptance
- Jobs
- Candidate sources
- Suite profiles
- Legacy compatibility windows
- Examples
- Install smoke
- Local Docker E2E
- Tunables
- Reusable live/E2E workflow
- Release-path chunks
- Plugin Prerelease
- QA Lab
- CodeQL
- Security categories
- Platform-specific security shards
- Critical Quality categories
- Maintenance workflows
- Dependency Audit
- Docs Agent
- Duplicate PRs After Merge
- Local check gates and changed routing
- Config baseline count ratchet
- Testbox validation
Related
Ordinary PR iOS smoke keeps app and test-bundle compilation while selecting its two simulator groups by source owner. With neither group selected, simulator preparation is skipped. OPENCLAW_CI_IOS_SIMULATOR_FULL=true restores full PR execution; selection and routing documents the owners and full hourly/release coverage.