openclaw/docs/ci.md
Peter Steinberger 0af08749bb
perf(test): tier process proofs and balance Windows CI (#153950)
* perf(test): tier process proofs and balance Windows CI

* fix(ci): import Windows planner directly and align tier guards

* perf(test): share Windows preparation and pack measured projects

* fix(ci): return explicit Windows shard fields

* docs(ci): clarify Docker proof coverage after tiering

Document PR boundary coverage, main owner selection, release survivor coverage, and the proof gap for coalesced main pushes. Confirm the existing guards consume the selected Docker inventory and dispatch exact-target release CI.

* style(test): format rebased Windows workflow assertion

* test(ci): satisfy planner inventory lint contracts
2026-09-21 06:05:45 +00:00

12 KiB

summary title read_when
CI job graph, scope gates, release umbrellas, and local command equivalents CI pipeline
You need to understand why a CI job did or did not run
You are debugging a failing GitHub Actions check
You are coordinating a release validation run or rerun
You are changing ClawSweeper dispatch or GitHub activity forwarding

This page is an index. CI is documented on nine pages, one per reader job. Open the page that matches your task.

For the published-upgrade regression gate, see selection and routing, runner budgets, and Package Acceptance baselines. Weekly validation is listed under Update Migration.

Docs-only main pushes skip CI and cache warming. The cache warmer publishes dependencies independently of long builds and maintains a bounded hosted seed in hybrid mode. Docker seed, QA Smoke, real-Gateway browser checks, and named process proofs run on selected main pushes and manual/release validation. Pull requests and exact-head PR fallback dispatches retain unit, boundary, build, and mocked-Gateway coverage. Windows retains its complete inventory across five measured file shards. See scope selection and capacity for the coverage trade-off.

Core-test-only PRs use targeted type checks only when every selected test exists in the checkout. Deleting a core test keeps the full type-check plan, including the existing core stripes on GitHub and hybrid profiles.

Core lint includes src/**/*.test-support.cjs in type-aware checks through the bounded src/tsconfig.json discovery project. Other source files retain the root TypeScript project; unrelated JavaScript files are not added to this test-support project.

Android native resource preparation uses the Mermaid renderer's filtered dependency install, including optional build tooling. Pnpm retains root dependencies but omits unrelated plugin packages; Gradle still builds the assets and runs the selected native tests and lint. Historical targets keep their compatibility path.

macOS Swift CI runs the app and independent package suites in separate native phases, retaining every test and the existing concurrency and timeout limits.

Native test builds retain coverage and source-line backtraces while omitting IDE indexes and full debugger type metadata. Local development builds keep their normal debug settings.

Short hybrid jobs use a 40-row base threshold and 45-row hosted admission limit, with unchanged coverage and Blacksmith fallback when optional work does not fit.

Windows keeps its complete explicit test inventory in five measured project-aligned shards, sharing each small project's setup within one job.

Real-Gateway browser checks use job budgets matched to their selected runner.

Control UI CI installs the Chromium revision pinned by Playwright even when the browser cache misses. Current targets use the installer's --require-playwright-chromium mode; historical targets retain their existing installer. Browser startup diagnostics include provider, page, WebSocket, and Chromium process events to diagnose a session-readiness timeout even when it is reported only after unrelated unit work finishes.

Browser extension CI launches the installed, patched Chrome MCP dependency directly.

Build, QA and test orchestration restore the same protected Node compile cache. The trusted warmer populates build tools before collecting test imports; ordinary CI remains restore-only.

In-process Gateway test configs use exclusive plan admission within existing packed jobs.

Changed-extension PR jobs use measured fallback rates and a 240-second packing budget within the landed 90-row compact, 130-row PR and 70-row push caps.

Compact planning reserves the actual appended plugin rows before applying those Node matrix caps, allowing existing hosted tooling compaction to use the remaining capacity.

Roomy serial Blacksmith Node jobs use measured Vitest worker sizing, with existing hosted, frozen-target, and overlapping-plan limits.

Source-only Linux Node shards can reuse content-validated compiled workers from the protected warmer; fixed preparation costs remain separate from test execution and runner capacity.

Vitest transform-cache fingerprints exclude the generated .ci-harness checkout so CI consumers and the protected warmer hash the same source inputs. Node bytecode caching remains enabled for ordinary Vitest runs; Vitest owns the worker-level coverage safeguard described in local testing.

Linux PR tests use Bun for the measured compatible lanes. Full Release Validation keeps their Node coverage and runs them on Bun too; see test runtime selection.

The complete startup corpus uses eight state test files so existing workers can share its release/config matrix. Its explicit fallback prepares the runtime once and uses four workers; historical frozen targets retain their legacy process layout.

Page Read it when
CI pipeline jobs The job table, the fail-fast order, and the Control UI size budgets.
Watch a CI run Wait on one pull request head, recover a stuck run, and pass the evidence gate.
CI checkout ownership Shared checkout anchors, fetch retry budgets, and trusted action policy.
CI scope and routing Why a job did or did not run: changed-scope detection and manual dispatch.
CI runner classes Trust-based runner routing, preflight queue recovery, Blacksmith classes, and runner backend modes.
CI capacity and shard weights The runner registration budget and the measured timings behind shard packing.
Release validation workflows Full Release Validation, live and E2E shards, Package Acceptance, install smoke, Docker E2E, and Plugin Prerelease.
Scheduled and maintenance workflows OpenClaw Performance, QA Lab, CodeQL, the maintenance jobs, and ClawSweeper activity forwarding.
Local checks and Testbox Reproduce a lane locally, keep the shrink-only ratchets, and run Crabbox or Testbox proof.

Where each section moved

Every section heading from the previous single-page version keeps its anchor here, so an existing link such as /ci#pipeline-overview still resolves. Each entry points at the page that now holds the content.