Commit graph

3338 commits

Author SHA1 Message Date
Dallin Romney
90563ee83b
ci: allow disabling PR fail-fast with a label (#163246)
* ci: allow PR label to disable fail-fast

* ci: align fail-fast label with final gate
2026-10-01 23:27:49 -07:00
Peter Steinberger
db20d75239
ci: pin the OpenClaw Bun fork b368 prerelease (#163251)
Pin the verified b368 Bun fork for Linux CI and admit three files whose CommonJS exports, post-script argv, and native pipe-output gaps are fixed. Keep mixed, broad, and V8-specific selections on Node; preserve complete oxlint diagnostic capture with a bounded fixture buffer.

Proof: no regressions in the eleven-selection fc90/b368 Linux comparison; final fast lane passed 16,608 cases on each pin; all 103 newly admitted cases passed through CI groups. Both Node-hidden smokes passed 10/10 with zero Node attempts. Final focused Node/Bun checks, local changed-file checks, and P2 reviews passed. WebKit remains unchanged at fb1167ebf2.
2026-10-01 23:53:23 -05:00
Dallin Romney
0fa2dddedb
ci(release): pin reachable live provider models (#163177) 2026-10-01 21:05:33 -07:00
Peter Steinberger
fb338924f5
ci(ios): cache verified Apple smoke build inputs
Reuse content-verified Mermaid assets and Watch RTC static slices, and cache SwiftPM sources and binary artifacts without disabling package resolution. Fresh rust-src timestamps invalidate Cargo target caches, so fingerprint the finished Watch slice instead.

PRs restore only; trusted main push and schedule runs publish exact input keys. Frozen targets retain their cold path. Keep every scheme, target, Swift lint phase, build-for-testing action, and simulator selection.

Local Xcode 27 smoke: 185.9s cold to 89.4s warm with fresh Swift build products and isolated package support caches; all 17 Swift compilation targets retained. P2 review is clean.
2026-10-01 21:02:28 -07:00
Peter Steinberger
7fc0cfb379
ci: batch compiler planning and start guards after preflight (#163196)
Read complete compiler memberships from one native snapshot while retaining
canonical graph ownership checks and conservative full selection. Keep the
serial CLI boundary owner and provide OPENCLAW_CI_TYPE_PLAN_SERIAL=true (or 1)
as an operator fallback; unset uses snapshot discovery. No repository setting
is created.

Move existing narrow-PR guards into the preflight-ready additional matrix,
reusing their commands, setup, runner and comparison base. Preserve full and
frozen layouts, coercion deduplication and the versioned observer count.

Four fresh Linux PR-input replays preserved every output field. Compiler
planning fell from 77-91s to 17-21s; complete materialization fell from 85-101s
to 25-32s. The materializer peaked at 14.8 GiB RSS, so the existing eligible
hybrid planner uses the 16-class. Hosted and trust fallbacks stay unchanged;
no jobs or permissions are added.
2026-10-02 03:09:15 +00:00
Dallin Romney
9ce1498432
ci(release): install Chromium for A-K live tests (#163178) 2026-10-01 19:54:21 -07:00
Vincent Koc
d6411f28a1
fix(ci): keep admitted Testboxes alive and enforce current workflow limits (#163021) 2026-10-02 09:51:59 +07:00
Vincent Koc
30aec15c05
fix(ci): run full type checks when selectors are absent (#163123)
* fix(ci): run full type checks when selectors are absent

* Merge branch 'main' into fix/ci-type-selector-fallback

Co-authored-by: Vincent Koc <vincentkoc@ieee.org>
2026-10-02 02:45:56 +00:00
Peter Steinberger
2ff7b76c8b
ci(android): overlap fork PR rows on Blacksmith
Canonical fork PR first attempts already use Blacksmith, but the Android
matrix cap still required a same-repository head. Admit all four normal rows
together instead of imposing a second 9.6-10.6 minute queue wave.

Keep two-way overlap for the GitHub backend override, retries, manual and
scheduled runs, and noncanonical repositories. Preserve runner routing,
row counts, registration allowances, cache trust, coverage, and deadlines.
Update the capacity docs and evaluate runner labels alongside concurrency
for fork backend/attempt combinations.
2026-10-01 19:39:57 -07:00
Peter Steinberger
4167c4acb1
ci: share prepared SDK declarations across PR checks (#163094)
* ci: share prepared SDK declarations across PR checks

* ci: preserve author-independent SDK producer routing

* ci: avoid anchors in SDK composite action

* ci: export additional checks to downstream harnesses
2026-10-01 17:39:55 -07:00
Peter Steinberger
ba0ff407ca
ci: pin the OpenClaw Bun fork fc90 prerelease (#163101)
CI still used the 17c9 Bun fork and kept catalog-retention coverage on Node because queryObjects was missing.

Pin fc90 with verified archive and executable hashes, retain WebKit fb1167ebf2, and admit the qualified retention test through the existing runtime owner. Keep V8-specific coverage on Node.

Linux A/B passed all eleven existing selections with 44,414 passing test executions per pin and no functional regressions. The enlarged fast-unit lane passed 16,587 tests; both Node-hidden smokes passed with zero Node attempts. Focused Node/Bun tests, changed-file checks, Codex P2 review, CI, and ClawSweeper passed.
2026-10-01 16:55:10 -07:00
Mitchell Etzel
dc794c0481
fix(ci): make PR runner capacity independent of author (#160958)
* ci: plan trusted fork PRs with their Blacksmith runner profile

Trusted fork pull requests (CONTRIBUTOR and above) already run on the same
Blacksmith labels as same-repository PRs, but preflight forced them into the
logical `github` planner profile. That profile plans from empty cold-start
timing hints and skips Blacksmith capacity promotion, so these PRs ran compact
Node bins on 2-CPU runners well past the admission budget.

Trusted fork first attempts now keep the configured profile. Untrusted forks
and fork retries (which route hosted) still use `github`.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01T4y3wMpVsBQLpzEEpvAySR

* ci: give every pull request author maintainer CI capacity

Non-maintainer pull requests ran on smaller or hosted runners while CI
timing budgets assume maintainer capacity, so their compact Node shards
regularly exceeded admission budgets. Drop author-association gating
from runner routing, the planner profile, matrix caps, parallelism and
hosted-offload admission so every PR first attempt, including forks
from first-time contributors, gets the same Blacksmith routes as a
maintainer PR.

Fork retries still route hosted because forks cannot read the backend
override, and cache trust stays restricted to openclaw/openclaw.

Approved by maintainer Patrick Erichsen.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01T4y3wMpVsBQLpzEEpvAySR

* ci: keep hosted lint stripes for forks while unifying Node planning

The unified profile put fork PRs on the all-in-one check-lint and
check-test-types jobs. Those jobs are sized for the trusted sticky disks
and caches forks cannot mount, and both were canceled at the 20-minute
limit in run 36579539665 while still linting and compiling.

Split the decision: forks keep the logical github check profile (hosted
lint/type stripes), while node_runner_backend keeps the configured backend
on fork first attempts, so every PR still gets the same Node planning,
32-vCPU promotion and measured timings regardless of author.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01T4y3wMpVsBQLpzEEpvAySR

* ci: give fork first attempts the 130-row Node parallelism

The 130-row Node max-parallel still required a same-repository head and a
non-github runner profile, so fork first attempts planned up to 130 rows
but ran them 96 at a time. Key the cap on the Node planner backend instead,
which is Blacksmith or hybrid for every PR first attempt regardless of
origin or author.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01T4y3wMpVsBQLpzEEpvAySR

* test(ci): source the package helper from the published-driver stub image helper (#163027)

The cell plan test fixture stubbed docker-e2e-image.sh without sourcing docker-e2e-package.sh; #162824 removed the cell script direct source that the stub relied on, so the main hourly failed with docker_e2e_prepare_package_tgz: command not found. The stub now sources its sibling package helper like the real helper.

Refs #162824 #162965

---------

Co-authored-by: Claude <noreply@anthropic.com>
Co-authored-by: Patrick Erichsen <patrick.a.erichsen@gmail.com>
Co-authored-by: Peter Steinberger <steipete@gmail.com>
2026-10-01 16:04:01 -07:00
Peter Steinberger
3ee29f3b82
ci: pin the OpenClaw Bun fork 17c9 prerelease (#163011)
CI still used the 57fad Bun fork before the upstream sync and fork fixes qualified in baseline v3. Pin the 17c9 prerelease and WebKit fb1167ebf2, update both Linux x64 checksums, and document the matching revisions while preserving runtime admission and checksum enforcement.

Validated the actual setup action and eleven Linux Bun selections: 44,376 passing case executions, zero candidate failures, and no regressions or added skips against 57fad. Both Node-hidden package smokes passed all ten steps with zero Node attempts. Full build, package integrity, local changed checks, exact-head CI, P2 Codex review, and ClawSweeper passed. No Node-only exclusions or blocker entries changed.
2026-10-01 14:44:09 -07:00
Peter Steinberger
7eebb1f8c5
ci: reduce published-driver PR validation cost (#162965)
The published-driver update cell ran serially after build-artifacts on every PR it was selected for, adding about 13.5 minutes of critical path. PR selection is now limited to the updater, activation, managed-service handoff, and canary-emitter owners (scheduled main and release validation still run the cell unconditionally); published-install caches have main-only writers with explicit cache-mode gating and read-only PR consumers; PR runs skip initial Doctor seeding and use shorter readiness waits. PR-mode cell measured at 3m32s with a cold image-builder cache. The cleanup tests no longer touch host Docker.
2026-10-01 14:27:07 -07:00
Peter Steinberger
48746800b0
refactor(installer): share standalone shell installer policy (#162652)
* refactor(install): share standalone shell installer policy

* refactor(installer): keep clone failures terminal in adapters

* fix(installer): type standalone assembler text inputs

* test(installer): fix workflow fixture lint
2026-10-01 13:01:25 -07:00
Vincent Koc
32dc6c861d
fix(ci): cap Testbox concurrency and idle spend (#162678)
* fix(ci): cap Testbox concurrency and idle spend

* test(ci): align workflow guards with Testbox admission

* fix(ci): declare optional Testbox admission inputs

* fix(ci): narrow Testbox CLI rejection errors

* test(ci): verify admitted Testbox workflow behavior

* fix(ci): reserve large Testboxes for memory-heavy proof

* fix(ci): default routine Testboxes to one hour

* docs(ci): clarify Testbox profiles and total job deadlines

Co-authored-by: Vincent Koc <vincentkoc@ieee.org>
2026-10-02 02:16:05 +07:00
Peter Steinberger
d6e4bd48c3
build(apps): generate native protocol models at build time (#162251)
* build(apps): generate native protocol models at build time

* test(apps): include native preparation in CI build contract

* fix(apps): share cached protocol generation across Apple destinations

* fix(build): type native protocol naming helpers

* fix(build): preserve native protocol tooling consumers

* fix(tooling): scope native protocol build audit roots

* fix(build): prepare native protocol in runtime-only builds

* fix(build): preserve native preparation package and workflow contracts
2026-10-01 11:43:36 -07:00
Peter Steinberger
4f0541eb46
fix: prevent delegated work from stopping silently (#162567)
* fix: prevent delegated work from stopping silently

* fix: preserve frozen bootstrap metadata and refresh task prompts

* test: retain a foreign manager in frozen hydration proof

* test: exercise Corepack bootstrap warnings during hydration

* test: align completion fixtures with required private replies

Verify meaningful private outcomes, retained completion authority after inline waits, and one real write with no recovery replay. Start the browser fixture completion deadline when its held model is released.

* fix(node): recheck upload cancellation after opening snapshots

* test: group subagent typechecks with session ownership
2026-10-01 18:39:01 +00:00
Peter Steinberger
778e1f37ae
fix(ci): fetch boundary bases for shallow PR checkouts
Reuse the extension-lint comparison-base action for PR boundary checks. Keep its bounded exact-SHA fetch and existing fallback policy, and leave checkout refs unchanged.

A depth-one merge stays an ancestry boundary even after its parent is fetched. Validate the pinned raw first parent, then compare the two trees; retain ordinary ancestry validation for other shapes. Shallow checkout regressions cover package, core and deleted public SDK changes, including blobless base inventory hydration.
2026-10-01 10:26:23 -07:00
Peter Steinberger
4e6eb74beb
fix(ci): avoid overlay copies in published-driver updates (#162858)
The published-driver cell timed out on scheduled main runs: the Docker fixture forced tens of thousands of OverlayFS copies on the candidate tree during the managed update. The cell now stages the update on a fresh private native volume per invocation (cleaned up on exit), which brings retention from 39 s to 12 s and the whole cell to about four minutes; checksum failures print both digests and reject before the update starts. The earlier "digest-mismatch: error" log line was an action setting echo, not a mismatch.
2026-10-01 10:14:24 -07:00
Peter Steinberger
a4a78b738e
fix(update): package activation recovery rejects Bun-hosted installs (#162586)
Package recovery treated checked Bun executables as Node and dropped configured macOS SQLite selection. Carry admitted runtime facts and the shared environment choice into durable standalone recovery, preserving version-1 journals, backup custody, and separately owned service restart.

Verified focused suites on Node and Bun, real macOS custom-library recovery from a clean shell, published-driver updates and Node-free interrupted recovery on AWS, static/import checks, and independent review. CI fixture and cell-lifetime repairs retain the existing assertions and product timeout policies.
2026-10-01 09:54:24 -07:00
Peter Steinberger
779880c635
fix(ci): reuse candidate artifacts for published-driver updates (#162765)
The published-driver update cell rebuilt the candidate inside its own job and overran its ten-minute bound on the 12:46Z main hourly, cancelling the whole run. The cell now consumes the checksummed candidate from the same run's build-artifacts job (explicit digest comparison, portable across GNU and BSD tools), runs the managed update in about five minutes, and reports a command timeout as a job failure naming the active phase instead of cancelling the workflow. Scheduled-main coverage and omission on unrelated PRs are verified by the plan tests.
2026-10-01 08:04:15 -07:00
Peter Steinberger
cd8907aa0c
ci: select affected extension packages for PR boundary checks (#162640)
PR boundary selection: check only extension packages the PR's own diff can affect. When core/SDK declaration inputs change, select directly touched extension packages, a fixed smoke set of at most three broad SDK consumers, and packages that import a changed public plugin-sdk entry directly; skip transitive declaration fan-out. Hourly/schedule and release keep the full boundary check and the negative canary. Kill switch: repository variable OPENCLAW_CI_BOUNDARY_SELECTION=full restores full PR selection (unset means aggressive).

Backtest: all 10 historical PR boundary failures remain selected. 20-PR replay: modeled boundary median 7:22 -> 2:55 (conservative 4:49).

Merged past one inherited red: published-driver-update / Published driver update was cancelled at its 10-minute job timeout. It is cancelled the same way on main in hourlies 36863974207 and 36869865743, and its owner is fixing it (reuse build artifacts; timeout becomes a failure with reason). All other 67 jobs, including tooling, passed.
2026-10-01 07:36:04 -07:00
Peter Steinberger
150e60a613
fix(ci): include update cell in evidence dependencies
Include published-driver-update in the release evidence sealer's complete direct workload dependency list. The job added in #162629 was already joined by ci-gate, but its omission broke the existing evidence inventory assertion.

Keep that assertion unchanged. No job, permission, timeout, or repository setting changes.

Validation: five full-file Linux repetitions passed (90/90 test executions). Full test types, check-changed, boundary lint, both preflight shapes and P2 passed. The whole tooling config was replayed across four native shards; its separate jsdom focus and subprocess deadline failures are covered by the already-landed c3425e1d7b and 23172d8b8c repairs.
2026-10-01 07:34:02 -07:00
Peter Steinberger
25972a6a28
ci: published-driver update cell for updater and identity paths (#162629)
Adds a CI cell that installs the latest published stable openclaw as the driver and runs a managed update to the candidate built from the current revision, asserting a finished run, the candidate version, Gateway readiness, and no canary, identity, or lease warnings. It is path-gated to the updater, lease/identity, state-database-open, plugin native admission, and startup-trace surfaces and counted by the aggregate gate; on the dispatch-fallback path (checkout revision differs from github.sha) it skips with a recorded reason. Motivation: the four 2026.9.7-only update regressions (#162131, #162130, #161746, #162047) were landed by PRs whose tests exercised encoder and decoder in-process; this cell fails on the pre-fix tree. The reusable workflow checks out github.sha with read-only permissions, no persisted credentials, no cache writes, and no secrets.
2026-10-01 05:38:00 -07:00
Vincent Koc
6c3bccb6c2
fix(ci): install ripgrep for baseline ratchets (#162651) 2026-10-01 19:11:56 +07:00
dependabot[bot]
3ad238cc06
chore(deps): bump the actions group across 1 directory with 5 updates (#162577)
Bumps the actions group with 5 updates in the / directory:

| Package | From | To |
| --- | --- | --- |
| [ruby/setup-ruby](https://github.com/ruby/setup-ruby) | `1.324.0` | `1.327.0` |
| [runs-on/action](https://github.com/runs-on/action) | `2.3.1` | `2.4.0` |
| [github/codeql-action/init](https://github.com/github/codeql-action) | `4.38.1` | `4.38.2` |
| [github/codeql-action/analyze](https://github.com/github/codeql-action) | `4.38.1` | `4.38.2` |
| [github/codeql-action/upload-sarif](https://github.com/github/codeql-action) | `4.38.1` | `4.38.2` |



Updates `ruby/setup-ruby` from 1.324.0 to 1.327.0
- [Release notes](https://github.com/ruby/setup-ruby/releases)
- [Changelog](https://github.com/ruby/setup-ruby/blob/master/release.rb)
- [Commits](a0102e0972...14594264cd)

Updates `runs-on/action` from 2.3.1 to 2.4.0
- [Release notes](https://github.com/runs-on/action/releases)
- [Commits](efac073ea2...dfae4d98c5)

Updates `github/codeql-action/init` from 4.38.1 to 4.38.2
- [Release notes](https://github.com/github/codeql-action/releases)
- [Changelog](https://github.com/github/codeql-action/blob/main/CHANGELOG.md)
- [Commits](1c5b675653...2892aa5e19)

Updates `github/codeql-action/analyze` from 4.38.1 to 4.38.2
- [Release notes](https://github.com/github/codeql-action/releases)
- [Changelog](https://github.com/github/codeql-action/blob/main/CHANGELOG.md)
- [Commits](1c5b675653...2892aa5e19)

Updates `github/codeql-action/upload-sarif` from 4.38.1 to 4.38.2
- [Release notes](https://github.com/github/codeql-action/releases)
- [Changelog](https://github.com/github/codeql-action/blob/main/CHANGELOG.md)
- [Commits](1c5b675653...2892aa5e19)

---
updated-dependencies:
- dependency-name: ruby/setup-ruby
  dependency-version: 1.327.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
  dependency-group: actions
- dependency-name: runs-on/action
  dependency-version: 2.4.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
  dependency-group: actions
- dependency-name: github/codeql-action/init
  dependency-version: 4.38.2
  dependency-type: direct:production
  update-type: version-update:semver-patch
  dependency-group: actions
- dependency-name: github/codeql-action/analyze
  dependency-version: 4.38.2
  dependency-type: direct:production
  update-type: version-update:semver-patch
  dependency-group: actions
- dependency-name: github/codeql-action/upload-sarif
  dependency-version: 4.38.2
  dependency-type: direct:production
  update-type: version-update:semver-patch
  dependency-group: actions
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-10-01 18:58:47 +07:00
Peter Steinberger
0689d93557
ci: select affected extension lint for pull requests
Select changed extension packages and consumers of changed public source through the existing import graph, including type-only imports. Keep the complete typed programs, native lint chunks, resource limits, artifact preparation, and non-extension checks.

Retain full extension lint for scheduled/hourly main and release validation, shared lint/type policy, uncertain prior source types, and the OPENCLAW_CI_EXTENSION_LINT_FULL repository switch. Read the pinned diff base so removing a global augmentation cannot hide its consumers. Publish selected package reasons in the check-plan summary.

Twenty recent PR path scenarios emit 54 -> 40 hosted extension-lint rows (25.9% fewer): four 3 -> 0, one 3 -> 1, two already 0 -> 0, and thirteen unchanged. On Linux with four CPUs and a 16 GiB limit, warm maximum full-stripe compute was 64.44 s and the qa-lab-only stripe was 14.58 s. Shared artifact preparation took 114.24 s separately; these are not end-to-end Actions job timings. Full typed programs are unchanged. Shared loose extension consumers still retain full coverage.

The bounded historical search found one source-attributed extension-lint failure among 42 inspected runs; all three failing packages remain covered. Ten causal runs were unavailable, so this is limited historical evidence.

Proof on Linux Testboxes: whole tooling plus both owning fast configs; full core/extension/root test types; final check-changed, typed lint, boundary guards, architecture, source contracts and dead exports; eight same/different-SHA native preflight cells; native before/after manifests for twenty PR scenarios; final helper parity for all twenty plus two probes; and P2 review. Initial new-fixture errors were corrected and replayed. The sole inherited suppression-inventory failure reproduces on the unmodified parent and passes with already-landed d346309979. No workflow dispatches or CI reruns. New helper tests cost 5.25 s locally, with their Linux replay included in the focused proof.
2026-10-01 04:18:57 -07:00
Peter Steinberger
72a976e273
fix(windows): recognize gateways with literal native arguments (#160680)
* fix(windows): recognize gateways with literal native arguments

* test(windows): keep native fixture callback void

* fix(windows): complete native process inspection boundaries

* test(windows): keep native process proof in its owning suite

* fix(windows): preserve updater arguments in native census

* test: inventory captured keyboard shortcut intrinsics
2026-10-01 03:07:34 -07:00
Peter Steinberger
288ace5c4e
ci(ios): select PR simulator smoke groups by source owner
Keep app and XCTest compilation on every admitted iOS smoke job. Select the
voice/media/typography and Access/chat lifecycle simulator groups from their
runtime, test, fixture and build owners, and omit simulator preparation when
neither group is selected. Preserve full scheduled/manual/release coverage.

Add OPENCLAW_CI_IOS_SIMULATOR_FULL to restore both PR groups, emit selection
reasons in preflight and job summaries, and include the planner in the trusted
preflight/platform harnesses. Retain every existing test case and assertion.

Test cost: the complete iOS workflow file passed 67 cases in 70.65s locally
while native compilation overlapped; the three integrated workflow/checkout
files passed 653 cases in 182.09s on Linux. Native build-only and lifecycle-only
paths passed with an unchanged app executable hash and no simulator use for build-only.
2026-10-01 02:55:48 -07:00
Peter Steinberger
1195747068
ci: pin the OpenClaw Bun fork prerelease and admit native compiler tests (#162569)
CI still used the separately maintained Bun artifact and kept native-compiler tests on Node. Pin the OpenClaw Bun fork prerelease at 57fadf566d with release metadata verification and independent archive/executable checksums. Admit the 17 compiler files and library through their existing runtime owners, retain Node siblings and dual-mode coverage, and require native PTY success in the Bun-only smoke.

Proof: all ten Linux AWS selections passed (44,534 case executions), Node focused tests passed 243/243, Bun routing tests passed 208/208, and changed-file, workflow, and import-cycle checks passed. The old-first full smoke was fail/pass/pass/pass, attributing Chrome's fresh-host first launch to a shared startup issue. Eight forced-Bun ledger global-stub failures reproduce unchanged on main; that tooling suite stays on Node.

Supersedes only the pin portion of #159988. The batch-2 pin bump remains separate.
2026-10-01 02:20:18 -07:00
Peter Steinberger
37ffe3ce94
feat: back up to external disks and Cloudflare R2 with storage locations (#161913)
Adds named storage locations as a generic, pluggable capability, with backup as its first consumer.

- Core storage owner (src/storage): storage.locations config, a location marker that binds identity (runtime never creates it, so unplugged disks and different disks at the same path are refused), client-side streaming encryption (scrypt key from a SecretRef passphrase, per-object HKDF keys, AES-256-GCM segments), and a built-in filesystem provider for external disks and mounts.
- Plugin SDK: api.registerStorageProvider plus manifest contracts.storageProviders; providers move opaque bytes only.
- Bundled cloudflare plugin: an r2 provider over the S3 API with conditional writes and bounded multipart uploads; auto-enabled when a location uses provider "r2".
- Backups: backup create --to <location> with verified archives, UTC retention, list/verify/restore --from, Gateway-owned offsite schedules (installed Git schedules unchanged), per-installation namespace claims fenced at publication and deletion, backup record for external jobs, backup.status RPC, Doctor/status hints, and a Systems page Backups section.

No config or state migration; the storage section is new and optional. Proof: live R2 and mounted-disk round trips, namespace takeover trace, and a published 2026.9.7 upgrade cell with an existing Git backup schedule.
2026-10-01 02:05:44 -07:00
Peter Steinberger
ccba7baf5d
ci: bound protected PR tests and retain source inventories
Limit supplemental protected-test expansion to depth two instead of whole owner areas. Preserve the previous PR selection behind OPENCLAW_CI_NODE_SELECTION=full and summarize selected files and rules. Keep six non-import inventory guards on source edits, including the wrapper and swap-fixture regressions from #162200 and #162246. Scheduled and ordinary release inventories stay complete.
2026-10-01 00:44:46 -07:00
Peter Steinberger
abb97239bc
perf(doctor): read update history through one shared snapshot (#162232)
* perf(doctor): read update history through one shared snapshot

`openclaw doctor --repair` spent 2.4 s on an empty `update_runs` table
because `noteStaleUpdateRuns` issued three independent shared-state reads,
each preparing its own private snapshot (a snapshot-staging worker thread
plus a read-only child). Read the interrupted candidate, active runs, and
history through one artifact-preserving snapshot, hand the pre-read
candidate to reconciliation, and re-read fresh only after reconciliation
writes. Notes are unchanged; write-side revalidation stays in the worker.

Add env-gated Doctor phase timings (`doctor.*`) through the existing
startup-trace owner so future cost regressions are attributable, and run
the built-CLI Doctor proof before the parallel verifier wave in CI because
its fixed 30 s per-command budget is load-sensitive.

Idle 32-core Mac, fresh install: stale-update phase 2.4 s -> 0.4-0.7 s,
snapshot-staging boots 4 -> 1 per run, whole Doctor run 18.1 s -> ~16 s.
Under heavy host load the old path took 17.5 s for that phase alone.

* test(ci): count the Doctor proof barrier in workflow guards

The built-CLI Doctor proof now waits before the parallel verifier wave, which adds a wait_checks barrier; the workflow guard counts barriers and now also asserts the Doctor barrier's position.
2026-09-30 22:48:46 -07:00
Peter Steinberger
80823b2b40
ci(ui): select browser e2e tests by changed source owners
Use the existing PR-exempt inventory, policy watches and import graph to
select Control UI browser proofs for changed route/component owners, while
retaining five cross-cutting smoke files and tests without proven ownership.
Shared UI, harness and build inputs retain full coverage. Publish each
selected file and its reasons; OPENCLAW_CI_UI_E2E_FULL restores full PR runs.
Scheduled main and full release keep all 654 Control UI files and the
separately owned 36 real-Gateway files.

A representative usage diff selects 239/654 Control UI files. Committed
weights project test work from 441.171 to 190.808 seconds; with a conservative
200-second setup reserve this is 6m31s, not a measured 5.5-minute result.
Natural PR timing remains follow-up. Ten causal historical failing PR runs
across nine PRs retain their failing files (zero misses); eight use shared
fallback, so narrow-selection backtest evidence remains limited.

Proof: Linux full test-types, full tooling with final affected-file deltas,
explicit-path check-changed, boundary lint, source contracts and architecture;
seven real manifest preflight cells cover same/different workflow SHA,
PR/schedule, kill-switch true/1 and full release. Independent P2 review clean.
Whole unit-fast: 1,498 files / 16,784 tests passed. Whole unit-fast-isolated:
138 files / 1,669 tests passed. Final affected planner proof repairs all
candidate failures from the whole-tooling run. Remaining tooling failures
are unchanged-parent PR review-expiry, Windows partition, and update-backup
fixture mismatches. Production behavior, browser assertions, screenshots,
workers and deadlines are unchanged.
2026-09-30 22:45:06 -07:00
Peter Steinberger
3e0672a035
ci: block new main-thread SQLite calls (#162366)
* ci: block new main-thread SQLite calls

* fix: preserve SQLite ratchet debt across staged checks and renames

* ci: ratchet total main-thread SQLite calls
2026-09-30 22:20:47 -07:00
Peter Steinberger
0fcd3102d7
build(android): generate localization projections at build time (#162225)
* build(android): generate localization projections at build time

Generate the native Kotlin lookup and XML rows as cached Gradle source outputs. Keep tool-display translations in the native inventory so clean builds preserve the existing localized bytes, and retain manual resources and locale validation.

* style(android): format localization generation task

* fix(android): retire generated lookup from locale publisher
2026-09-30 21:54:53 -07:00
Peter Steinberger
be052821d6
chore: block new wall-clock test timeout races (#162357)
Tests must not race real timers (docs/help/testing/writing-tests.md, Cost
budget and Flake triage), yet withTestTimeout and raceWithTimeoutResult call
sites grew from 123 (2026-09-01) to 478 (2026-09-26). #161691 audited them,
fixed the worst files, and added the timer-free replacements
awaitGateBeforeSettlement and withinTest. This stops new uses.

check:test-timeout-race-ratchet keeps per-file counts in
config/test-timeout-race-baseline.txt (172 files, 403 sites) and only lets
them shrink. It parses every repository code file that mentions either helper
and counts each identifier reference except import specifiers: plain and
generic calls, namespace calls, aliases, re-exports, and local copies such as
the private raceWithTimeoutResult in fetch-guard.ssrf.test.ts. Comments and
strings do not count; test/helpers/promise.ts owns the helpers and is
excluded. Failures point authors at awaitGateBeforeSettlement, withinTest, or
fake timers through the owner's clock seam; removed sites require --prune.

The per-file count lifecycle (merge-base comparison, verified renames, base
drift allowance, prune, shrink demand) moves from the assertion-safety ratchet
into scripts/lib/shrink-ratchet.mts so both ratchets share one owner. The
assertion-safety output and exit codes are unchanged.

Wiring: scripts/check.mts preflight, check:changed routing for any code file
or the baseline, and one added line each in the existing PR baseline-ratchet
step and the main-push ratchet step (ci.yml +164 bytes, no new steps).

Proof (Linux Testbox, pre-rebase tree; baseline refreshed after rebase): new ratchet passes in under 1 s; injecting a call into
a baselined file and a new test file fails with the guidance; assertion-safety
reports its unchanged totals; tsgo:scripts, tsgo:test:root, and
check:changed --base origin/main pass; the affected test/scripts files pass.
New test file: 4 s with --maxWorkers=1.

Release note context: maintainer tooling only; no user-visible change.

CI: run 36812237216 green except checks-windows-node-test-3, which fails the same
package-update-swap.windows.test.ts assertion on main (scheduled run 36807764485);
fix owned by #162361.

Related: #161691
2026-09-30 21:39:29 -07:00
Peter Steinberger
de845cf527
ci: retain extension boundary on Blacksmith during overflow (#162289)
Keep the extension package-boundary check on Blacksmith during optional hybrid/runson hosted overflow. Hosted boundary jobs could not restore the self-hosted compiled-declaration archives (restore key includes runner.environment) and took 22-23 minutes; 2 of 25 boundary jobs in a 40-run census ran hosted. Explicit GitHub overrides, retry/manual/trust fallbacks, compiler checks, canary, deadlines and concurrency are unchanged.

Merged past one inherited red: checks-windows-node-test-3 src/infra/package-update-swap.windows.test.ts "preserves the package after persistent EPERM on linux (retries=false)" came from b5555b0bd9 (#162231), is red on main in hourly 36811800156, and is fixed on main by d7b029347f (#162352). This PR touches no Windows or updater code.
2026-09-30 21:30:50 -07:00
Dallin Romney
4f3945f904
fix(release): require signed publication tags (#162323)
* fix(release): require signed publication tags

* fix(release): accept SSH-signed tag retries

* fix(release): reject lightweight signed-commit tags

* fix(release): re-sign local publication tags

* fix(release): pin signed tags across publication

* test(release): stage signed-tag finalization helper

* test(release): model signed finalization tag

* test(release): model signed Android tag resolution

* test(release): model signed finalization refs
2026-09-30 21:27:24 -07:00
Josh Avant
ce2c70a3f0
feat(android): distribute daily builds through Firebase (#162371) 2026-09-30 23:20:47 -05:00
Peter Steinberger
55fe1b4889
ci: let canonical PR rerun matrices finish
Disable native Node matrix fail-fast for every openclaw/openclaw PR
attempt. Run 36804915849 attempt 2 cancelled 57 jobs after an inherited
main failure, preventing the remaining green proof needed by the
explicit prior-CI admin landing route.

Keep first-attempt monitoring, runner caps, routing, timeouts, and other
matrices unchanged. Qualify cancellation against each run's tested
workflow: retain historical expressions and accept the new expression
only for PRs in other workflow repositories. Align CI and landing docs.

Local proof: cancellation verifier 41 tests, workflow control 14 tests,
monitor 65 tests, hourly CI 22 tests, focused Node planning 1 test,
runner-cap and workflow-size guards 2 tests. The new regression failed
on the original workflow. Workflow sanity, formatting, and focused lint
passed; ci.yml is 404072 bytes under the 480000-byte budget. Codex P2
review found no actionable findings. No CI dispatch or rerun requested.
2026-09-30 20:41:33 -07:00
Peter Steinberger
97e5ea7be9 ci: reuse iOS smoke test build products
Build current PR smoke products once for testing, then run both focused simulator groups without rebuilding. Preserve test selectors, Debug settings, destination, separate group logs and historical/non-smoke actions.

Validation: cold/warm Xcode proof passed all 139 focused tests with unchanged app hashes; missing-product and preparation-failure controls fail explicitly. Full test-types, whole tooling config, changed checks, boundary lint, workflow validation, both manifest harness shapes and P2 review passed on the scoped candidate.
2026-09-30 16:07:45 -07:00
Josh Avant
e09dfc8897
feat(android): add daily internal testing builds (#161812) 2026-09-30 18:01:07 -05:00
Peter Steinberger
eb65a6c5f1
ci: overlap iOS simulator preparation with builds
Select the simulator before compilation, then boot and slim that exact
simulator alongside the app build. Join preparation before XCTest, retain
its failure log, and bound a hung join without changing test selection or
build settings.

Validate both focused groups on cold and warm native runs (139 tests each),
plus preparation failure and timeout controls. Linux proof includes full
test types, the whole tooling config, changed checks, source contracts,
and both preflight harness shapes for PR and scheduled-main inputs.
2026-09-30 14:48:59 -07:00
Peter Steinberger
7045908edb
refactor(config): deslop config eighth pass (#162057)
## What Problem This Solves

The maintainer-requested eighth config cleanup removes redundant schema construction and helper plumbing left after the earlier passes.

## User Impact

No user-visible change. Config fields, generated schemas, defaults, environment precedence and preservation, Doctor normalization, redaction, and persisted state retain their existing contracts. The protected SQLite accessor files and the 698-line config environment owner are untouched.

## Why This Change Was Made

- Construct 202 strict Zod objects directly instead of constructing and then cloning each object to make it strict. Permissive objects, catchalls, refinements, field metadata, and non-immediate strict chains stay intact.
- Use the shared promise-cache owner for config observation roots, retaining successful identities and evicting only the rejected promise.
- Remove a plugin normalization cache used once, duplicate schema types, and copied suppression records; keep Doctor migration cloning explicit.
- Share history option capture and inline single-use transcript helpers without changing storage, authority, or projection behavior.
- Remove a redundant override cast and guard, inline the Nix error's single-use formatter, and shrink the assertion allowance for the removed cast. The core schema also fits its actual line limit now, so its grandfathered suppression and baseline entry are removed.

Production diff: 1,018 added / 1,356 removed, **338 net lines removed**. No tests were added or weakened. Coverage records 152 production files over 200 lines read, with 83 protected SQLite accessor files explicitly excluded.

## Evidence

- Independent Codex review completed with no actionable P0-P2 findings.
- Blacksmith Testbox `tbx_01m3skp23vxamk2c0ggwd6mx4r`, [run 36747330340](https://github.com/openclaw/openclaw/actions/runs/36747330340): all changed-source hashes matched the candidate based on `870c5b6b7f`.
- Full config suite: 471 files passed, 3 skipped; 5,661 tests passed, 10 skipped. Focused config, history, and shared-helper tests also passed.
- `pnpm build` passed. Both import-cycle checks reported zero cycles. Filesystem import-boundary suite: 23 passed.
- Schema generation, generated channel metadata, and config docs baseline checks passed with no generated changes.
- `check-changed` passed core and all test typechecking, dead-export scans, and its other guards. Its final lint failure identified an obsolete max-lines suppression after the file shrank; that suppression and baseline entry were deleted, and the affected lint and max-lines checks passed separately. The removed cast's assertion allowance was also reduced.
- Final base/candidate schema comparison: all **835,669 bytes identical**, SHA-256 `f222f2a568c39b9d3ef4ef09e595701161562faf07a08190d304f608d92f4b0a`. Both refreshed cycle checks again reported zero. Final lint and suppression checks ran locally at reduced priority after remote transport failures; the full config suite, build, and typechecks ran on Testbox.

### Fixes found along the way

The first hosted run exposed a formatting assumption in performance target metadata detection: it required exactly four spaces before the canonical `mediaModels` field. Direct strict-object construction changed that indentation while leaving the schema identical. Both canonical marker checks now ignore leading whitespace; metadata remains unevaluated and all trust, legacy-layout, and pinned-ref rules remain unchanged.

The existing sparse-checkout regression failed on the first PR head and passed after the repair. Fresh Blacksmith Testbox `tbx_01m3svxdk2ajybnbxw6s7eng2n`, [run 36764456521](https://github.com/openclaw/openclaw/actions/runs/36764456521): all 56 performance workflow tests passed (12.52s wall), both cycle checks reported zero, and workflow sanity passed. Independent review of the repair found no actionable P0-P2 issues. No tests, assertions, retries, timeouts, or snapshots were changed.

### Inherited main failure

The SDK surface-budget failures in `test/scripts/plugin-sdk-surface-report.test.ts` also occur on current hourly main [run 36760032986, job 110040129161](https://github.com/openclaw/openclaw/actions/runs/36760032986/job/110040129161), main SHA `112df95f5e`. Both main and the first PR run report 4,595 exports / 2,699 callables against budgets of 4,594 / 2,698, failing the same three assertions. This PR changes no SDK entrypoint inventory, budget, or exported name. The reporter traverses config types, but the counted export invariant already fails with the same totals on main; its budgets remain with the main-CI coordinator.


### Final head and compatibility assessment

At `5b39f39a50e31942efb4edb9b9636a8e7a1b9df6`, [CI run 36766041963](https://github.com/openclaw/openclaw/actions/runs/36766041963) completed. The repaired performance-workflow shard passed. The only failed test job is `checks-node-compact-small-15`, containing the same three SDK surface-budget failures documented on main above; the CI-gate failures are downstream of it. Security-fast, dependency review, and security-sensitive review passed.

The review's data-model flag names `config-write-guard.ts`, where this diff only inlines the existing Nix error formatter. No storage operation, migration, serialization format, updater marker, or lifecycle step changed. Streaming normalization retains the same clones and precedence; plugin normalization retains the same input and result after removing a cache used once. The complete generated schema is byte-identical, and config/Doctor-facing suites passed. This introduces no new migration or upgrade contract. A published-updater integration run is not claimed; the path-based data-model classification does not describe the actual change.
2026-09-30 13:01:28 -07:00
Peter Steinberger
cfd219c519
fix(release): skip pending ClawHub publications and surface recovery for failed ones (#161985)
The ClawHub release planner and prepared-artifact resolver read the new public publication-state endpoint (/api/v1/packages/{name}/versions/{version}/publication). Only absent versions are republished; pending ones are skipped, and failed ones are excluded, with the recover command printed to the step summary. A 404, or a 200 without a state field, falls back to the legacy version probe.
2026-09-30 09:35:54 -07:00
Peter Steinberger
ebe57ef28a
ci(android): restore headroom for phone test rows
Move third-party app lint back to the existing Wear row and budget
Blacksmith phone tests across up to four isolated JVMs. Each JVM receives
its share of the available CPUs; the existing 1 GiB heaps remain unchanged.

On the same eight-CPU Linux Testbox, clean project outputs and warm
dependency caches reduced the affected phone row from 446.13 to 233.34
seconds of Gradle wall time. Wear moved from 20.49 to 174.36 seconds;
the two rows together fell from 466.62 to 407.70 runner-seconds.
All 3,665 third-party and 3,533 Play tests passed, along with Wear/shared
tests and the selected Android lint tasks.

Keep all four normal rows, all six full-validation rows, test assertions,
Gradle tasks, hosted settings and timeout budgets. Actions job-wall
measurement remains a follow-up for the next natural hourly.

Validation: complete 905-file tooling inventory across the interrupted and
resumed runs; 897 files passed, seven retained existing skips, and three
assertions in pr-merge-prior-ci-timeout.test.ts failed identically on the
unmodified parent (25 passed, three failed). Full test-types, main/PR/full
preflight smoke, workflow validation, boundary lint, six-file check-changed
and direct-API P2 review passed. No test or assertion was weakened.
2026-09-30 07:51:27 -07:00
Peter Steinberger
94c99f73e0
ci: defer published upgrade survivor to hourly verification
The intact published-upgrade-survivor runs in every admitted hourly main CI run
and Full Release Verification through its normal_ci child. Preserve the exact
legacy-operator-state scenario with auto-auth when the target declares it.

No PR owner changes select this survivor, including updater, Doctor,
state-migration, or direct survivor changes. The other five Docker seed lanes
keep their existing PR owner selection.

Frozen historical targets retain their supported base fallback; invalid
catalogs fail. No survivor scenario case or assertion is removed.
2026-09-30 02:57:50 -07:00
Peter Steinberger
3487cc00a7
ci: use smoke packages for Docker seed PRs
Reuse the existing main smoke package profile for owner-selected PR Docker seed checks. Keep runtime builds, SDK declarations, tarball validation, and upgrade survivor assertions; full declarations remain in hourly cache warming and Full Release Validation.
2026-09-30 02:57:49 -07:00