Commit graph

785 commits

Author SHA1 Message Date
Ayaan Zaidi
824b12c7de
fix(qa): make Telegram delivery-failure proof deterministic (#155899)
## What Problem This Solves

Telegram Test Server proof could not deterministically reject one final Bot API request, and starting a prebuilt QA provider through the source launcher could rebuild during the credential lease and miss readiness.

## User Impact

No production behavior or configuration changes. Maintainers can exercise final-send and deletion failures against a real Telegram Test Server user without changing the application under test or introducing a second lease.

## Why This Change Was Made

The existing local proxy gains one method/occurrence/body-filtered, pre-upstream Bot API rejection. The QA runner keeps the development launcher for `--source-gateway` and uses the exact built entry for its prebuilt provider, explicitly enabling the private source-only QA command. The same runner still owns the credential, Test Server proxy, Gateway, user recorder, and cleanup.

## Evidence

- 36 focused Node tests pass. The exact-head ClawSweeper review identified a file-download failure when no rejection was armed; the proxy now requires an armed fault before comparing methods. The HTTP-boundary regression proves `/file/bot…` downloads are forwarded before, during, and after a one-shot rejection, with the rejected request never forwarded. Focused command: `node --test .agents/skills/telegram-e2e-userbot/scripts/telegram-test-api-proxy.test.mjs .agents/skills/telegram-e2e-userbot/scripts/scenario.test.mjs .agents/skills/telegram-e2e-userbot/scripts/run-mock-sut-user-e2e.test.mjs`; 36 passed, 0 failed, runner duration 2.66 seconds, wrapper wall 3.36 seconds.
- A loopback private QA provider started from each independently built baseline/candidate entry in under four seconds and served `/debug/requests`.
- Telegram Test Server: an independently recorded 20-second tool turn and one injected final-send rejection produced a visible baseline failure and a candidate terminal status without rerunning the tool. The comparison uses one frozen harness revision for both builds.
- One timed-out username discovery and a separate failed provider setup are retained as setup failures, not product verdicts; the bounded native chat preparation for the final runs remained outside this PR.
- After integrating `main`, the 17 focused wrapper-import/source-closure regressions passed, including the SQLite reader modules required by current main. Exact-head hosted CI's `openclaw/ci-gate` job passed; the security review status is awaiting its automatic reevaluation.

The production repair and its real-flow evidence are in #155906.

Merged current `main` (`86d353b748`) to align the older branch with its SQLite reader and PR wrapper inventory. The two earlier cherry-picked upstream CI fixes retain Peter Steinberger's commit authorship; they are already on `main` and disappear from the PR diff. The QA change remains six files, with no production behavior added.
2026-09-23 09:12:32 +05:30
PollyBot13
8a53ee3e16
fix: handle endpoint-protected plists during Gateway activation (#149307)
* fix: handle endpoint-protected plists during Gateway activation

* fix: refuse interrupted launchd label extraction

* test: drain direct chat projections before fixture teardown

* test: extract oversized gateway fixture helpers

* test: await owned catalog queue mutations

---------

Co-authored-by: Patrick Erichsen <patrick.a.erichsen@gmail.com>
2026-09-22 20:37:25 -07:00
Peter Steinberger
f4a5f85241
fix(ci): warn on locale drift before release validation dispatch
The release CI skill refused dispatch when generated locales lagged source
PRs, although source/generated isolation deliberately delegates translation
to serialized post-merge refreshes. Report that lag as advisory and allow
validation to start.

Keep strict locale jobs and PR policy unchanged. The FRV diagnostic summary
reads existing child evidence to show both locale results and warn on
failures without changing the failed validation state or adding checks.

Regression fixtures fail before the summary exists and preserve failed,
successful, and missing evidence afterward. Testbox validates strict locale
parity plus 20 standalone runs and three tooling-config runs with workflow
siblings. Independent review found no actionable P2-or-higher findings.

Measured changed test file cost: 6.35 seconds wall with maxWorkers=1.
2026-09-22 19:48:31 -07:00
Vincent Koc
72094fee91
chore(autoreview): sync canonical review policy and runtime fixes (#156039) 2026-09-23 09:23:37 +08:00
Josh Avant
fdae00fe82
chore(qa): qualify execution identity across live boundaries (#156033)
* test(qa): repair forced restart and message inspection harnesses

* test(e2e): qualify installed execution identity persistence

Extend the packed npm onboarding owner with audit opt-in, one deterministic local turn, installed CLI inspection before and after Gateway restart, and isolated-state/privacy assertions.

Qualification from 267bb9d2: new shell-flow proof fails before the harness change at the missing opt-in. All 53 focused support tests pass; final changed shell/SQLite selection took 26.94s with one worker. Changed checks and P0-P2 autoreview pass; the full export scan was reused after brace-only lint fixes. Packed Docker proof remains a remote handoff.

* test(qa): qualify live execution and channel identity

* test(audit): qualify execution identity lifecycle gaps

* test(qa): align suppression audit with required replies

* test(qa): qualify Telegram participant identity

* test(qa): provision Telegram identity fixtures

* test(qa): qualify private-production Telegram identity

* test(e2e): preserve installed execution identity proof

* test(qa): await post-delivery memory maintenance

* fix(qa): repair qualification CI boundaries
2026-09-22 20:15:02 -05:00
Peter Steinberger
5ac833c172
ci: pack test jobs and add opt-in Spot cron routing (#155403)
* ci: reduce extension fanout and preserve fast artifact builds

* ci: add opt-in RunsOn spot routing

* test: repair CI qualification fixtures and record Spot measurements

* ci: pack measured Node jobs and qualify main routing

* ci: price compact packing from complete tooling measurements

* docs: refresh CI counts after runtime tier integration

* fix(ci): keep the type-shard size default internal
2026-09-22 13:08:30 -07:00
Dallin Romney
ab4c2cd058
fix(qa): reject stale Telegram credential archives (#155601)
* fix(qa): reject stale Telegram credential archives

* test(qa): align Telegram readiness fixtures

* fix(qa): preserve Telegram readiness boundaries

* fix(qa): validate restored Telegram chat lists
2026-09-22 11:44:21 -07:00
Peter Steinberger
dd0003ac5c
fix(release): forward the stable soak waiver through the candidate helper (#155798)
Three release-tooling fixes surfaced by the 2026.9.6 stable candidate (Full Release Validation run 35742760135). Each is a separate commit.

## 1. Forward the stable soak waiver through the candidate helper

**Problem.** `pnpm release:candidate` runs the publish preflight in-process with a hardcoded `stableSoakWaiver: ""`. For a stable tag (for example `v2026.9.6`) whose Full Release Validation ran with `release_profile=beta` and `run_release_soak=false`, the `<consumer>.soak` and `core-npm.performance` gates in `scripts/lib/release-publish-gates.mts` FAIL instead of WARN, so the helper throws `Publish preflight failed` even though `docs/reference/RELEASING.md` documents the operator `stable_soak_waiver` for exactly this case and `pnpm release:publish-preflight` already accepts `--stable-soak-waiver`.

**Solution.**
- Add `--stable-soak-waiver <reason>` to the candidate helper's option table and usage, mirroring the standalone preflight option name.
- Forward `options.stableSoakWaiver` into the embedded `runReleasePublishPreflight` call. The normal-route command is rendered by the existing `buildReleasePublishDispatchCommand`, which already emits `-f stable_soak_waiver=<reason>` when set; no duplication.
- Render the same input in the prepared-route `publish_inputs` payload (`buildPublishCommand`), which `scripts/openclaw-release-ready.mjs` already accepts.
- Document the option in `docs/reference/RELEASING.md` and the maintainer skill reference.

Without the option the behaviour is unchanged (empty waiver, soak required).

## 2. Pack bundled dependencies in the prepared npm bundle (unblocks 2026.9.6)

**Problem.** The "Prepare publishable npm package" job failed with `ERR_PNPM_BUNDLED_DEPENDENCIES_WITHOUT_HOISTED` from `pnpm --dir <root> pack`. Root `package.json` gained `bundleDependencies: ["chrome-devtools-mcp"]` in #154215, and pnpm refuses to pack bundled deps under the isolated linker. `scripts/package-openclaw-for-docker.mts` already passes `--config.node-linker=hoisted`; the default `runPack` in `scripts/npm-prepared-bundle.mjs` did not.

**Solution.** Add `--config.node-linker=hoisted` to the pnpm pack args. Scripts stay enabled (`prepack` must run; no `ignore-scripts`). Verified locally on the release candidate: `OPENCLAW_PREPACK_PREPARED=1 pnpm --dir . pack --config.node-linker=hoisted --pack-destination /tmp/x` succeeds and the tarball contains 346 `node_modules/chrome-devtools-mcp/` entries.

## 3. Raise the npm pack unpacked-size budget to 320 MiB

**Problem.** The release-checks job `install_smoke_release_checks / installer_smoke_update` failed with `candidate.tgz unpackedSize 312428525 bytes exceeds budget 246415360 bytes`. The 2026.9.6 package is 296.6 MiB unpacked vs 214.6 MiB for 2026.9.5.

**Why this is a budget bump, not a bloat fix.** The growth is intentional product work, not accidental duplication: +46.2 MiB `dist/worker/sqlite-store.worker.mjs` (portable cloud SQLite worker bundle, #154711) and +12.6 MiB bundled `node_modules/chrome-devtools-mcp` (#154215), plus ~4.6 MiB `worker.mjs` growth and 3.6 MiB `dist/state`. Operator-approved release-prep exception for 2026.9.6; a follow-up should shrink the sqlite-store worker bundle.

**Solution.** Raise both defaults from 235 MiB to 320 MiB, kept equal: `scripts/lib/npm-pack-budget.mts` and `scripts/test-install-sh-docker.sh` (the `OPENCLAW_INSTALL_SMOKE_PACK_UNPACKED_BUDGET_BYTES` override is unchanged).

## Impact

Stable candidates validated on the beta profile without soak can be completed by the candidate helper with the operator's approved waiver; the prepared npm bundle packs again with bundled dependencies; and the install smoke budget admits the 2026.9.6 package while still bounding accidental growth.

## Evidence

- `test/scripts/release-candidate-checklist.test.ts`: the coordinator matrix now passes `--stable-soak-waiver` for stable `latest` cases and asserts the waiver reaches the preflight input and the printed normal/prepared command; negative-checked (reverting the forwarding line fails 5 cases).
- `test/scripts/npm-prepared-bundle.test.ts`: new test intercepts the default `runPack`'s pnpm invocation and asserts the exact args include `--config.node-linker=hoisted` with `OPENCLAW_PREPACK_PREPARED=1` and no `ignore-scripts`.
- `test/release-check.test.ts` and `test/scripts/test-install-sh-docker.test.ts`: budget boundary cases updated to 320 MiB (exact budget passes, one byte over fails).
- `node scripts/run-vitest.mjs run` wall times (local, macOS): release-candidate-checklist + release-publish-gates + release-publish-preflight-interface + release-publish-preflight-evidence: 4 files, 189 passed in 16.71s; npm-prepared-bundle: 40 passed in 3.62s; release-check + test-install-sh-docker: 176 passed in 63.92s (57% transform). The new hoisted-linker test adds one in-process fixture pack (no real pnpm), well under a second. CI timings: the same files run inside the sharded `checks-node-compact-*` lanes on the PR head; no lane exceeded its budget (see PR checks).
- `OPENCLAW_TESTBOX=1 pnpm check:changed` (hosted Testbox) on head 7dfaf713a80: passed, all lanes including the test-root lint lane.
2026-09-22 16:10:03 +00:00
Vincent Koc
b6bd69d6e8
docs(agents): clean up finalized task artifacts (#155716) 2026-09-22 12:39:30 +00:00
Ayaan Zaidi
d5fc539e74
feat(discord-e2e): check bot readiness without QA Lab (#155598)
Related: #153456

## What Problem This Solves

Discord E2E agents need a quick way to verify the shared bot credentials and channel access without starting QA Lab or confusing API access with a full Gateway test.

## User Impact

Maintainers with the existing Convex login and provisioned Discord pool can check both bot identities, the pinned SUT application, the leased guild text channel, and observed history access using one read-only command. Empty history is reported as inconclusive (exit 2), because Discord returns an empty list both for an authorized empty channel and when Read Message History is missing. No Discord messages or other objects are created. Convex lease acquisition, heartbeat, and release still occur.

Only the Discord skill changes. QA Lab, Gateway, release workflows, dependencies, app permissions, and application storage schemas are unchanged.

## Why This Change Was Made

The command reuses the existing credential-lease and bounded-response helpers. It checks lease health before and after reads, stops on cancellation or lease loss, hides private response/error content, and fails when lease release is not confirmed.

Mutation testing remains with the existing `channelE2e` lifecycle rather than adding another send/edit/delete runner or receipt ledger. The skill now clearly separates read-only readiness, owned bot mutations, Gateway/model behavior, and actual human interactions. It also distinguishes stored public replies from client rendering and explains why the bot pair cannot prove human-to-bot DMs or user interactions.

The private `result.json` is a named test result containing identities and checks, not a new application database or migration contract. Tokens and channel history are not saved or printed.

## Evidence

- Final live Discord readiness returned `status: ready`: all eight checks passed through a real Convex lease, both bots returned valid history message metadata from the leased channel, and the broker confirmed release. The command uses GET requests only and starts no Gateway, recorder, or model.
- Twenty-three isolated CLI controls passed, including both empty-history cases, malformed message metadata, content-redacted messages, and failure/release-error precedence over an inconclusive result, alongside identity/access, cancellation, lease-loss, privacy, response-bound, and read-only controls.
- The empty-history defect was reproduced through the original CLI: HTTP 200 with an empty list incorrectly exited 0 and claimed history access. The corrected command exits 2 with `status: inconclusive` for both a denied-history bot and an authorized empty channel, without claiming either is a proven permission failure. This follows Discord's [documented response semantics](https://docs.discord.com/developers/resources/message#get-channel-messages); live app permissions were not changed to create the controls.
- Script syntax, focused Oxlint, Oxfmt, scoped Markdown lint, local links, and whitespace checks passed.
- The existing mutation and Gateway lifecycle was source-audited, not rerun or changed. No new message, thread, file, reaction, Gateway, model-tool, or client-rendering coverage is claimed by the readiness check.

Full hosted CI remains required before landing; local proof does not replace it.

Co-authored-by: Ayaan Zaidi <hi@obviy.us>
2026-09-22 16:28:50 +05:30
Ayaan Zaidi
48769b4583
feat(slack-e2e): check shared user OAuth without QA Lab (#155547)
Related: #153456

## What Problem This Solves

Slack E2E agents need a simple way to check the shared user OAuth account and exercise user-authored text without running or changing QA Lab.

## User Impact

Maintainers with the existing Convex login and provisioned Slack pool can run a read-only Slack readiness check, or explicitly create, read, edit, and delete one owned message. No Gateway, model, QA application, release workflow, dependency, or shared runner changes are included. Existing bot-based Gateway recipes remain unchanged.

## Why This Change Was Made

The Slack skill now separates user API proof, bot/Gateway proof, and actual Slack client interaction. Its small opt-in command reuses the existing lease owner and retains private receipts. Cancellation stops new test actions; cleanup needs the live lease and the exact message receipt. Uncertain writes are not replayed, and failed cleanup is not a pass.

A provisioned official user OAuth grant is required; the command does not create tokens or change scopes. Slack may attach app attribution to a user-authored API message, so a passing smoke does not prove human-client ingress. Reconstructed approval images are also identified as API-data renderings, not Slack client screenshots.

## Evidence

- Live leased Slack smoke passed: pinned human/workspace and driver app checks, channel history, stored user/app authorship, stored edit, and stored deletion. Slack reported app attribution; cleanup completed and the broker confirmed lease release.
- Twelve isolated CLI controls passed: read-only mode, successful lifecycle, missing user, wrong workspace, ambiguous and rejected creates, foreign app attribution, cancellation after create, lease loss after create, stale edit readback, denied deletion, and unconfirmed deletion. No duplicate create or unauthorized cleanup request was sent; synthetic secret errors were not printed or retained.
- Script syntax, focused Oxlint, Oxfmt, scoped Markdown lint, and whitespace checks passed.
- Full local lint preparation could not run in the nested checkout because of the existing ancestor dependency boundary. The broad Markdown command also found an unchanged extra blank line in `docs/internal/conversation-history-redesign.md`; the two changed Markdown files pass the same rules. Hosted CI remains required before merge.

This is API-user text proof only. No new Gateway, Slack client UI, button, slash-command, or model-tool behavior is claimed.

Co-authored-by: Ayaan Zaidi <hi@obviy.us>
2026-09-22 13:50:27 +05:30
Ayaan Zaidi
43ed8cee3f
feat(qa): add Convex-backed Discord and Slack E2E skills (#153471)
Related: #153456

## What Problem This Solves

Agents need reusable Discord and Slack E2E workflows with the same Convex-login setup as the Telegram userbot skill, rather than separate secret setup and ad hoc channel probes.

## User Impact

Developer-only: repository skills provide same-lease readiness, reusable YAML flows, native message/file/thread actions, correlated Gateway replies, and explicit evidence/cleanup instructions. Telegram reuses the extracted broker-discovery owner. Curated QA defaults and production channel behavior stay unchanged.

The dedicated Slack QA Driver now has the operator-approved reaction/file read/write scopes and was reinstalled. Its existing pooled token remains valid; no workspace-admin privileges, user-token scopes, or production-app changes were needed. The setup manifest includes those scopes.

## Why This Change Was Made

The skills extend the existing QA Lab transport, Gateway, scenario, and credential-lease owners instead of adding another runner. Opt-in `--doctor` and repeatable `--scenario-file` use Convex CI leases and deterministic `mock-openai` by default, while preserving explicit overrides.

Native-write flows do not retry ambiguous writes. Cleanup retains lease authority through Gateway shutdown, captures final native receipts before temporary state removal, and removes only owned fixtures. Native and cleanup Slack requests have bounded settlement with no write replay. Unanswered Gateway mutations remain explicit uncertainty and preserve runtime evidence instead of disappearing. Discord recording remains active through shutdown; threads are archived rather than deleted when the lease lacks thread-management permission.

Native API receipts are not model-tool or visual proof. Human slash commands, component clicks, modals, ephemeral replies, and rendering retain explicit client-testing requirements. Broader testing capabilities are deferred to follow-up PRs.

## Evidence

- Real Discord full native lifecycle passed on `ea1b053620`: messages/replies/pagination/edit, reaction add/remove, public thread, attachment upload/delete, a fresh correlated SUT reply, continuous recording, and successful owned cleanup.
- Real Slack full native lifecycle passed on `db88f0ac1c`: all five steps passed, including quiet ingress, threaded Gateway reply, stored edit/delete and reaction add/remove verification, and exact file upload/readback. Cleanup reported zero remaining owned messages/files/reactions and zero pending, uncertain, or failed operations.
- Both native runs used real transports, a temporary Gateway, `mock-openai`, and Convex CLI discovery with both broker environment variables unset. Telegram/shared discovery regressions: 30 passed.
- Earlier focused QA set: 301 passed; subsequent affected-consumer set: 298 passed (overlapping, not additive). Extension typecheck, changed-TypeScript type-aware lint, source ratchets, docs checks, and whitespace checks passed during preparation.
- Landing reconciled one pinned main snapshot, `eac221ae78`, preserving both lifecycle guards. After restoring its frozen dependency versions, the merged Gateway lifecycle, Slack ownership/adapter, and scenario deadline suites passed: 4 files, 52 tests. Earlier native artifacts are retained as pre-reconciliation proof, not relabeled as executions of the merge head.
- Repaired-head Slack lifecycle passed on `62ed1326be41`: all five steps, reaction/file readback, final receipt capture, and clean owned teardown. This rerun covers the bounded-request and uncertain-capture repairs, on the reconciled base.
- Both teardown regression controls failed on the pre-repair implementations (missing HTTP deadline; unanswered mutations dropped). The repaired owner/consumer set passed all 73 tests across six files. Targeted type-aware lint, extension typecheck, docs MDX and whitespace checks also passed.
- Hosted CI will not be awaited, per the operator's explicit landing instruction; no full CI success is claimed.

Co-authored-by: Ayaan Zaidi <hi@obviy.us>
2026-09-22 09:03:22 +05:30
Peter Steinberger
be0ee80928
test(ci): split workflow coverage by responsibility (#155001)
Co-authored-by: Peter Steinberger <steipete@gmail.com>
2026-09-21 17:50:29 -07:00
Peter Steinberger
836aeaa718
ci: defer maintainer tooling on product-only PRs (#153962)
* ci: defer unrelated report composition to release validation

* ci: keep unrelated tests out of release owner watches

* ci: preserve exact report tier owner routing

* ci: defer full tooling family on product-only PRs

* test: align CI guards with tooling owner routing

* ci: retain canonical UI consumers with the tooling tier
2026-09-21 22:31:16 +00:00
Peter Steinberger
6869ace3d3
chore: move Apple CI to Xcode 27 (#154332)
* ci: validate Apple builds and tests with Xcode 27

* fix(ci): include Swift selection in trusted platform checkouts

* fix(ci): retain Periphery index layout and report Swift crashes

* fix(test): adapt native validation to Xcode 27 runtime and indexing

* test: run Quick Chat presentation flows with Swift Testing

* docs: clarify Xcode analyzer compatibility requirements

* test: diagnose early AppKit test process exit

* test: give rendered Mac tests an AppKit event loop

* test: await AppKit event processing before rendered tests

* test: wait for the rendered Quick Chat model picker

* fix(ci): isolate Apple test logs and await fixture readiness

* test: retain all default-profile capture evidence

* fix(ci): repair Apple qualification diagnostics and fixtures

* test: surface Quick Chat sends rejected before transport entry

* test: capture suspended Quick Chat tasks during CI stalls

* test: read attributed accessibility titles without abandoning Swift tasks

* test: normalize accessibility titles across native fixtures

* test: find named-profile model controls by accessible name

* fix: retain lost Windows PID authority through shutdown retries

Co-authored-by: Peter Steinberger <steipete@gmail.com>
2026-09-21 13:57:07 -07:00
Peter Steinberger
8f5be799d9
improve: reduce isolated Gateway CI test time (#154366)
* ci: use measured Gateway test worker limits

Admit up to eight workers for the Gateway isolated/database-worker cohort only on roomy serial self-hosted jobs, with a 28 GiB memory floor and the existing two-worker fallback. Preserve sibling caps, timing history, coverage and matrix budgets.

* test(ci): preserve measured Gateway placement expectations

* test: identify compiler lifetimes across PID reuse

* test(ci): respect proof tiers in Gateway policy coverage

* test(ci): retain synthetic Gateway placement ownership

Use the upstream synthetic Gateway fixture for recipient and worker-cap assertions after rebase. Keep the fixture isolated from the changing project inventory and preserve every placement and execution-policy assertion.

* test(ci): preserve Gateway caps through dense packing

Keep the upstream settled serial row, prediction, inventory, and neighboring packing assertions while checking the measured Gateway and unmeasured sibling worker limits at their group owner.

* ci: keep session history benchmark on Node

Retain the observed-failing Node process and SQLite lifecycle benchmark in the existing Node-required runtime partition. Preserve the complete test, profiles, assertions, and deadline, with routing coverage for bun-compatible and dual policies.

Co-authored-by: Peter Steinberger <steipete@gmail.com>
2026-09-21 11:11:06 -07:00
Peter Steinberger
9eb594a3de
fix(ci): use tooling worker capacity (#153820)
* test: verify owned listener cleanup

Observe native listener and child close events at operation settlement instead of rebinding released ephemeral ports. Preserve held-port, refusal, HTTP, and abort assertions. Linux tooling endurance exposed the UI preflight race; repair the three matching cleanup oracles as well.

* fix(tooling): tolerate process exit during lifecycle sampling

* test: observe owned resources during tooling cleanup

* ci: run tooling files in parallel

Use admitted file workers when balancing and pricing tooling shards, retaining the longest-file floor and canonical artifact owners. Defer unrelated source inventory discovery for precise tooling plans.

* test: preserve feasible tooling admission fixtures

* test: balance saturated tooling capacity fixture
2026-09-21 14:24:33 +00:00
RoboClaw
32fbb3c6f3
fix(release): keep two extended-stable months eligible (#154777)
Co-authored-by: roboclaw-bot <309084314+roboclaw-bot@users.noreply.github.com>
2026-09-21 04:41:53 -07:00
Peter Steinberger
c79a03cd8a
fix: PR landing fails on exhausted GitHub quota (#154558)
* refactor: simplify PR landing and test fixtures

* fix: recognize GitHub quota exhaustion with CORS headers
2026-09-21 08:19:59 +00:00
Peter Steinberger
736219b5cd fix: recover conflicted PRs after accepted auto-merge
Add explicit retained auto-merge cancellation before head repair, preserving request history and captures. Reconcile uncertain responses and concurrent merges, then admit reviewed replacement heads through completed ordinary gates. Validate with 40 focused native recovery and CLI tests.
2026-09-21 00:54:57 -07:00
Peter Steinberger
f202df3c89 docs: follow pending PR merges through completion 2026-09-21 00:37:49 -07:00
RoboClaw
ca6a7d9818
feat(release): add non-Latest extended-stable releases (#154515)
* feat(release): add non-Latest extended-stable releases

Reconcile #120522 with current main while preserving qualified artifacts, publication approvals, active-line checks, and supported recovery routes.

* feat(release): add non-Latest extended-stable releases

OpenClaw-Publication: c2238e25-fc84-40bf-ba4d-a6d9d11f4a5b

---------

Co-authored-by: roboclaw-bot <309084314+roboclaw-bot@users.noreply.github.com>
2026-09-21 00:29:24 -07:00
RoboClaw
89cc3c7cec
fix(testbox): keep login-shell validation in the selected checkout (#153280)
* fix(testbox): keep login-shell validation in the selected checkout

Restrict the Blacksmith SSH convenience cd to interactive shells and verify exact physical cwd before Testbox ready registration. Cover source identity, synchronized patches, subdirectories, interactive startup, unknown hooks, and lease freshness.

Co-authored-by: steipete <58493+steipete@users.noreply.github.com>

* fix(testbox): keep login-shell validation in the selected checkout

Worked on by:
- @steipete

Co-authored-by: steipete <58493+steipete@users.noreply.github.com>
OpenClaw-Publication: aebc5627-0b5b-4a46-9212-cccb5ac2ee5d

---------

Co-authored-by: steipete <58493+steipete@users.noreply.github.com>
2026-09-21 07:01:16 +00:00
Peter Steinberger
0af08749bb
perf(test): tier process proofs and balance Windows CI (#153950)
* perf(test): tier process proofs and balance Windows CI

* fix(ci): import Windows planner directly and align tier guards

* perf(test): share Windows preparation and pack measured projects

* fix(ci): return explicit Windows shard fields

* docs(ci): clarify Docker proof coverage after tiering

Document PR boundary coverage, main owner selection, release survivor coverage, and the proof gap for coalesced main pushes. Confirm the existing guards consume the selected Docker inventory and dispatch exact-target release CI.

* style(test): format rebased Windows workflow assertion

* test(ci): satisfy planner inventory lint contracts
2026-09-21 06:05:45 +00:00
RoboClaw
1395926530
docs: reserve Team live updates for Night Watch (#154425)
Keep investigation and reviewed repairs delegable while reserving live execution and exact-transaction handback for the designated owner.

Co-authored-by: steipete <58493+steipete@users.noreply.github.com>
Co-authored-by: vincentkoc <25068+vincentkoc@users.noreply.github.com>
2026-09-20 21:49:29 -07:00
Peter Steinberger
62f5950ba4
improve: publish CI dependencies promptly and warm hosted JavaScript caches (#154156)
* ci: publish dependencies independently and warm hosted caches

* ci: warm hosted transforms through their consumer entrypoints

* test(ci): guard compiled worker ownership after warmer split

* docs(ci): record rejected Windows store cache experiment

* docs(ci): describe Vitest-owned bytecode coverage safeguards

* ci: expose cache warmer entrypoints to dependency checks
2026-09-20 21:36:44 -07:00
RoboClaw
0fbded75d2
docs: keep update requests owned through acceptance (#154111)
* docs: keep update requests owned through acceptance

Co-authored-by: steipete <58493+steipete@users.noreply.github.com>
Co-authored-by: vincentkoc <25068+vincentkoc@users.noreply.github.com>

* docs: notify Team only around actual update downtime

Keep non-action outcomes private and summarize the exact accepted main range after verified recovery. Preserve update ownership and distinguish rollback from successful deployment.

Co-authored-by: steipete <58493+steipete@users.noreply.github.com>
Co-authored-by: vincentkoc <25068+vincentkoc@users.noreply.github.com>

* docs: own confirmed updater repairs through acceptance

Qualify defects at the failing path, use isolated worktrees and subagents where useful, and land tested reviewed repairs without bypassing protection.

Co-authored-by: steipete <58493+steipete@users.noreply.github.com>
Co-authored-by: vincentkoc <25068+vincentkoc@users.noreply.github.com>

---------

Co-authored-by: steipete <58493+steipete@users.noreply.github.com>
Co-authored-by: vincentkoc <25068+vincentkoc@users.noreply.github.com>
2026-09-20 21:18:58 -07:00
Peter Steinberger
cdd01a1944
improve: hand pending PR checks to GitHub auto-merge (#154307) 2026-09-21 03:33:59 +00:00
Peter Steinberger
5941d999dc
ci: reserve plugin rows in compact Node admission (#154126)
Apply the remaining final Node budget before hosted tooling compaction, counting dist descriptors separately while preserving existing caps and coverage.
2026-09-20 18:02:58 -07:00
Peter Steinberger
5f97ef7536
fix(ci): reuse protected compile cache in build jobs (#154113) 2026-09-20 16:45:49 -07:00
Peter Steinberger
77ef790c1b
fix(ci): reduce Swift test build and fixture overhead (#154023)
* perf(ci): reduce Swift test build and fixture overhead

* test(ci): retain optimized Swift arguments in launcher proof

* test(state): import subagent record codec from its owner

Restore the identifier test after main moved the binder out of the SQLite
store. Retain the existing assertions and SQLite reader import.

Use the minimal import correction from openclaw/openclaw#154014.

Co-authored-by: Dallin Romney <dallinromney@gmail.com>

---------

Co-authored-by: Dallin Romney <dallinromney@gmail.com>
2026-09-20 14:18:37 -07:00
Peter Steinberger
2bf6257853
fix(ci): keep package builds out of the macOS app test budget (#153911)
* ci: run Swift package tests alongside the macOS app

* ci: account for both hosted Swift test phases

* test(ci): verify cache ownership across Swift phases
2026-09-20 11:55:08 -07:00
Peter Steinberger
ce8b5078ef
ci: collect measured timings for PR-only tooling (#153676)
* ci: collect measured timings for PR-only tooling

* test: align timing fixtures and workflow routing
2026-09-20 07:12:46 -07:00
Peter Steinberger
61bbd59ea8
ci: budget changed-extension envelopes from measured costs (#153515)
* ci: budget changed-extension envelopes from measured costs

* ci: fit measured extension envelopes within landed row caps
2026-09-20 04:24:50 -07:00
Peter Steinberger
123b8acdcd
ci: raise compact and Node matrix row caps slightly (#153519)
Raise compact plans to 90 rows and final Node matrices to 70 push / 130 PR rows. Preserve overflow rejection, runner policies, workers, timeouts, and test coverage.

Update capacity documentation and boundary tests. Keep the dependency-free preflight fixture focused on its import contract without relaxing its deadline or forbidden-import guard.

Validated by the green CI run on the refreshed PR head, exact scope checks, and the completed independent review.
2026-09-20 03:01:58 -07:00
Peter Steinberger
ba8a0e9bad
fix(test): restore Parallels 27.0.2 guest commands (#153469) 2026-09-20 01:14:17 -07:00
RoboClaw
87fdeff3bd
docs(release): preserve handoff and recovery guidance (#153322)
Co-authored-by: steipete <58493+steipete@users.noreply.github.com>
2026-09-19 18:29:33 -07:00
RoboClaw
741d65bcf7
docs(crabbox): document Blacksmith directory downloads (#153232)
* docs(crabbox): note Blacksmith directory download workaround

Co-authored-by: steipete <58493+steipete@users.noreply.github.com>

* docs(crabbox): document Blacksmith directory downloads

Worked on by:
- @steipete

Co-authored-by: steipete <58493+steipete@users.noreply.github.com>
OpenClaw-Publication: d47d6f6c-4aa6-4a07-9ec2-39f392317cb3

---------

Co-authored-by: steipete <58493+steipete@users.noreply.github.com>
2026-09-19 16:20:58 -07:00
Peter Steinberger
bef4c20d45
docs: accept chat screenshot attachment delivery 2026-09-19 12:35:03 -07:00
Peter Steinberger
dcb4847e48
feat(release): check publication gates before dispatch (#152470)
* feat(release): check publication gates before dispatch

* fix(release): accept empty draft bodies in publish preflight

* fix(release): wire publish preflight command and shared gate proof

* fix(release): discover draft state and await committed archive proof

* fix(ci): preserve release metadata types and catalog test lifetime

* fix(release): reuse resume authority and settle UI test phases

* style(release): satisfy typed metadata and UI fixture lint
2026-09-19 01:41:43 -07:00
Peter Steinberger
9c831c7551 chore: retire CLAUDE.md files
Remove legacy Claude instruction filenames and aliases. Preserve retained AGENTS.md instructions and retire obsolete alias guidance where needed.
2026-09-18 18:23:39 -07:00
Peter Steinberger
24caac494e
chore(autoreview): sync mirror to agent-skills a7e91e1 (#152039)
Canonical commit: a7e91e188fa0c3d692ac69b3137f24c6c3a2d2c9
Canonical PR: https://github.com/openclaw/agent-skills/pull/262
2026-09-18 13:04:27 -07:00
Peter Steinberger
f78391a5e7
fix(test): isolate revoked Telegram proxy connections (#151925) 2026-09-18 09:32:20 -07:00
Peter Steinberger
b7b55e3c1d
fix(test): clean up triage fixtures when startup fails (#151894)
* fix(test): clean up triage fixtures when startup fails

* fix(test): expose triage preload dependency to analysis
2026-09-18 09:03:28 -07:00
Peter Steinberger
d41099ad9d
fix(release): allow an explicit operator soak waiver for stable publication (#151880) 2026-09-18 08:17:11 -07:00
Peter Steinberger
c209576127
fix(test): make Telegram lease expiry checks deterministic (#151729) 2026-09-18 05:36:52 -07:00
Peter Steinberger
76604d06e9
fix(qa): preserve lease kind and failed spawn state (#151644) 2026-09-18 03:49:44 -07:00
Peter Steinberger
8a8ecf114b
test: publish launcher PID before cancellation readiness (#151642) 2026-09-18 02:32:03 -07:00
Peter Steinberger
fdbb2455ac
fix(qa): restore Telegram skill integration checks (#151632) 2026-09-18 01:56:51 -07:00
Ayaan Zaidi
ea6ad975a5
fix(skills): consolidate Telegram E2E setup and lifecycle (#151571)
## What Problem This Solves
The Telegram userbot harness had diverged into separate copies, and setup rejected authenticated Convex installations available through bunx or npx. This consolidates useful improvements into the repository-owned skill; all changes are isolated to that skill.

## User Impact
Maintainers get bounded CLI/auth discovery, readiness on the actual run's lease, isolated group setup, reliable cancellation/recovery, and recorder evidence that never blindly resends an unconfirmed message.

## Why This Change Was Made
Keep one canonical harness and concise instructions. Try existing `convex`, `bunx --no-install convex`, and offline npx before asking for authentication. Preserve the established broker binding, credential isolation, rich-message recording, and existing evidence fields. No OpenClaw product runtime or configuration contract changes.

## Evidence
- 130 isolated Node tests passed initially; 76 affected cases passed after review corrections, including the real Node/Python uncertain-send boundary and recovery from an actual restored archive.
- 25 Python driver tests and 9 recorder tests passed.
- Live documented runner: baseline failed Convex discovery; candidate discovered authenticated bunx and delivered the requested Telegram reply. A separate real run delivered its reply, then SIGTERM returned 143 with one successful broker release, removed Gateway scratch, and no remaining proof listeners.
- Skill files formatted; `git diff --check` passed.
- Explicit forum chat/topic selection remains; the unsupported broker-backed shortcut was removed rather than changing broker payloads.

CI waiting and a further review pass are explicitly waived by the maintainer after the code findings were fixed and the remaining live-proof gap was closed. Unawaited checks are not represented as passed.

Co-authored-by: Ayaan Zaidi <hi@obviy.us>
2026-09-18 13:45:03 +05:30