Commit graph

102935 commits

Author SHA1 Message Date
Peter Steinberger
8a981770fe
fix(tooling): verify Darwin zombie groups after EPERM
Darwin killpg excludes zombies and returns EPERM when no signalable group
member remains. Strict normal-exit cleanup mistook this terminal state for
a surviving process group, even when later drainage observed termination.

Reuse the existing zombie census for Darwin, including BSD state flags,
and reconcile a group reaped during the census with a fresh ESRCH probe.
Both observation and termination require positive completion evidence;
live, mixed, and uninspectable groups retain their cleanup failures.
Keep snapshot work inside the existing escalation and drainage deadlines.

Native same-uid Z and ZN groups reproduce EPERM for signal 0 and SIGKILL.
The original owner fails four regression cases; the repair passes all 13.
Shared owner/output tests pass (146 passed, one platform skip). The full
mac-elevation-host shard passes all 123 cases in 383.63s; final metadata
replay passes all 13 selected cases. Independent review is scoped-clean.
Production delta: +44 lines for bounded Darwin termination evidence.

Test cost: node scripts/run-vitest.mjs test/scripts/managed-child-process.termination.test.ts --maxWorkers=1: 5.83s wall; node scripts/run-vitest.mjs test/scripts/managed-child-process.tree.test.ts --maxWorkers=1: 5.05s wall.
2026-09-30 18:04:04 -07:00
Peter Steinberger
56a1ac4b06 fix(state): remove duplicate maintenance scope type export
A concurrent main fix for the #162200 check:madge-import-cycles failure
moved the same types and used an inline scope re-export. Rebasing the
standalone re-export onto that change retained both exports and caused
TS2300 at lines 20 and 28.

Keep the standalone type re-export and remove the duplicate inline export.
Both affected modules now exactly match the previously validated candidate.

Validation:
- node scripts/run-tsgo.mjs -p tsconfig.core.json --incremental --tsBuildInfoFile .artifacts/tsgo-cache/core.tsbuildinfo (TS2300 before; passed after)
- pnpm check:madge-import-cycles (0 cycles)
- node_modules/.bin/oxfmt --check src/state/openclaw-state-maintenance-context.ts src/state/openclaw-state-db-async-lifecycle.ts (passed)
- node scripts/run-oxlint.mjs --tsconfig config/tsconfig/oxlint.core.json src/state/openclaw-state-maintenance-context.ts src/state/openclaw-state-db-async-lifecycle.ts (passed)
- git diff --check (passed)
- Byte-for-byte comparison against the candidate that passed all 95 focused tests and 27 test-type shards (identical for both modules)
- Independent Codex autoreview (no actionable findings through P2)
2026-09-30 18:02:23 -07:00
Peter Steinberger
17b86a8238 fix(state): break the maintenance-context import cycle
Fix check:madge-import-cycles after #162200 introduced the cycle
openclaw-state-db-async-lifecycle.ts ->
openclaw-state-maintenance-context.ts ->
openclaw-state-db-async-lifecycle.ts through a type-only import.

Move OpenClawDatabaseMaintenanceScope and AgentSchemaMigration into the
maintenance context, and preserve the lifecycle module's public scope type
export. Runtime behavior and existing caller imports remain unchanged.

Validation:
- pnpm check:madge-import-cycles (1 cycle before; 0 cycles after)
- pnpm check:import-cycles (0 runtime value cycles)
- node scripts/run-tsgo.mjs -p tsconfig.core.json --incremental --tsBuildInfoFile .artifacts/tsgo-cache/core.tsbuildinfo (passed)
- node scripts/run-tsgo-core-test-shards.mjs (all 27 shards passed)
- node scripts/run-vitest.mjs src/state/openclaw-state-db-async-lifecycle.test.ts src/state/openclaw-agent-db.storage-migration.test.ts src/state/openclaw-state-db-readonly.worker-lifecycle.test.ts src/state/openclaw-state-lease-async.maintenance.test.ts src/state/openclaw-agent-execution.creation-witness.test.ts src/state/openclaw-state-db-resource-scope.test.ts src/state/agent-deletion-journal.snapshot.test.ts src/state/openclaw-state-db-readonly.snapshot-admission.test.ts (95 tests passed)
- node_modules/.bin/oxfmt --check src/state/openclaw-state-maintenance-context.ts src/state/openclaw-state-db-async-lifecycle.ts (passed)
- node scripts/run-oxlint.mjs --tsconfig config/tsconfig/oxlint.core.json src/state/openclaw-state-maintenance-context.ts src/state/openclaw-state-db-async-lifecycle.ts (passed)
- git diff --check (passed)
- Independent Codex autoreview (no actionable findings through P2)
2026-09-30 17:59:29 -07:00
Josh Avant
284b735a03
fix(android): allow internal releases to commit to Google Play (#162235) 2026-09-30 19:59:07 -05:00
Adarsh
f053128775
docs: macOS VM install step fails on a fresh VM without Node (#160239)
* docs: install Node before OpenClaw in the macOS VM guide

A fresh macOS VM ships without Node.js or npm, so the guide's npm install
step fails. Use the installer script, which provisions a supported Node
runtime, and keep the npm path for VMs that already have Node. Also check
that the Gateway background service exists, since the guide later runs the
VM headless.

* docs: stop the foreground Gateway before installing the VM service

Quick start onboarding keeps the Gateway in the foreground, so the guide must
tell users to press Ctrl+C before installing and verifying the background
service, matching the getting-started sequence.
2026-09-30 17:56:17 -07:00
Peter Steinberger
3e05f6b380 fix(state): break the maintenance scope type import cycle
#162200 left openclaw-state-maintenance-context.ts importing OpenClawDatabaseMaintenanceScope from the lifecycle module that imports values from it, so pnpm check:madge-import-cycles fails on main. The leaf now owns the scope type and its migration helper type; the lifecycle module re-exports the scope type, so importers are unchanged.
2026-09-30 17:54:56 -07:00
Linze Shi
0338bb3adb
docs: fix session tool gateway timeout link (#159796) 2026-09-30 17:45:21 -07:00
Peter Steinberger
45a3d17238
fix(sessions): keep one agent's queue cleanup from cancelling another agent's work (#162179)
* fix(sessions): scope provider, placement, and rewind queue cleanup

Use the existing lifecycle queue owner for provider precautions, worker
placement dispatch and move barriers, and transcript rewind/branch switches.
Preserve other agents' work in shared raw-key lanes while cancelling the
selected agent and session incarnation.

Revalidate existing authority at cleanup and retain rewind's original session
ID after history rotation. Boundary regressions exercise real settlement,
barrier, and Gateway handler entry points and fail again when the production
changes are temporarily reverted.

* fix(sessions): scope broad stop and interrupt queue cleanup

Cancel all incarnations only within the selected agent and conversation when no session ID is supplied. Keep dispatch, move, and rewind queue settlement unconditional after committed effects. Retain subagent raw-key cleanup pending target ownership in the registry.

* fix(sessions): keep keyless interrupt clearing its own lane

* test(sessions): fail keyless interrupt regression fast

* test(sessions): document keyless runtime fixture

* test(sessions): expect scoped interrupt lane clear

Co-authored-by: Peter Steinberger <steipete@gmail.com>
2026-09-30 17:45:05 -07:00
RoboClaw
9ad5d51876
fix(ui): archived sessions remain visible until reload (#162208)
* fix(ui): archived sessions remain visible until reload

Use canonical session observations for team rosters and retain lifecycle visibility for adopted catalog keys. Preserve bounded enriched paging, Undo, and observer error recovery.

Co-authored-by: steipete <58493+steipete@users.noreply.github.com>

* fix(ui): archived sessions remain visible until reload

Worked on by:
- @steipete

Co-authored-by: steipete <58493+steipete@users.noreply.github.com>
OpenClaw-Publication: 0661ad88-9479-4b9d-9525-97ae385da6dc

* test(ui): wire canonical sessions into refetch-volume fixture

Co-authored-by: steipete <58493+steipete@users.noreply.github.com>

* test: align roster and workspace recovery fixtures

Co-authored-by: steipete <58493+steipete@users.noreply.github.com>

---------

Co-authored-by: steipete <58493+steipete@users.noreply.github.com>
2026-09-30 17:43:41 -07:00
Peter Steinberger
b2d26e4a61
fix(code-mode): retire warm workers with their host
## What Problem This Solves

Completed Code Mode workers could remain alive after their creating Gateway or CLI command closed. The warm pool and idle-expiry timers had no host lifetime.

## User Impact

Closing a host now waits for its Code Mode workers to retire. Other live hosts keep their workers, and failed native cleanup remains owned for a later retry. Standalone executor calls close completed workers immediately instead of retaining an unowned warm pool. Suspended continuations retain their existing contract.

## Why This Change Was Made

Warm pools capture the existing resource host when created. Each host owns a scheduler scope and its admitted workers; idle expiry uses that scope. Reuse stays within the creating host. The four-worker bound, five-minute idle period, fresh VM contexts, and memory-pressure eviction remain unchanged.

Production: **+65 / -7 / net +58** (`src/**` and `extensions/**`, excluding tests and support). This growth supplies host isolation and joined retirement; this cut is not a LOC reduction. There are no persisted format or configuration changes. Update behavior: existing state remains compatible; the retirement ordering takes effect after restart.

## Evidence

Independent reviews of the production change and the test-fixture correction are clean through P2. Candidate: `6c475248c2ab5b21a3eb0b9e121b38e76d5bffe3`.

- The final source tree passed all 30 tests across the five Code Mode lifecycle/native sibling files in **22.31s pnpm wall**. The native warm-pool fixture now supplies a real host and joins its teardown; the real workers, test bodies, and worker-count assertions are preserved.
- The original production implementation failed the new standalone-worker retirement assertion as intended. The initial 14 lifecycle cases passed in **15.33s pnpm wall**. Full production changed checks passed in **1414.68s**; the final test-only delta passed its changed checks in **205.14s**.
- The final candidate build passed in **434.34s**. Its canonical package passed tarball integrity checks.
- The two native warm-pool failures during development were deterministic fixture mismatches: standalone calls no longer retain a pool, while those tests assert host-owned reuse. This was not a flake fix; no assertions, timeouts, retries, or worker implementations were weakened.

Blacksmith Testbox leases: `tbx_01m3t3g3afjjrmnqjhr0mqasv0` ([native and production proof](https://github.com/openclaw/openclaw/actions/runs/36779653190)) and `tbx_01m3taxfbahdc6c63a38n0fq50` ([final build, fixture checks, and update cell](https://github.com/openclaw/openclaw/actions/runs/36792095804)). The first lease completed at two hours during the fixture changed gate, after all 30 tests passed. The unfinished gate was completed on the fresh lease; the disconnected run is not counted as a passing gate.

The real published `2026.9.7` updater installed the exact candidate in **151s**. Both installed Gateways passed health and joined SIGTERM; the original disabled cron job, workspace bytes, authored configuration, and SQLite integrity were preserved. Doctor applied its existing implicit-main-roster normalization (`agents.entries.main = {}`). The first verification incorrectly expected the whole agents object to remain byte-shaped; source inspection confirmed the existing migration, and a separate read-only verification of the same original artifacts asserted that exact normalization and all remaining preservation checks. The update was not rerun. This cell covers installed update and shutdown; the worker tests above cover warm-worker custody.
2026-10-01 00:41:27 +00:00
Peter Steinberger
cf78e64434
fix(plugins): distinguish settled disposal faults from retained cleanup
Report settled inspection disposal errors without marking their managed
resources as retained. Keep rejected instance/cache prerequisites and drain
timeouts classified as retained cleanup. Preserve CLI resource release,
original error causes, borrowed ownership, and exactly-once disposal.

Fix the product regression introduced by bfdb432570 (#160181). Update
caller assertions to require the disposal failure while preserving successful
process close and subsequent healthy publication; document the contract.

Validation on an isolated AWS lease at main 0c1952c56e:
- Original 129-file plugin CI composition reproduced both failures before
  the fix, then passed 1,257 tests with one existing skip in 59.45 seconds.
- Four affected files passed all 60 tests in three consecutive runs
  (40.87s, 23.71s, 23.78s including runner overhead).
- New settled/rejected/pending inspection regressions took 79ms/19ms/17ms
  in the final focused run; the settled case fails on the original code.
- Core and plugins-platform test typechecks, scoped lint, oxfmt --check,
  git diff --check, and independent P0-P2 review passed.

Pre-commit formatting ran on the lease: oxfmt --check passed for all five
changed files. The local hook is skipped because this sparse worktree has
no node_modules.
2026-09-30 17:36:35 -07:00
Vincent Koc
9d79e79fa0
refactor(ollama): consolidate redundant test coverage (#162214)
Consolidate duplicate coverage at the owning boundaries and repair discovery, cancellation, cache, timing, and stream-order assertions. Preserve framing and reader lifecycle contracts in focused suites.

Related: #139428
2026-10-01 07:34:35 +07:00
Peter Steinberger
504ee39733
fix(test): memory flush suite teardown fails with ENOTEMPTY (#161536)
* fix(test): memory flush suite teardown fails with ENOTEMPTY

The runMemoryFlushIfNeeded suite stores each case's session database
inside its per-test rootDir. Session entry writes kick deferred SQLite
maintenance (reclamation Worker) and transcript reads retain history
Worker connections; both retire only when the cached agent database
handle closes, which the suite never did. A pass that opened the
database while fs.rm was recursing recreated files in rootDir, so the
final rmdir failed with ENOTEMPTY.

Close the agent database owners under rootDir with
closeOpenClawAgentDatabasesAsync before removing it.

* fix(test): drain memory suite agent databases once per suite

Closing each case's agent database owners in afterEach waits for the
maintenance and history Workers that case spawned, adding about 20 s
to the file on an 8-vCPU Linux runner. Keep one case directory per
test under a suite root and drain every owner under that root once in
afterAll before removing it; removal still never races a live owner.
2026-09-30 17:32:43 -07:00
AmelioMansour
170d2b5646
docs: clarify Telegram Node networking references (#159240) 2026-09-30 17:30:23 -07:00
Peter Steinberger
588b5d7d50
fix: avoid slow requester checks during Doctor repairs (#162200)
Doctor requester-authority validation copied the entire SQLite database in a synchronous worker on every authority callback (195 snapshot launches, 31 s for the configured-owner case; candidate Doctor 74 s). Maintenance now owns one reusable live reader; admission, physical file identity, and revocation are re-checked on every reused read, and the reader joins maintenance and database cleanup. Configured-owner case 31.4 s -> 3.6-4.6 s; snapshot launches <= 8 during admission, zero during maintenance.

Refs #161949
Refs #161867
2026-09-30 17:26:18 -07:00
Peter Steinberger
7e7c8d4ffe
test(infra): gate snapshot recovery on sibling close admission
After a staging Worker exits, native cleanup fences all sibling roots.
The fixture started the first retry before its sibling requested cleanup,
so a fast native reply could legitimately reject the first close before
the pending assertion.

Hold invocation of the real stop until both close requests are admitted,
then preserve the blocked-microtask proof that servicing only the sibling
finishes both closes. Extract the progress barriers into test support to
respect the existing file line cap without removing coverage.

The fast-stop probe fails the original pending assertion with unrequested
sibling cleanup and passes the fixed fixture. All 11 owner tests and the
5 retained worker-task sibling tests pass. Production LOC is unchanged.

Test cost: 273.37s wall (cold worker rebuild); node scripts/run-vitest.mjs src/infra/sqlite-snapshot-staging-owner.test.ts --maxWorkers=1
2026-09-30 17:20:58 -07:00
Mike Mondragon
32cc3f063d
docs: document image replay cost for vision models (#159205)
* docs: document image replay cost for vision models

Native-vision sessions skip image understanding by default, so raw
images re-attach on later turns inside the replay window at full
image-token cost. Document that default, its cost consequence, the
explicit tools.media.models[] entry that opts into description-based
replay suppression, and the preferredModel/enabled boundaries that do
not trigger it.

* docs: qualify image description and replay boundaries

---------

Co-authored-by: Patrick Erichsen <patrick.a.erichsen@gmail.com>
2026-09-30 17:20:12 -07:00
Peter Steinberger
94e17d521e
fix(macos): native tests flake on busy runners when fixed wait deadlines expire (#161616)
* test(macos): wait on signals instead of wall-clock deadlines

Native tests outside the Dashboard suites still bounded positive-outcome
waits with private 2-5 s deadlines (ContinuousClock loops, AsyncTimeout
wrappers, a sleeping timeout inside the Quick Chat observation wait, and
DashboardHTTPFixture listener readiness). On a saturated runner those
deadlines expire before the awaited work runs, e.g. the Cron source
adoption test failing at CronGatewayOwnershipTests.swift:236.

Generalize DashboardTestWait.state into TestWait: `state` re-checks
signal-less state every 10 ms, `observed` wakes on Observation changes
to tracked storage, and AsyncTestSignal lets fixtures wake waiters when
they record requests, payloads, listener states, or closed connections.
None has its own deadline; every migrated suite declares `.timeLimit`,
whose clock excludes parallel queue time. Absence windows are unchanged.

* test(macos): size wait time limits for saturated runners

Suite time limits count time a test spends queued on TestIsolation or a
starved main actor. On a saturated three-core runner single tests took up
to five minutes without any wait, so one-minute limits fired on scheduler
starvation instead of lost signals. Give the limit one owner, a shared
10-minute testWaitLimit trait beside TestWait, and use it for every suite
that relies on deadline-free waits, including the Dashboard suites.

* test(macos): keep synthetic browser sessions valid past the CI job

Dashboard browser-session fixtures minted credentials that expired 300 s
after creation. Under the longer wait limit a starved runner let one
cookie-store test outlive its fixture and fail with .expired. These tests
are not about expiry, so give their sessions a one-day validity shared as
Date.fixtureSessionExpiry.
2026-09-30 17:19:08 -07:00
Linze Shi
0ad243f936
docs: fix mapped hooks section link (#159062) 2026-09-30 17:15:46 -07:00
Peter Steinberger
b4dcb584b6
fix(test): prepare tunnel manager imports before lifecycle test deadlines (#161637)
createManager in node-worker-tunnel.test-support.ts lazily imported
node-worker-tunnel.js inside the test body (since #160415). That graph
reaches src/infra/runtime-process-entrypoints.ts, a compiled-subprocess
declaration whose Vite load hook waits for the invocation-wide worker
generation, so the first lifecycle test absorbed the whole preparation:
55 s with a warm local cache and a 120 s timeout cold under load.

Import the tunnel manager statically so preparation happens at
collection, outside test deadlines (the #156011 policy), and make
createManager synchronous at its 20 call sites. No timeout, retry, or
assertion changed.
2026-09-30 17:15:29 -07:00
Peter Steinberger
0c1952c56e
fix: root Vitest runs use the matching browser provider (#162137) 2026-10-01 00:13:31 +00:00
Peter Steinberger
76da51eae5 chore(ui): reuse the shared bundle for queued update tests
Queued update recovery tests mock build identities and protocol errors without
using source-module serving. Borrow the invocation's existing production preview
and remove the file's redundant private-server classification.

Preserve all 13 cases, assertions, and captures. Four-CPU Linux measurements fell
from 66.627s in the private project to 51.318s in an isolated bundled run; the full
parallel config measured 55.583s. Every selecting PR retains the same coverage.
2026-10-01 00:12:42 +00:00
Linze Shi
80f68b8862
docs: fix SDK migration section links (#158229) 2026-09-30 17:10:24 -07:00
Peter Steinberger
33ae32bfc1 test(gateway): assert describe precedes projection drain
The 2,048-session regression exceeded its 100 ms dirty-describe bound
while answering with 2,047 unrelated rows still dirty and no transcript,
usage, or bounded-tail reads during materialization. The earlier main
commit also exceeded the absolute bound under local load.

Capture dirty-row counts inside each actual RPC response and require one
correct successful reply independently for the initial and invalidated
requests. Preserve the zero-read, fixture, index-reuse, and cleanup
assertions, timing diagnostics, and existing timeout. The test protects
keyed reads completing ahead of unrelated bulk materialization.

Validation on Linux: matched parent/head comparison 20 runs each; dirty
describe medians 14.691/13.788 ms. Forced full-drain dependency fails the
reply-time assertion. The repaired file passes 20 pressured runs; whole
gateway-core passes 8,928 tests; all three original core/client shard
replays pass. Check-changed (types/lint/guards), boundary lint, and P2
review pass. No timeout or skip changes.

Measured single-file wrapper wall time under pressure: median 34.140s.
The original hourly case cost 10.035s; timing diagnostics remain available.
2026-10-01 00:07:19 +00:00
Deepak Jain (DJ)
d2965c01b0
docs(nextcloud-talk): explain message admission and reply flow (#157715)
Signed-off-by: Deepak Jain <deepujain@gmail.com>
2026-09-30 17:04:41 -07:00
Linze Shi
7db2084787
docs: fix utility model migration links (#157441) 2026-09-30 17:00:40 -07:00
Vito Cappello
6daee6e4bd
fix(gateway): restore CLI-backed exec with secret egress proxy (#160760)
* fix(gateway): pass admitted run instance to loopback-mediated exec

* test(gateway): cover egress-enabled MCP exec grants

Prove the real MCP HTTP grant, cached tool construction, egress-enabled process launch, retired-grant rejection, and a later run in the same session. Route the fixture through the existing native database-worker test group.

Co-authored-by: VACInc <3279061+VACInc@users.noreply.github.com>

* test(ui): wait for transcript resize before measuring centering

The real layout owner publishes viewport width asynchronously after resize. Wait for rendered width readiness, then sample all centers atomically while retaining visible-box and one-pixel assertions. The focused catalog E2E now passes in secretless AWS; the real-browser mechanism reproduces the prior 160px phase and still rejects a 2px displacement.

Co-authored-by: VACInc <3279061+VACInc@users.noreply.github.com>

* fix(ci): restore Doctor plugin repair lint budget (#160715)

Co-authored-by: VACInc <3279061+VACInc@users.noreply.github.com>

* fix(ci): restore config and node adapter lint budgets (#160843)

* fix(ci): keep config env and worker tests within lint budgets

* test(gateway): share node capacity and rejection fixtures

Co-authored-by: VACInc <3279061+VACInc@users.noreply.github.com>

* refactor(talk): combine transcript early-return guards

Preserve echo-first short-circuit evaluation while restoring the existing relay file-length budget. No runtime behavior or lint limit changes.

Co-authored-by: VACInc <3279061+VACInc@users.noreply.github.com>

* refactor(ui): share model setup activation payloads

Reuse the same discovery-to-activation projection for prepared and directly selected candidates, retaining optional-field omission and excluding discovery metadata. Restore the existing model setup page line budget without changing UI behavior or lint limits.

Co-authored-by: VACInc <3279061+VACInc@users.noreply.github.com>

* test(gateway): join chat execution before metadata assertions

The operator.write verbose-level fixture asserted agent.wait success after a one-second observation window without joining its detached run. Reuse the existing execution observer before the same scoped RPC, preserving the timeout and every metadata/permission assertion.

Verified the complete chat file on CI's Node 24.19.0 in network-isolated execution: 41 tests passed. Independent P0-P3 review and targeted lint passed.

Co-authored-by: VACInc <3279061+VACInc@users.noreply.github.com>

* test(gateway): join recap producer lifecycle events

Wait for actual model entry before inherited connection drain and for the relocated recap post-commit publication before inspecting persisted state. Preserve model cancellation, old-owner fencing, exact target/store checks and existing persistence assertions. New waits follow test cancellation rather than timing a cold worker through an observation poll.

Verified both edited files (25 tests), the original receiving-merge peer groups (377 and 164 tests), focused lint, both owning type graphs and independent P0-P3 review.

Co-authored-by: VACInc <3279061+VACInc@users.noreply.github.com>

* fix(infra): retain explicit SQLite launch paths when cwd is unavailable

Keep supplied launch and transport facts, resolve relative arguments only against a genuine captured cwd, and anchor native inspections to selected absolute operation paths when ambient cwd is gone. Refuse unresolved cwd-dependent inputs rather than redirecting them. Prove both native transports after real directory removal; cold Node Worker bootstrap remains a separate lifecycle limitation.

Co-authored-by: VACInc <3279061+VACInc@users.noreply.github.com>

* test(gateway): coordinate monitor fault injection with worker admission

Use the existing managed state write transaction for persistent trigger creation and removal. Controlled scheduleUnowned interleaving reproduces both raw-DDL lock failures and passes both managed-DDL variants without changing reload assertions or timeouts. Ensure final trigger cleanup cannot skip service cleanup.

Co-authored-by: VACInc <3279061+VACInc@users.noreply.github.com>

* test(gateway): keep monitor fault injection in fixture support

Preserve the exact managed DDL statements and cleanup lifecycle while shrinking the over-cap reload test. Line-cap, suppression, assertion and import-cycle guards pass; the 115-case file, focused lint, owning types and independent P0-P3 review pass.

Co-authored-by: VACInc <3279061+VACInc@users.noreply.github.com>

* test(update): select host runtime for Homebrew guidance

Keep the Homebrew-specific guidance fixture independent of the container used for isolated proof. Container cases and all guidance assertions remain unchanged.

Co-authored-by: VACInc <3279061+VACInc@users.noreply.github.com>

---------

Co-authored-by: VACInc <3279061+VACInc@users.noreply.github.com>
Co-authored-by: Peter Steinberger <steipete@gmail.com>
2026-09-30 16:59:39 -07:00
Peter Steinberger
e603e31e7c
fix(test): splash skeleton geometry check flakes when the viewport grows (#161509)
* test(ui): wait for the resized shell canvas before sampling skeleton geometry

page.setViewportSize resolves before the rendering update in which the
shell viewport owner publishes the new --shell-viewport-height. When the
viewport grows (900 -> 1440), the canvas stays capped at the old height
until that frame. Separate boundingBox calls could straddle the commit
(content at 900, composer at 1400: the -500 seen on scheduled main CI),
and the single-frame read from 531f0a229c could still sample the stale
900px canvas inside a 1440px viewport.

Wait on a ResizeObserver until the app canvas matches the requested
viewport height, then sample the geometry from that committed layout.

* test(ui): keep splash geometry comment within print width
2026-09-30 16:59:35 -07:00
Juampi
32680440ce
docs(firecrawl): link the Firecrawl site and where to get an API key (#157257) 2026-09-30 16:54:55 -07:00
Fede Kamelhar
c4e32ccfe1
docs(tts): explain that restricted tool profiles exclude the tts tool (#155950)
The tts catalog entry belongs to no built-in profile, so minimal, coding,
and messaging all remove it with no documented way back. Document the
tools.alsoAllow grant on the agent-tool page and the profile table, and
note that automatic TTS is unaffected by tool profiles.

Refs #126688
2026-09-30 16:52:23 -07:00
Peter Steinberger
02cf3f0d1d
fix(ui): keep Enter on a hover-opened submenu owner from selecting a row (#161393)
With the pointer resting on "Assign to…", Web Awesome opens its submenu on hover. Pressing Enter on the focused owner item then fell back to the submenu's active row and assigned the session to the first person (the current user). The dropdown key handler now only falls back to the level's active item when no dropdown item holds focus. The people-search browser test hovers before Enter and fails without the fix. It flaked in CI whenever the shared browser left the mouse over that item.
2026-09-30 23:49:52 +00:00
xiaziqiang
aae1a3adf3
docs: clarify session metadata and conversation history readers (#155449) 2026-09-30 16:47:30 -07:00
Peter Steinberger
c30cb63bb6
refactor(cli): centralize noninteractive Gateway option validation (#162203)
Keep Gateway option syntax admission in onboarding preflight and remove duplicate lower-level checks. Inline token persistence choices at their only assignment site, and let the local caller own whether daemon installation was requested.

Preserve late environment and saved-config validation, custom bind ordering, SecretRef retention, rerun precedence, environment-only password behavior, service auth admission, and runtime-pin ownership. Extend the existing preflight test table without adding runtime seams or waits.
2026-09-30 23:47:26 +00:00
RoboClaw
6dc7dbeee8
fix: resume compacted turns after a Gateway restart (#160128)
* fix: resume compacted turns after a Gateway restart

Reuse the validated persisted input when compaction has removed its model-context copy. Clear one-shot suppression for distinct queued inputs, preserving the transcript turn and ownership guards.

Co-authored-by: steipete <58493+steipete@users.noreply.github.com>

* fix: resume compacted turns after a Gateway restart

Worked on by:
- @steipete

Co-authored-by: steipete <58493+steipete@users.noreply.github.com>
OpenClaw-Publication: 9bb76237-cd52-47d4-8178-4be2a0fca461

* test: collect plugin retention fixtures in native tasks

Keep the existing eight GC calls and all retention assertions, but invoke collection before resolving the immediate callback so Bun does not retain completed Promise frames during the collection.

Co-authored-by: steipete <58493+steipete@users.noreply.github.com>

* fix: resume compacted turns after a Gateway restart

Worked on by:
- @steipete

Co-authored-by: steipete <58493+steipete@users.noreply.github.com>
OpenClaw-Publication: 0dde0ca5-2371-4363-b454-6504a78a8052

* fix: resume compacted turns after a Gateway restart

Worked on by:
- @steipete

Co-authored-by: steipete <58493+steipete@users.noreply.github.com>
OpenClaw-Publication: 83686719-0a5c-4ba0-96b9-39e28eeb7125

---------

Co-authored-by: steipete <58493+steipete@users.noreply.github.com>
2026-09-30 16:45:05 -07:00
Peter Steinberger
aac84b8a98
refactor(doctor): remove superseded CLI backend field migrations (#162153)
Remove session-argument and reliability sub-transforms that ran after the registered whole-map CLI backend migration. Preserve the supported-release map migration, configured model, and original config backup.

Independent review, ordered-registry tests, type checks, lint, and published-driver update proof passed on Testbox. Hosted CI failed unrelated Gateway abort cases. The fixture FIFO deadlock is fixed on main by 79f0e9bf12; the separate swarm-slot assertion remains unresolved. Exact tested-merge input and call-path qualification is recorded in the PR. Production +0/-32/net -32.
2026-09-30 16:37:16 -07:00
Peter Steinberger
f46120bca9
fix(doctor): preserve prior cron run-log archives (#162152)
* fix(doctor): preserve prior cron run-log archives

Reuse the shared hash-checked cron archiver after the history transaction commits. Preserve original bytes, including malformed UTF-8, without replacing an earlier archive. Surface archival failures through the existing Doctor warning path.

* refactor(doctor): retain existing run-log reads

Keep the original synchronous read API in this Doctor one-shot. The shared-archiver repair does not need an asynchronous read cutover. Preserve Buffer hashing and the registered entrypoint read-failure contract.
2026-09-30 16:33:54 -07:00
Peter Steinberger
dac3c15a0d
test(core,plugins): remove low-value tests (batch d119) (#162202)
* test(pr): deslop t0431 tests

* test(plugins): deslop t0430 tests

* test(mattermost): deslop t0415 tests

* test(telegram): deslop t0427 tests

* test(process): deslop t0424 tests

* test(state): deslop t0434 tests

* test(voice-call): deslop t0437 tests

* test(telegram): deslop t0428 tests

* test(gateway): deslop t0443 tests

* test(agents): deslop t0436 tests

* test: preserve lifecycle and authorization contracts in d119
2026-09-30 23:33:21 +00:00
Erick Kinnee
5c4041708d
fix(plugins): let session_end hooks read ended transcripts (#161451)
* fix(plugins): expose bounded ended-session transcripts

* fix(plugins): bound ended transcript reads by content and retirement

* fix(plugins): avoid ended transcript dependency cycle

* fix(sessions): discover closing resets without payload budgets

* test(gateway): decouple sign-in retry from provider prompts

Keep real WebSocket ownership, activation and revoked retry assertions while dropping incidental prompt order and timing diagnostics. Removing post-await authority validation makes the retained no-session assertion fail. Tracks local task openclaw-1e6.

---------

Co-authored-by: jalehman <550978+jalehman@users.noreply.github.com>
2026-09-30 16:33:04 -07:00
Hannes Rudolph
ca17064696
docs: publish release notes for v2026.9.7 (#162188)
* docs: publish release notes for v2026.9.7

* docs: publish release notes for v2026.9.7
2026-09-30 17:20:44 -06:00
Colton Harris
5ca201627d
docs(plugins): correct the rendered Control UI descriptor surfaces (#147599)
* docs(plugins): correct the rendered Control UI descriptor surfaces

registerControlUiDescriptor accepts session, tool, run, settings, tab,
and widget, but only tab and widget are consumed. The overview table
advertised four inert surfaces and omitted widget, one of the two that
actually render.

* docs: include rendered link-reader descriptors

---------

Co-authored-by: Patrick Erichsen <patrick.a.erichsen@gmail.com>
2026-09-30 16:15:58 -07:00
Peter Steinberger
fa9e014304
refactor(cli): share Gateway configuration prompts and decisions (#162191)
* refactor(cli): share Gateway configuration prompts and decisions

* test(cli): read setup port from shared defaults
2026-09-30 16:14:03 -07:00
Peter Steinberger
503de53d45
fix(update): explain the bridge release for retired config (#162150)
Share Doctor retired-format detection with candidate update admission, preserving the existing guard set and rejected-value privacy. Refusals now name the 2026.9.5 Doctor bridge.

Validated with independent review, focused admission tests, all 27 test-type graphs, lint, and published 2026.9.7 updater cells on Testbox. The package proof used the combined tree; the subsequent explicit-undefined correction passed focused validation. Hosted CI failed an independently attributed Gateway session-abort slot-release assertion outside the changed Doctor/update call paths; the failure and evidence are retained in the PR. Production +77/-59/net +18.
2026-09-30 16:12:00 -07:00
qingminlong
1edf7ea3ab
fix(docs): restore gateway CLI option anchors (#147038) 2026-09-30 16:10:26 -07:00
Peter Steinberger
97e5ea7be9 ci: reuse iOS smoke test build products
Build current PR smoke products once for testing, then run both focused simulator groups without rebuilding. Preserve test selectors, Debug settings, destination, separate group logs and historical/non-smoke actions.

Validation: cold/warm Xcode proof passed all 139 focused tests with unchanged app hashes; missing-product and preparation-failure controls fail explicitly. Full test-types, whole tooling config, changed checks, boundary lint, workflow validation, both manifest harness shapes and P2 review passed on the scoped candidate.
2026-09-30 16:07:45 -07:00
Kimi Yu
6f91eda9c7
fix(codex): restore persona on remote app-server connections (#162156) 2026-09-30 16:05:21 -07:00
Frosmans
3de2d4aa03
docs: explain skill readiness and visibility (#146951)
* docs: explain skill readiness and visibility

* docs: distinguish skill inventory from visibility
2026-09-30 16:03:34 -07:00
Peter Steinberger
b543922ae2
refactor(agents): share CLI candidate binding lifecycle (#162167)
Share placement-scoped CLI binding acquisition and successful result settlement across command and channel candidates. Move fork consume, restore, and successor binding writes to the existing guarded CLI binding-store owner.

Preserve writable versus read-only lookup, lifecycle and writer guards, native fork recovery, channel event ordering, and exceptional cleanup. Read command session state in the worker owner and recheck live authority and media activity after awaited reads to prevent replay.

Validation: focused CLI and session-store suites, channel/fallback/placement cross-owner tests, adverse media-replay control, both cycle checks at zero, and independent review clean through P2. No schema, configuration, protocol, or public SDK changes.
2026-09-30 23:02:54 +00:00
Hannes Rudolph
40c70a80fa
docs: publish release notes for v2026.9.6 (#162187) 2026-09-30 17:02:29 -06:00
Josh Avant
e09dfc8897
feat(android): add daily internal testing builds (#161812) 2026-09-30 18:01:07 -05:00
Peter Steinberger
5b9e29c20c
refactor(channels): finish shared draft-stream ownership (#162174)
Move draft generation, atomic receipt publication, sticky stale-create retirement, and current-message reset/retirement into the existing shared lifecycle. Discord uses the generation operations; Matrix reuses message reset while keeping its live-marker policy.

Preserve existing SDK parameters and operations; the API report confirms only additive members on the existing helper. Validation: 273 draft/entry-point tests, 1153 plugin contract tests, changed-file checks, both cycle checks at zero, and independent review through P2.

Hosted CI hit the pre-existing chat.abort-errors fixture deadlock also present in main run 36780090668. Main already contains the owning fixture fix, 79f0e9bf12. The PR records this inherited failure and its incidental SDK barrel import; security checks passed.
2026-09-30 15:59:41 -07:00