Commit graph

23587 commits

Author SHA1 Message Date
Peter Steinberger
4e9e68140a
fix(zalouser): avoid Gateway stalls during credential persistence (#156003)
* fix(zalouser): avoid Gateway stalls during credential persistence

Move QR credential saves and logout revocation onto the shared SQLite worker,
await their durable completion, and keep restoration behind pending revocation
on its originally selected state directory. Preserve the credential format,
revocation markers, refresh CAS, and named older-host capability fallback.

Carry a live hosted-wizard assertion through setup and into the final write
grant so a detached QR login cannot persist after its owner is retired.

Validation: 172 focused tests across worker storage, registered setup, synthetic
SDK transport, and core wizard lifecycle; complete independent P2 review clean.
Repository gates: 34 passed, 3 unchanged core gates reused. The inherited unused
TSGO_CORE_TEST_MAX_ROOTS export is the sole remaining gate failure on this base;
canonical fix 1619aff478 supplies its repair.

Regression controls prove the former host-SQL logout path, a missing final
wizard-lifetime guard, and state-root drift across a pending logout. No live
provider account, operator Gateway restart, schema change, or protocol bump.

* test(gateway): await recap publication before projection assertion

Wait for the existing owner's next publication after releasing the model result,
assert its captured state without filtering successful outcomes, and check one
direct describe response. Preserve the durable watermark and model-call assertions.
No production logic or test deadline changes.

Validation: 22 cases on the exact tested CI merge with Node 24.21; real writer
queue probe preserves the old updating recap until write settlement, then exposes
the published current recap. Owning type graph, typed lint, formatting, and P2
review pass. The initial CI log redacts the returned summary, so the change does
not claim that its historical failure was proved timing-only.

* test(discord): scope policy reload fixture activation

Limit the synthetic policy fixture to its Discord plugin so real reloads do
not cold-load unrelated runtime plugins. Require the reload to report
applied while preserving the active-turn, held lookup and revocation flow.

Exact CI merge-source proof passed locally in 15.3s. Root fixture types,
typed lint, formatting and independent review passed. No production code
or test timeout changed.
2026-09-23 08:23:57 -07:00
Vincent Koc
645f1bc4ef
feat(browser): add portable Lightpanda deployment and benchmarks (#154360)
* feat(browser): add opt-in Lightpanda semantic profiles

* fix(browser): reject direct selectors for Lightpanda profiles

* feat(browser): add portable Lightpanda deployment and benchmarks

* docs(browser): translate lightweight browser page title

* fix(browser): verify stale targets with a valid navigation control
2026-09-23 15:14:16 +00:00
Peter Steinberger
46e894f891
refactor(plugins): deslop codex and agent-harness plugins (#156297)
* refactor(codex): deslop codex

* refactor(plugins): deslop agent-harness provider plugins
2026-09-23 07:58:14 -07:00
Vincent Koc
b7eb012f64
feat(browser): add opt-in Lightpanda semantic profiles (#154359)
* feat(browser): add opt-in Lightpanda semantic profiles

* fix(browser): reject direct selectors for Lightpanda profiles
2026-09-23 22:50:37 +08:00
Peter Steinberger
9db35b7514
fix(agentsapi): package the bundled plugin icon, activity mark, and agent-runtimes category (#156549) 2026-09-23 14:08:59 +00:00
Peter Steinberger
482a4b2c49
perf(workers): accelerate warm cloud startup (#156380)
* perf(workers): reduce warm cloud startup work

* fix(workers): preserve snapshot recovery and refresh rules

* fix(workers): avoid lifecycle self-contention during admission

* fix(state): retain warm authority reads for Incognito lifetimes
2026-09-23 06:52:10 -07:00
Peter Steinberger
c8d3f81045
fix: isolate subagent concurrency per session (#156381)
* fix(agents): scope subagent concurrency to spawning sessions

Give each immediate spawning session its own configured execution budget,
including nested orchestrators, instead of making unrelated sessions share
one Gateway-wide queue. Preserve queue ordering, cancellation, hot limit
updates, and idle cleanup; retain the same parent identity during compaction.

Show aggregate activity with an explicit per-session limit in diagnostics.
Keep the existing child admission cap and Codex-native scheduling separate.

Verify with focused queue/runtime/UI tests and an isolated real-provider
Gateway case that holds one session's slot while another session and its
nested child run, then awaits all completion acknowledgments.

Closes #156370

* test(telegram): match sticker fixture to ingress dispatch

Register the native sticker pipeline's provider runtime prerequisite for
local and CI execution. Preserve the same Request JSON decoder while
exercising the registered bot handler, as durable webhook ingress does
after acknowledgment, and join the held turn before fixture cleanup.

The original CI file order reproduced the grammY timeout. Runtime
preparation alone was insufficient; the test's generic webhook adapter
imposed a deadline that is absent from the production dispatch boundary.
Keep the cache, media, and model-admission assertions unchanged, with the
real webhook acknowledgment contract covered by its existing test.

* test: repair packaging and runtime-owner CI fixtures

Copy the newly imported check-limits helper into the trusted packaging
harness so its dependency-free startup assertion reaches the CLI.

Compare prepared Telegram runtime files against the same config that
supplies the expected files. Keep worker-envelope coverage, packing,
concurrency and execution-budget assertions unchanged.

Both regressions reproduced before their fixes. The packaging owner passed
64 tests and ten targeted packing/prerequisite checks passed afterward.

* test(matrix): await follow-up adoption while active turn is held

Replace the 500 ms post-routing poll with resolver-owned adoption completion. Preserve foreground FIFO settlement by releasing the active turn before joining handlers, and settle fixture gates on timeout.

Validation: all four owner cases pass (37.31 s wrapper, 3.32 s bodies); selected changed checks and independent P0-P2 review pass.

* test(ci): recognize stopped Linux fixture process groups

Reuse the existing all-thread process-group assertion after detached fixture leaders are joined. Kernel kill probes also include exited zombies awaiting reaping; retain rejection of live descendants and uncertain observations.

Validation: 64 Docker scheduler cases and 18 retained-zombie/census cases passed on Linux Testbox; selected changed checks and independent P0-P2 review pass. The original CI process state did not reproduce in the original-order diagnostic.

* test(ui): count roster refreshes independently of child lookups

Install the browser clock before navigation and match roster sessions.list requests by includeGlobal. Preserve the 4999/+2 ms event window and avatar checks, and require exactly one roster refresh after the boundary.

Validation: reproduced the failing count and traced its extra request to a spawnedBy child lookup while the roster timer remained pending. All 14 browser cases pass (38.15 s; changed case 1.063 s). Selected changed gates and independent P0-P2 review pass.
2026-09-23 06:46:57 -07:00
Peter Steinberger
d0de886394
test: advance admitted Telegram media retry waits (#156413) 2026-09-23 06:23:22 -07:00
Peter Steinberger
4205a47a63
fix(telegram): defer unused image discovery and settle test updates (#156482)
* test(telegram): settle updates before fixture cleanup

* perf(media): defer registry discovery for explicit image models
2026-09-23 13:21:14 +00:00
Serg
4f43e47518
fix(xai): give new Grok releases thinking levels and image input (#156397)
Recognize plain, latest-alias, and dated Grok release IDs using xAI's documented capability ranges, so newer releases receive supported thinking levels and image input without another exact-ID update. Keep Grok 4.20 and variant suffixes on their existing conservative paths, and share the release rule with Responses tool defaults.

Validation: 56 focused regression tests passed on the refreshed head; independent review found no P0/P1 issues.

Co-authored-by: Takhoffman <781889+Takhoffman@users.noreply.github.com>
Co-authored-by: Tak Hoffman <781889+Takhoffman@users.noreply.github.com>
2026-09-23 13:18:28 +00:00
Peter Steinberger
5a15cf46eb
test: use native equality for Matrix archive bytes (#156490) 2026-09-23 06:12:01 -07:00
Peter Steinberger
d61bf615bb
perf(infra): keep hot Workers warm and expose churn
## What Problem This Solves

Intermittent tasks repeatedly recreate healthy Workers after the existing 30–60 second idle timeout, paying isolate startup costs and leaving resident memory behind after teardown. Existing metrics cannot distinguish which scripts are starting and exiting.

## User Impact

Hot task pools retain one Worker for five minutes of inactivity after an idle-retired Worker is promptly needed again. Node Code Mode retains its existing one-worker cache for five minutes. Operators can graph `openclaw_worker_started_total{script}` and `openclaw_worker_retired_total{script,reason}` through the existing diagnostics heartbeat.

## Why This Change Was Made

The existing retirement owner now owns idle timers and the single warm slot. Other slots keep their ordinary timeout; native cleanup, cancellation, pressure retirement, and rotation retain their existing ownership barriers. Existing idle GC collects released payloads in place. Code Mode still checks the runtime entry and heap limit before reuse.

The existing Worker resource registry owns cumulative lifecycle counts, including bundled direct Worker constructors. Counts advance on successful creation and confirmed native exit, survive exporter restart, and use bounded labels. Eval/unknown Workers remain `other`; nested Workers and V8-native threads remain outside the parent registry. Memory projection moves into a focused Prometheus module to keep the existing service below its growth ratchet.

No new configuration, database changes, dependencies, or managed-service environment changes. Updating installs the new runtime behavior through the normal restart; no migration or operator-state mutation is needed.

## Evidence

Real production-owner calls, ten verified synthetic tasks at 70-second simulated idle intervals (630 seconds observed), with real Worker creation and native teardown:

| Script | Starts before → after | Starts/min before → after | Task-run wall before → after |
| --- | ---: | ---: | ---: |
| `cron-stream-matcher.worker.js` | 10 → 2 | 0.952 → 0.190 | 485 → 111 ms |
| `git-operation.worker.js` | 10 → 2 | 0.952 → 0.190 | 3,076 → 824 ms |
| `code-mode-node.worker.js` | 10 → 1 | 0.952 → 0.095 | 858 → 187 ms |

The original source fails both new warm-window regressions; the pool creates a third Worker and Code Mode creates a second pool. In-place GC proof releases a 138 MiB payload to a 7.4 MiB heap while reusing the Worker.

### Allocator experiment and limits

Requested the largest advertised Blacksmith Linux class, 32 vCPUs, through a temporary proof-only workflow branch. The resulting guest exposes **8 online CPUs**, so this does **not** fulfill ≥32-vCPU validation or prove production memory behavior. Node 24.19.0, glibc 2.39, 32 concurrent Workers, 1,000 identical tasks touching 8 MiB each:

| Fresh process | Worker starts | Wall | RSS minus aggregate heapTotal increase after all Workers exit |
| --- | ---: | ---: | ---: |
| Create/terminate each task | 1,000 | 8.11 s | +639.59 MiB |
| Same churn, `MALLOC_ARENA_MAX=2` | 1,000 | 7.54 s | +423.86 MiB |
| Reuse 32 Workers | 32 | 3.08 s | +512.05 MiB |

The arena control reduces final residue 33.7%; reuse reduces it 19.9%. The churn proxy's post-100 slope is 102.63 KiB/task, versus 19.26 KiB/task with the arena control. A comparable retained-worker proxy slope is **not interpretable**: V8 heap capacity shrinks 992 MiB while RSS remains about 716.5 MiB, producing negative subtraction values and a false slope. Retained raw RSS grows only 0.152 MiB from tasks 400–1,000. These short exploratory results do not establish a many-core native-memory slope or allocator-contention tradeoff; consequently this PR does not set `MALLOC_ARENA_MAX`. Live deployment/production attribution remains a follow-up.

### Validation

Provider: `blacksmith-testbox`; profile: `openclaw-check`; lease: `tbx_01m36makaczyy9pds4yt5t2pem`; [workflow run](https://github.com/openclaw/openclaw/actions/runs/35834466664). One active remote command at a time, normal synchronization throughout. All test/typecheck/build work ran remotely.

- `node scripts/check-changed.mjs --base HEAD -- <all 25 changed paths>`: passed in 43m18s, including all 25 core-test typecheck graphs, core/extension production and test types, lint, documentation, SDK boundaries, dead exports, and import cycles.
- `pnpm test <file> --maxWorkers=1` for every changed test file, then `node scripts/run-vitest.mjs` for the five affected lifecycle/transport siblings: passed; 153 tests across 11 files. Final combined remote command: 59.2s. Original-source targeted runs failed both new regressions for the intended extra Worker/pool creation; candidate runs pass.
- The production-owner before/after benchmark command passed, including synthetic output assertions and native-exit cleanup for all three scripts. The allocator benchmark completed all three fresh-process cases with identical checksums.
- `pnpm build` followed by `OPENCLAW_LOCAL_CHECK=0 node --import tsx scripts/profile-extension-memory.mts --extension telegram --skip-combined --concurrency 1`: passed in 3m54s combined. All 95 plugin distributions built; bootstrap and 53 native control-plane module checks passed. Telegram isolated import exited cleanly, 160.09 MiB peak RSS (113.85 MiB above the profiler's empty-process baseline).
- Independent Codex review: scoped-clean, no actionable P0–P2 findings. `git diff --check`: passed. Net production growth: 146 lines; no test-only production seam.

Measured single-worker test command cost (seconds, including setup):

| Changed file | Wall |
| --- | ---: |
| `src/infra/worker-task-pool.test.ts` | 20.33 |
| `src/infra/worker-cpu.test.ts` | 7.48 |
| `src/logging/diagnostic-memory.test.ts` | 6.58 |
| `src/agents/code-mode-node.lifecycle.test.ts` | 5.21 |
| `extensions/diagnostics-prometheus/src/service.event-loop.test.ts` | 1.81 |
| `extensions/telegram/src/telegram-ingress-worker.test.ts` | 5.18 |

The two added behavior regressions use a fake idle clock; the real-worker pool case took 45ms and the Code Mode lifecycle case 35ms. CI timing will be updated after the head run exists. Crabbox reported an external runner-portal sync timeout after successful changed-check, final-test, and build commands; their underlying commands and reported run status succeeded.

Setup failures were diagnosed rather than counted as product failures: the transport checkout initially lacked `tsx` (fixed with remote frozen install), has no cgroup-v2 `cpu.max`, and lacks `origin/main` (changed checks use the actual original `HEAD` base plus exact changed paths). The initial benchmark fixture awaited an unrelated/unreferenced Worker exit; the corrected fixture uses the pool's resource-release boundary. The first changed gate identified line-cap growth; moving idle timing into its owner and extracting memory projection resolved it without an exception.
2026-09-23 12:53:07 +00:00
Vincent Koc
904f9c03a3
feat(diagnostics): trace completed harness commentary (#156343)
* feat(diagnostics): trace completed harness commentary

* fix(diagnostics): keep commentary event type internal

* fix(diagnostics): account for commentary in stability projection

* refactor(diagnostics): separate event fields and snapshot queries

* test(ui): provide scroll pane identity in ownership fixture
2026-09-23 20:33:15 +08:00
Peter Steinberger
cfb5802705
test(telegram): prevent idle socket races in HTTP dispatch tests (#156466)
Close fixture responses so grammY cannot reuse a socket as the native HTTP server expires its idle lifetime. Preserve production ambiguity handling and the terminal-delivery assertions.

Refs #155040. Reproduced the PR #156189 run 35834470541 typing-only failure at the native idle boundary; the same case passes with the fixture header. Linux proof includes 30 delayed-response stress trials, 243 tests across 13 files, check:changed, and independent review.
2026-09-23 12:03:32 +00:00
Peter Steinberger
71cf3a4a16
perf(session-share): keep chat catalog startup responsive (#156400)
Share concurrent Session Share node listings under service-owned authority and bound progressive chat-startup waits to five seconds. Retain compatible pending pages and publish completed refreshes through the existing catalog lifecycle, preserving current caller filtering. Targeted metadata, pagination, and non-progress clients retain complete-response semantics.

A synthetic six-caller progressive burst behind a 30-second node response improves from 30000 ms p99 and six RPCs to 5000 ms and one RPC. Testbox validation passed 41 focused tests, 64 existing gateway/service tests, typechecks, lint, and architecture checks. Independent review found no remaining actionable issues.
2026-09-23 11:49:31 +00:00
Vincent Koc
bab3136d3c
fix(plugins): withhold Incognito content from observation hooks (#156417) 2026-09-23 19:41:31 +08:00
Shakker
17d2faf6b7
feat: enforce role model limits in native agents and Visitor Access (#154893)
Carry operator role model limits through native Codex parents, child agents, restored work, reviews, and model-calling tools, with current-source revocation and lifecycle cleanup. Require an explicit Visitor model policy while preserving the documented staff-mode limitation and existing configured service authority.
2026-09-23 12:28:30 +01:00
Peter Steinberger
af7420cc36
test(telegram): await typing request before renewal assertion (#156454) 2026-09-23 04:21:37 -07:00
Peter Steinberger
6dc9d86447
fix(qa): release terminal children after requester settlement
The mock released child results when the parent's HTTP response was sent,
before the parent execution owner closed. Timestamp-prefixed all-settled
inputs also missed the fixture matcher and produced a generic response.

Correlate pending children with the exact runtime parent and use the existing
scenario wait loops to verify sessions.list reports that parent done, inactive
and not aborted. Keep the existing budgets and every direct-delivery, exact-send,
privacy and restart assertion. Recognize timestamped settlement inputs without
accepting quoted history. Migrate all mock-server and scenario callers together.

The six changed standalone test files each pass 20 full runs, with 91 focused
Gateway/QA cases and 43 HTTP sibling cases passing. Linux original profile 4
passes three times (27/27 scenarios, zero skipped), plus native Telegram and
QA-channel/empty flows. Actions proof: 35846151730. Single-worker file walls:
parser 2.47s, gate 2.48s, handoff 11.67s, routing 15.33s, surface 24.26s.

Types, scoped type-aware lint with a negative canary, formatting and P2 review
pass. Full lint declaration preparation hit ancestor-install isolation and used
the requested scoped substitute. Installed-package upgrade/rollback was not run
because no candidate tarball was configured; its migrated callbacks typecheck.
2026-09-23 04:15:55 -07:00
siri1410
12fe32ebbb
fix(discord): remove stale progress drafts after queued follow-ups (#156186)
Closes #149640

## What Problem This Solves

Fixes: a stale processing progress message stays in Discord when a follow-up queued behind an active turn creates a draft after its inbound dispatch has returned.

## User Impact

The draft is cleared when the queued turn settles, the same as Telegram and Slack. No configuration changes or new plugin APIs are required. Failed, held, ambiguous, and partial queued finals remain a shared gap: their origin-route delivery outcomes are not available to channel settlement cleanup. Preserving drafts for those outcomes needs a separate core fix and is not claimed by this PR.

## Why This Change Was Made

Discord now uses the existing queued-follow-up settlement callback to run its draft cleanup, as Telegram and Slack already do. Cleanup belongs to the draft lifecycle rather than a new final-delivery hook in the public reply options.

### Maintainer decision

Ayaan's decision, September 23, 2026, as relayed by the coordinating maintainer:

> Land #156186 with parity to Telegram/Slack.

This accepts settlement-based cleanup for this Discord-only fix. The hosted finding about failed, held, ambiguous, and partial queued finals is acknowledged as a shared gap for a separate core fix, not dismissed as already solved. This PR changes only two Discord files; it changes no Slack files.

## Evidence

- The contributor's original late-queued-draft regression failed against the pre-fix production code.
- The regression adapted to the existing settlement callback also failed before the repair: the draft ID remained `preview-next` after settlement.
- Complete Discord draft-progress file: 30 tests pass (22.88 seconds). Draft-recovery file: 21 tests pass (20.23 seconds), including ordinary dispatcher delivered-error and failed-final retention. Those recovery tests do not establish retention for queued origin-route failures.
- Live Discord proof passed on `4857854a66de8824f88f87607fcaace265cbef3f`, using Convex-leased bots, an isolated Gateway, and the deterministic mock provider. The second tool turn was sent while the first progress draft was visible. Discord Gateway events showed its distinct progress draft, final reply, and draft deletion; REST history then confirmed the final remained and the draft was gone.
- The live doctor's readiness check passed. Owned fixture cleanup completed with zero failures and no unresolved operations. No manual Discord rendering or screenshot claim is made.

Redacted live event sequence (UTC, September 23):

| Time | Observed Discord event |
| --- | --- |
| 05:53:09.126 | First turn's progress draft created |
| 05:53:09.695 | Follow-up trigger sent while the first turn was still running |
| 05:53:18.758 | First turn's final reply created |
| 05:53:21.428 | Queued turn's distinct progress draft created |
| 05:53:25.485 | Queued turn's final reply created |
| 05:53:25.928 | Queued turn's progress draft deleted |
| 05:53:29.527 | Fixture teardown completed successfully |

Co-authored-by: Ayaan Zaidi <hi@obviy.us>
2026-09-23 16:30:48 +05:30
Peter Steinberger
760dc1498d
fix: avoid blocking channel operations on stored account credentials (#156267)
* refactor(channels): await operational account resolution

Prefer optional asynchronous account and configured-state hooks for runtime
operations while retaining synchronous SDK compatibility. Read Matrix
credentials through the worker-backed store and honor the supplied environment
when selecting its default account.

Preserve bootstrap metadata decisions and ordered target validation. Recheck
request, registry, and configuration authority around logout preparation and
teardown so reloads cannot redirect an older logout onto its successor.

* fix(channels): preserve logout results across async account changes
2026-09-23 03:27:50 -07:00
Ayaan Zaidi
fad0547d78
fix(catalog): new models never reach installs when one models.dev provider disappears (#156352)
## What Problem This Solves

Fixes: new models (for example GPT-6 Sol and Luna from #155967) never reach existing installs because hosted catalog publishing has failed on every scheduled run since 2026-09-21, after models.dev stopped publishing the `kimi-for-coding` provider.

## User Impact

User impact: once this lands and the catalog republishes, installs on older builds get current model metadata on their next Gateway restart, as intended, without a release. A renamed or broken upstream provider now degrades only that provider's models.dev hydration instead of freezing catalog updates for all 44 providers. No config, schema, or runtime change.

## Why This Change Was Made

- The Kimi plugin mapped `kimi` to models.dev `kimi-for-coding`, which no longer exists. models.dev split it into `kimi-code-plan-global` (api.kimi.ai) and `kimi-code-plan-cn` (api.kimi.com). OpenClaw's provider uses `https://api.kimi.com/coding/`, so the mapping now points at `kimi-code-plan-cn`. Its four model ids match the manifest.
- The publisher threw for any missing or malformed mapped upstream provider, which failed the whole publication. It now logs a warning naming that provider and publishes its manifest rows unhydrated. A models.dev outage or a non-object response stays fatal and still preserves the last published artifact.
- Docs describing hydration failures are updated to the new contract.

## Evidence

- Production failure: `openclaw/catalog` "Publish catalog" runs on 09-22 and 09-23 end with `models.dev catalog missing or malformed for provider kimi-for-coding` / `[publish-model-catalog] FAILED (exit 1)`. The published `models/v1/catalog.json` was last changed on 09-18 and has no `gpt-6-sol`.
- Downstream symptom: a Gateway built before #155967 kept a synthesized placeholder for configured `openai/gpt-6-sol` (128K, text-only, no reasoning) after restart, so its thinking level clamped to `off`.
- Live dry run of the workflow command (`publish-model-catalog.mts --pricing`) against current models.dev and pricing feeds:
  - `origin/main` (`e7fbe5389e0`): reproduces `missing or malformed for provider kimi-for-coding` / `FAILED (exit 1)`.
  - This branch: exit 0, `providers=44 models=1062`, `models.dev provider=kimi added=0 filled=0 skipped=0`. The written bundle contains `openai/gpt-6-sol` and `gpt-6-luna` with `reasoning: true`, text+image input, 1,050,000 context window and 272,000 context tokens. The 10 remaining warnings are pre-existing pricing gaps.
- End to end on a Gateway image built before #155967 (fresh state, no network, catalog served locally): `openclaw models refresh` against the currently published catalog reports `unchanged (44 providers, 1033 models; generated 2026-09-18)` and has no `openai/gpt-6-sol` row. Against this branch's published output it reports `updated (44 providers, 1062 models)`, and `models list` returns `openai/gpt-6-sol` as text+image with a 1,050,000 context window and 272,000 context tokens.
- Tests: `test/scripts/publish-model-catalog.test.ts` 73/73 pass. The new case fails on the original script with the production error message, and passes with the fix. Outage and malformed-feed cases still reject without changing the bundle.
- Test cost: `pnpm test test/scripts/publish-model-catalog.test.ts --maxWorkers=1` took 21.1 s wall on an Apple M3 Ultra (Vitest duration 18.61 s, 73 tests; 62% transform, 34% tests). CI run 35837009655 on this head passed all 117 jobs; its logs report no per-file timing for this test.

Co-authored-by: Ayaan Zaidi <hi@obviy.us>
2026-09-23 15:45:25 +05:30
Vincent Koc
f34c36899c
fix: exclude incognito turns from automatic memory capture (#155594)
Co-authored-by: Vincent Koc <vincentkoc@ieee.org>
2026-09-23 17:53:28 +08:00
Peter Steinberger
99cfeee840
fix(anthropic): stop retaining MCP configs in the session catalog (#156266)
Project bounded Desktop metadata before cache and overlay retention, preserving the independent CLI index contract. Detach trimmed strings so short display values cannot keep large backing strings alive.

Synthetic full catalog proof reduced additional retained heap from 64.49 MiB to 0.482 MiB while preserving all 48 visible rows and metadata. Fixes #155754.

Co-authored-by: Peter Steinberger <steipete@gmail.com>
2026-09-23 02:44:44 -07:00
Peter Steinberger
3e4ab6bed8
refactor: retire pre-June import and verification compatibility (#156285)
* refactor: retire pre-June import and verification compatibility

Remove pre-June task, flow, and plugin-state sidecar imports, obsolete
runtime chunks, package/installer validation exceptions, the old MCP
attachment fallback, and the April self-upgrade lane with its orphan helpers.

Leave retired data files untouched and document migration through 2026.6.1.
Preserve June-and-later contracts and September delivery recovery receipts.

Refs #156190

* docs: route legacy upgrades through 2026.9.5

* test: await Telegram fixture lifecycle events

Replace the setup stopwatch with the actual stop event or terminal run outcome. Keep cancellation assertions and outer execution bounds, and prove early terminal outcomes fail promptly.
2026-09-23 02:22:22 -07:00
Vincent Koc
ca78e11c09
fix(qa): stabilize runtime tool evidence (#120353)
* fix(qa): wait for native patch transcript completion

Punchcard-Session: golden-valley-workshop-br

* fix(qa): make sessions spawn fixtures deterministic

Punchcard-Session: golden-valley-workshop-br

* fix(qa): bound sessions spawn evidence

---------

Co-authored-by: Dallin Romney <dallinromney@gmail.com>
2026-09-23 01:37:27 -07:00
Vincent Koc
7f61c6cf84
feat: extend Ultra harness mode across supported runtimes (#155393)
* feat: extend Ultra harness mode across supported runtimes

* fix: preserve Ultra across model capability boundaries

* test: align remaining session fixtures with Ultra harness mode

* test: bind Ultra session fixture to prepared provider policy

* test: align Ultra capability fixtures

Keep the full Gateway assertions while isolating observed native efforts from known model capability floors. Use an unsupported native effort in the compaction clamp fixture now that Ultra is a supported harness mode.

* fix: honor configured Ultra capabilities across runtimes

Compose configured model overrides with captured catalog rows for CLI and cloud execution in omitted, merge, and replace modes. Keep route-bound capabilities with their transport owner, including when applying a previously unspecified route.
2026-09-23 15:56:42 +08:00
Sarah Fortune
051c8871bc
feat(agentsapi): enable native live web search (#156306)
Co-authored-by: Sarah Fortune <sarah.fortune@gmail.com>
2026-09-23 07:54:02 +00:00
Ayaan Zaidi
de5e31b895
refactor(telegram): remove redundant test coverage (#155040)
Consolidate Telegram tests at their behavioral owners and remove redundant mock inventories and private test seams while retaining distinct routing, authorization, replay, media, and streaming regression checks.

Align retained delivery tests with the landed shared final-delivery lifecycle. Respect replyToMode off when finalizing a quoted preview so the answer is edited in place instead of deleted and re-sent.

Exact-head hosted CI passed. The PR records current-head Telegram Test Server proof with reverted controls and the remaining validation limitations.

Work session: https://team.openclaw.ai/chat/roboclaw/dashboard/7320cd5b-9414-4cf6-99fe-0abb9a51df58

Co-authored-by: obviyus <22031114+obviyus@users.noreply.github.com>
Co-authored-by: steipete <58493+steipete@users.noreply.github.com>
2026-09-23 00:50:06 -07:00
Peter Steinberger
fdf8d86f8f
fix(crabbox): avoid rejecting valid macOS worker signatures (#156301)
Feed captured codesign metadata directly to grep so successful matches cannot leave printf failing with SIGPIPE under pipefail. Extend the existing desktop fixture with a bounded diagnostic tail that exposes the race deterministically.
2026-09-23 07:47:24 +00:00
Vincent Koc
4875b96d5b
fix(codex): deliver native child results after parent yields (#156247)
* fix(codex): deliver native child results after parent yields

Retain accepted native child completion ownership beyond the admitting turn and through Gateway restart drain. Bind event persistence, task transitions, history recovery, and delivery to the original assignment receipt, releasing execution holds after terminal persistence and the first handoff.

Keep thread-only wait snapshots from mutating successor assignments. Document the optional exact-assignment transition contract for custom detached runtimes.

Release-note: Fix native Codex child completions lost after parent yield or during Gateway restart drain.

* fix(tasks): keep transition contracts independent of execution

* test(tasks): isolate legacy dispatch from native settlement

* fix(codex): validate task settlement before native admission

* test(tasks): prove native completion authority at final effects

* test(tasks): pass context to completion authority scenarios

* test(tasks): settle completion request owners before assertions

* test(copilot): complete task runtime capability fixture

* test(tasks): match completion harness runtime contracts

* test(tasks): align authority harness with typed lint

* test(copilot): extract native task failure fixture
2026-09-23 07:39:35 +00:00
Peter Steinberger
741a1c84cd
feat: use protected model credentials in Crabbox applications (#156207)
* feat: use protected model credentials in Crabbox applications

* refactor: use the shared model egress lazy runtime owner
2026-09-23 07:36:15 +00:00
Peter Steinberger
0fa2ffd812
fix(models): identify native app catalog sign-in failures (#156278)
* fix(models): identify native app catalog sign-in failures

* fix(models): preserve catalog outcomes for arbitrary runtime IDs

* test(acpx): assert structured native catalog results
2026-09-23 07:32:00 +00:00
Peter Steinberger
dc063e7bc0
chore(deps): update fs-safe to 0.18.2 (#156262) 2026-09-23 07:16:19 +00:00
RoboClaw
5e5f9c2755
fix(anthropic): finish Opus 5.5 defaults and thinking display (#156093)
* fix(anthropic): finish Opus 5.5 defaults and thinking display

Co-authored-by: fuller-stack-dev <263060202+fuller-stack-dev@users.noreply.github.com>

* fix(anthropic): preserve authored model selections during CLI setup

Co-authored-by: fuller-stack-dev <263060202+fuller-stack-dev@users.noreply.github.com>

* test(qa): assert rolling Opus defaults without redundant resolution

Replace the ineffective real-time deadline probe with deterministic fake-clock coverage beyond the watchdog grace.

Co-authored-by: fuller-stack-dev <263060202+fuller-stack-dev@users.noreply.github.com>

* test: align remaining Opus default fixtures

Keep explicit version pins, assert immutable numeric pricing, and consolidate PDF selection fixtures without dropping cases.

Co-authored-by: fuller-stack-dev <263060202+fuller-stack-dev@users.noreply.github.com>

---------

Co-authored-by: fuller-stack-dev <263060202+fuller-stack-dev@users.noreply.github.com>
2026-09-22 23:59:18 -07:00
Sarah Fortune
51a27d44d5
refactor(agents): share native harness tool execution (#154217)
* refactor(agents): share native harness tool result bookkeeping

* refactor(agents): share host tool invocation lifecycle

* fix(agents): clean shared invocation lint findings

* fix(agents): narrow media attachments with the shared record guard

* refactor(agents): keep tool helpers in the focused runtime seam

* test(agents): colocate shared harness tool coverage

* fix(agents): resolve shared tool refactor review findings

* style(agents): format shared tool runtime

* fix(codex): avoid shadowing tool start time

---------

Co-authored-by: Sarah Fortune <sarah.fortune@gmail.com>
2026-09-22 23:57:07 -07:00
Peter Steinberger
2e9a2eceb1
refactor(node-host): await prepared workspace storage (#149519)
Carry the prepared workspace worker cut onto current node launch, workspace,
and Codex lifecycle owners. Preserve synchronous external plugin acquisition,
transaction-time authority, and durable retirement semantics.

Join invocation and duplex cancellation through the runtime handoff and fence
legacy synchronous acquisition while async workspace mutation admission waits.

Private source checkpoint: fresh scoped review is clean through P2 and 43
selected ordinary/regression cases pass. Full changed checks and final build
remain incomplete; assertion baseline shrink maintenance and the current-main
schema 18 carry follow before qualification. The historical specific
reproduction remains held and was not replayed.
2026-09-22 23:44:38 -07:00
Vincent Koc
5b77505b37
refactor(policy): share sandbox allowlist findings (#155074)
Co-authored-by: Vincent Koc <vincentkoc@ieee.org>
2026-09-23 14:26:56 +08:00
Vincent Koc
6a9af74a1c
refactor(qa): share Slack native data scenarios (#155096)
Co-authored-by: Vincent Koc <vincentkoc@ieee.org>
2026-09-23 14:21:09 +08:00
Galin Iliev
e4f6503081
fix(codex): avoid duplicate replies after streamed completions (#156144)
* fix(codex): avoid duplicate replies after streamed completions

* fix(codex): retire preview activity before cleanup

Emit the superseded Activity event before removing an orphan final-answer preview. Cover terminal-snapshot-only completion and preserve the authoritative final transcript.

Co-authored-by: galiniliev <5711535+galiniliev@users.noreply.github.com>

---------

Co-authored-by: galiniliev <5711535+galiniliev@users.noreply.github.com>
2026-09-22 23:11:03 -07:00
Peter Steinberger
b981f930dc
fix: keep session writes responsive during archive cleanup (#156014)
* fix: keep session writes responsive during archive cleanup

* fix: load archive history reader lazily

* fix: resume budget cleanup after pooled worker checkpoints

* fix: keep archive maintenance fixtures on their host owner
2026-09-23 06:09:35 +00:00
Peter Steinberger
72ff919499
fix(codex): release idle catalog decoder memory (#156202) 2026-09-22 23:07:00 -07:00
Vincent Koc
e816496c4e
refactor(signal): share transport port reservations (#155166)
Co-authored-by: Vincent Koc <vincentkoc@ieee.org>
2026-09-23 13:56:28 +08:00
wangmiao0668000666
80bab7e23c
fix(mattermost): route threaded channel reactions to the thread session (#155812)
Fixes #156008
Related: #127605 (Mattermost slice)

## What Problem This Solves

A reaction on a Mattermost channel-thread message reaches the parent channel session instead of the thread containing that message.

## User Impact

Thread reactions now reach the same session as their messages. Top-level reactions remain in the parent session with default threading off. Configured threading follows existing message routing; failed post lookups retain the parent-session fallback. No configuration, public API, or stored-data changes are required.

## Why This Change Was Made

Reaction payloads contain a post ID but no thread root. Resolve the post through the existing bounded resource-cache pattern, then let the existing event plan choose the session using the same post identity as the message path.

## Evidence

- Exercised the real REST client, monitor resource cache, reaction handler, and thread resolver against a loopback Mattermost HTTP stub. The runtime event sink records the selected session.
- With pre-fix production sources, the thread assertion fails: the observed key is `mattermost:default:channel:chan-1`. With this head, it is `mattermost:default:channel:chan-1:thread:root-1` and the server observes the post lookup.
- Top-level/default-off and unresolved-post controls retain the parent session on both sides.
- Candidate regression and resource suites: 31 tests passed. The supplemental loopback proof adds three passing scenarios.
- Simplification retained the existing routing owner and bounded metadata-cache pattern; no competing policy or public seam. Production delta: +41/-2; retained test delta: +196/-0.
- Exact-head hosted CI run 35762478168 is green; hosted ClawSweeper is ready for `2f2942660abd633dcd4775d03b7dff7d4c547f43`.
- No live Mattermost tenant or WebSocket transport was exercised. Documentation already describes routed reaction events and existing thread settings; no documentation or release-owned changelog change was needed.

Co-authored-by: Ayaan Zaidi <hi@obviy.us>
2026-09-23 11:12:23 +05:30
Sarah Fortune
31e9a02836
refactor(agentsapi): use the official OpenAI SDK (#156116)
* refactor(agentsapi): use the official OpenAI SDK

* fix(agentsapi): narrow SDK message items and update dependency lock

* refactor(agentsapi): use SDK client defaults

* test(agentsapi): identify requests without URL matching

* refactor(agentsapi): import SDK types directly

* fix(agentsapi): honor resolved reasoning effort

* docs(agentsapi): explain configured reasoning effort

* fix(agentsapi): preserve the official SDK endpoint

* fix(agentsapi): honor transport abort deadlines

* fix(agentsapi): preserve admitted SDK authentication

---------

Co-authored-by: Sarah Fortune <sarah.fortune@gmail.com>
2026-09-22 22:40:36 -07:00
Alix-007
837a17c1fb
fix(openai): preserve GPT-Live startup after recoverable provider errors (#155534)
Related: #155190 (provider-startup portion only; telephone-call teardown remains separate).

## What Problem This Solves

Fixes GPT-Live sessions failing to start when a recoverable provider error arrives before `session.started`.

## User Impact

Recoverable warnings no longer prevent a subsequent readiness event from starting the voice session. Fatal authentication errors still reject startup immediately. No configuration or migration changes are needed.

## Why This Change Was Made

The bridge now checks the existing fatal-auth classification before rejecting startup, logs a redacted warning for recoverable startup errors, and continues waiting within the existing readiness timeout. Established-session error behavior and redaction remain unchanged. The classifier is unchanged: status 401 (number or string), or codes `authentication_error`, `invalid_api_key`, `invalid_token`, and `token_expired` are fatal; other error events, including `missing_scope` without status 401, are nonfatal.

Provider-error dispatch stays in a small private sibling module because the bridge is already at its enforced line cap. The regression checks readiness and terminal behavior rather than exact warning wording.

## Evidence

- Real bridge `connect()` replay using a real `ws` client and a TCP loopback WebSocket server, production event parsing/lifecycle, and the existing media-runtime test adapter. The server sent `error` followed by `session.started` after receiving `session.start`.
- Baseline `4d50b52bce`: `missing_scope` rejected startup; connected=false, ready callbacks=0. Candidate: startup resolved; connected=true, ready callbacks=1, warning=1, error callbacks=0, close callbacks=0.
- `authentication_error` rejected startup on both baseline and candidate, with connected=false and ready callbacks=0 despite the subsequent readiness frame.
- `node scripts/run-vitest.mjs extensions/openai/realtime-quicksilver-bridge.test.ts --maxWorkers=1`: 37 passed; measured command wall time 38.71 seconds (Vitest 34.73 seconds). Earlier exact contributor-head CI changed-extension test-shard step took 133 seconds.
- Replay ran under Bun 1.4.2; this is synthetic provider-stream transport proof, not live OpenAI availability, full Gateway, telephone-call, or media-worker packaging proof. No real provider credentials were used.

Co-authored-by: Ayaan Zaidi <hi@obviy.us>
2026-09-23 10:52:54 +05:30
Vincent Koc
6dda26e46c
fix(codex): report failed inference WebSocket handshakes (#155710)
* fix(codex): return HTTP errors for failed WebSocket upgrades

Return sanitized 502/504 responses for upstream connection and deadline failures instead of dropping the pending local upgrade. Keep native retries, HTTPS fallback, and provider auth errors unchanged.

* fix(codex): honor handshake deadline before forwarding failures

* test(codex): bound native inference fixture payloads

* Merge commit '5b7e61fe36' into fix/codex-websocket-handshake-20260922

Co-authored-by: Vincent Koc <vincentkoc@ieee.org>
2026-09-23 05:19:45 +00:00
Marvinthebored
b373c9a9bc
fix(talk): preserve Grok transcripts and confirmed voice consults (#155838)
Grok (xAI) Talk saved each spoken sentence as several duplicate or truncated user messages, because every cumulative transcription snapshot was persisted as a final turn. Longer spoken replies were also cancelled when they overflowed the browser's 10 s playback queue. Tool-call consults did not tell the agent which blocked call the user had confirmed.

The xAI provider now previews input snapshots and commits one final per utterance at a real boundary: next speech, response end, session close, or 1.5 s of quiet after a late recognition. It fences output, cancel errors and buffered tool calls from retired responses. The relay forwards snapshot mode and the saved transcript id, so the Control UI hides a live caption once its saved row arrives. Browser playback allows 60 s / 4,096 sources, and the 20 ms relay frame contract is unchanged. Tool-call consults receive the same blocked-call retry context and confirmation-id reply as native delegation. A same-action retry in a new run reuses the pending challenge without extending it. `onTranscript` gains an optional `{ textMode: "snapshot" }` metadata argument.

Proof: live isolated Gateway and Control UI Talk runs on xAI Grok against main.
- Three utterances were saved as 3 turns instead of 22. Long answers played to the barge-in instead of being cut by overflow.
- A spoken "Yes." wrote the confirmed file exactly once; "No." wrote nothing.
- Non-affirmations, expired challenges, superseded or consumed ids, and ids from a closed session were all rejected before the exec ran.

Co-authored-by: Marvinthebored <peter@lindsey.jp>
Co-authored-by: VACInc <3279061+VACInc@users.noreply.github.com>
Co-authored-by: Ayaan Zaidi <hi@obviy.us>
2026-09-23 10:44:33 +05:30
Sarah Fortune
40cc8c023a
fix(agentsapi): preserve conversation context in native messages (#156167)
* fix(agentsapi): include conversation context in native messages

* refactor(agentsapi): use private attempt context helper

* refactor(agents): keep inbound context type with shared owner

* fix(gateway): preserve context for active harness messages

---------

Co-authored-by: Sarah Fortune <sarah.fortune@gmail.com>
2026-09-22 22:11:32 -07:00
Josh Avant
094f165494
fix(workboard): run automations after workers finish (#156130)
* fix(workboard): run automations after worker authority closes

* test(workboard): update registration cleanup failure injection
2026-09-22 23:46:41 -05:00