Commit graph

99056 commits

Author SHA1 Message Date
Vincent Koc
d9bf2a328e
improve(i18n): reuse source hashes in bulk catalog verification (#152657)
* improve(i18n): reduce locale catalog hashing overhead

* improve(i18n): reuse source hashes in bulk catalog verification

* Merge branch 'main' into improve/i18n-prepared-catalog-source

Co-authored-by: Vincent Koc <vincentkoc@ieee.org>
2026-09-23 07:40:00 -07:00
Vincent Koc
a87f01c46c
feat(models): publish v2 catalog beside unchanged v1 (#156530)
* feat(models): publish v2 catalog beside unchanged v1

* test(models): satisfy catalog publisher fixture lint
2026-09-23 14:37:23 +00:00
Peter Steinberger
6619a0e5eb
fix(memory): reject reads outside a replaced authorized directory (#156222)
Related: #154370

## What Problem This Solves

Fixes memory reads returning contents outside an authorized directory when that directory is replaced between path admission and the actual read.

## User Impact

Memory reads and transient retries stay bound to the admitted filesystem root. Existing contained workspace aliases, strict extra-directory rules, hardlinks, literal filenames, Markdown/glob restrictions, file-size behavior, and missing-file versus I/O-error results remain supported.

## Why This Change Was Made

The memory reader now retains the existing fs-safe Root through the content read and retry, replacing separate path checks followed by an unrelated absolute-path read. This removes duplicate containment helpers without changing configuration, stored formats, or public memory APIs. The existing `memory_get` and remote-worker routes continue to call the same reader.

## Evidence

- Two real directory-substitution cases returned outside contents on the original implementation and now refuse the substituted data. Existing compatibility controls also passed on the original source.
- Final focused proof passed 40 cases in 73.022 seconds. The changed reader suite passed all 15 cases with `node scripts/run-vitest.mjs packages/memory-host-sdk/src/host/read-file.test.ts --maxWorkers=1 --reporter=verbose` in 59.985 seconds wall time. Cold preparation differs between runs; no speed comparison is claimed.
- Coverage includes contained workspace aliases, rejected extra-root aliases, hardlinks, literal `~`, Markdown/glob admission, transient `EAGAIN` retries, propagated `EIO`, and missing/error distinctions.
- All selected check components passed across the initial runs and targeted lint corrections, including core and all 25 test type graphs, dead exports, formatting, and the remaining guards. The final ten-command recovery run passed in 583.501 seconds; the original failing check command is not represented as a pass.
- Earlier review findings about retry and I/O propagation were fixed. The final review's Markdown-widening claim was rejected after source inspection: `matchesDirectory` still requires `.md`, its directory branch requires that predicate, and the retained `note.txt` rejection case passes. Raw review findings were preserved; there are no accepted actionable findings. The final lint edit only omitted an identical default `void` type argument.
- Proof used synthetic local files on macOS. No Windows or live-Gateway execution is claimed.
2026-09-23 07:36:09 -07:00
Peter Steinberger
f2d7a0d210
fix(agents): republish model owners after a superseded auth refresh (#156561)
Reuse the auth publication queue and shared typed lifecycle predicate for one fresh-generation retry. Preserve pending gates and record a terminal retry failure so readers settle with a diagnostic reason.

Refs #156178. Thanks @Captain69G for the detailed report.
2026-09-23 14:31:34 +00:00
Vincent Koc
d8bf4703c6
feat(ui): explain connection access in Profile (#156304)
* feat(ui): show connection permissions in profile

* fix(ui): defer profile connection access copy

* improve(ui): explain Profile access in plain language

* test(ui): preserve personal editor through access reconnect

Co-authored-by: Vincent Koc <vincentkoc@ieee.org>
2026-09-23 22:31:06 +08:00
Peter Steinberger
ecf2fc7c17
feat(ci): add bounded retries for declared flaky release jobs (#156463)
* feat(frv): retry declared flaky jobs once per child

Freeze exact flaky-job declarations into the execution plan and admit one
automatic retry wave from child attempt one to two. Retain intent and
rejection witnesses across parent recovery without replaying requests.

Coordinate manual retries with automatic owners, preserve confirmed
no-effect outcomes through later operator attempts, and finish batch
provenance admission before issuing sibling mutations. Refuse unsupported
frozen tooling before creating refs or dispatching validation.

* fix(frv): bind immutable retry cache and workflow fixtures

Limit retry-intent cache saves to parent attempt one after the intent
witness succeeds. Qualify only that exact proof-cache path, key, workflow,
and job under cache isolation.

Refresh frozen dispatch fixture defaults and bind the expanded collector
dependencies and artifact downloads to their owning jobs.

* docs(frv): clarify when retry rejection evidence exists
2026-09-23 07:29:37 -07:00
Vincent Koc
318fc3c41f
fix: align Inbox notification ages and controls (#156560) 2026-09-23 14:26:40 +00:00
openclaw-mantis[bot]
3b77f9653c
chore(ui): refresh control ui locales (#156559)
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
2026-09-23 14:26:21 +00:00
Peter Steinberger
6e424042ad
fix(tests): await SDK retention maintenance (#155659) 2026-09-23 07:25:52 -07:00
Jason (Json)
a5fe21a5e7
fix(worktrees): recover interrupted clean removals (#156529)
* fix(worktrees): recover interrupted clean removals

* refactor(worktrees): name retained Git configuration reader precisely

* test(worktrees): settle recovery fixture state before cleanup
2026-09-23 08:22:35 -06:00
Vincent Koc
6b3eb02c54
feat(catalog): define the model-first v2 feed contract (#156518)
* feat(catalog): define the model-first v2 feed contract

* docs(catalog): explain v2 validation assertions

* fix(models): preserve v2 provider identities during sanitization
2026-09-23 14:19:06 +00:00
Vincent Koc
ab5b994413
improve(i18n): reduce locale catalog hashing overhead (#152355)
* improve(i18n): reduce locale catalog hashing overhead

* Merge branch 'main' into investigate/p13-sha256-helper-20260919

* Merge branch 'main' into investigate/p13-sha256-helper-20260919

Co-authored-by: Vincent Koc <vincentkoc@ieee.org>
2026-09-23 22:15:34 +08:00
Peter Steinberger
917029cc35
refactor: reuse CLI process exit fixtures (#156550) 2026-09-23 07:15:19 -07:00
Peter Steinberger
c755700867
refactor(sessions): share canonical row selection (#156495) 2026-09-23 07:13:47 -07:00
Peter Steinberger
9db35b7514
fix(agentsapi): package the bundled plugin icon, activity mark, and agent-runtimes category (#156549) 2026-09-23 14:08:59 +00:00
Peter Steinberger
9c85b63490
fix(workers): keep running turns on their original session store (#156452)
* fix(workers): keep running turns on their original session store

* fix(workers): stop new transcript submissions after cancellation

* test(workers): assert cancellation stops new transcript writes
2026-09-23 07:08:24 -07:00
Shakker
282fc78eaf
feat: support scoped session reading and organization (#155184)
Allow scoped users to read visible conversations and organize their own existing sessions while preserving separate execution permissions.
2026-09-23 14:56:25 +01:00
Peter Steinberger
482a4b2c49
perf(workers): accelerate warm cloud startup (#156380)
* perf(workers): reduce warm cloud startup work

* fix(workers): preserve snapshot recovery and refresh rules

* fix(workers): avoid lifecycle self-contention during admission

* fix(state): retain warm authority reads for Incognito lifetimes
2026-09-23 06:52:10 -07:00
Peter Steinberger
73788ab0d3
fix(ci): never fail a release image build on the build-limit warning relay (#156444) 2026-09-23 06:50:07 -07:00
Peter Steinberger
668eef7522
perf(secrets): reuse upstream TLS trust and connections (#156357)
Parse the immutable proxy CA bundle once and give each process grant an origin-pooled HTTPS agent. Revoke active and idle connections with their grant.

A 1,000-request local TLS rig reduced SecureContext creation and handshakes from 1,000 to 1, and main-thread CPU from 5.048 to 0.383 ms/request with identical forwarded payloads. All 99 focused proxy tests pass on Blacksmith Testbox.
2026-09-23 06:48:56 -07:00
Peter Steinberger
c8d3f81045
fix: isolate subagent concurrency per session (#156381)
* fix(agents): scope subagent concurrency to spawning sessions

Give each immediate spawning session its own configured execution budget,
including nested orchestrators, instead of making unrelated sessions share
one Gateway-wide queue. Preserve queue ordering, cancellation, hot limit
updates, and idle cleanup; retain the same parent identity during compaction.

Show aggregate activity with an explicit per-session limit in diagnostics.
Keep the existing child admission cap and Codex-native scheduling separate.

Verify with focused queue/runtime/UI tests and an isolated real-provider
Gateway case that holds one session's slot while another session and its
nested child run, then awaits all completion acknowledgments.

Closes #156370

* test(telegram): match sticker fixture to ingress dispatch

Register the native sticker pipeline's provider runtime prerequisite for
local and CI execution. Preserve the same Request JSON decoder while
exercising the registered bot handler, as durable webhook ingress does
after acknowledgment, and join the held turn before fixture cleanup.

The original CI file order reproduced the grammY timeout. Runtime
preparation alone was insufficient; the test's generic webhook adapter
imposed a deadline that is absent from the production dispatch boundary.
Keep the cache, media, and model-admission assertions unchanged, with the
real webhook acknowledgment contract covered by its existing test.

* test: repair packaging and runtime-owner CI fixtures

Copy the newly imported check-limits helper into the trusted packaging
harness so its dependency-free startup assertion reaches the CLI.

Compare prepared Telegram runtime files against the same config that
supplies the expected files. Keep worker-envelope coverage, packing,
concurrency and execution-budget assertions unchanged.

Both regressions reproduced before their fixes. The packaging owner passed
64 tests and ten targeted packing/prerequisite checks passed afterward.

* test(matrix): await follow-up adoption while active turn is held

Replace the 500 ms post-routing poll with resolver-owned adoption completion. Preserve foreground FIFO settlement by releasing the active turn before joining handlers, and settle fixture gates on timeout.

Validation: all four owner cases pass (37.31 s wrapper, 3.32 s bodies); selected changed checks and independent P0-P2 review pass.

* test(ci): recognize stopped Linux fixture process groups

Reuse the existing all-thread process-group assertion after detached fixture leaders are joined. Kernel kill probes also include exited zombies awaiting reaping; retain rejection of live descendants and uncertain observations.

Validation: 64 Docker scheduler cases and 18 retained-zombie/census cases passed on Linux Testbox; selected changed checks and independent P0-P2 review pass. The original CI process state did not reproduce in the original-order diagnostic.

* test(ui): count roster refreshes independently of child lookups

Install the browser clock before navigation and match roster sessions.list requests by includeGlobal. Preserve the 4999/+2 ms event window and avatar checks, and require exactly one roster refresh after the boundary.

Validation: reproduced the failing count and traced its extra request to a spawnedBy child lookup while the roster timer remained pending. All 14 browser cases pass (38.15 s; changed case 1.063 s). Selected changed gates and independent P0-P2 review pass.
2026-09-23 06:46:57 -07:00
Peter Steinberger
6f558f8ed7
perf(test): overlap independent publication admission fixtures (#156527) 2026-09-23 06:46:11 -07:00
Orion
930bd34747
fix(doctor): apply ambient-owner fallback to dream diary and all memory target handlers (#156437)
Apply the existing ambient-owner fallback in the shared Doctor memory target resolver, matching doctor.memory.status. Older clients that omit agentId can read the dream diary and run the five maintenance handlers when a configured owner is unambiguous. Explicit selection still wins, and ownerless multi-agent fleets still receive the selection error.

Thanks @Orionation for the fix and @Cyb3rb1ade for reporting the legacy-client failures. The configured-owner case is addressed; selecting an agent remains necessary for ownerless fleets.

Validation: scoped review found no accepted/actionable P0 or P1 findings, the author supplied a resolver comparison, and exact-head hosted CI passed. No fresh after-fix Gateway trace was captured for this landing. Production +6/-1; tests unchanged.

Closes #138369

Co-authored-by: Jonathan Sieling <jonathan@tailoredmonkey.com>
2026-09-23 06:45:18 -07:00
openclaw-mantis[bot]
6f90a9a332
chore(i18n): refresh native locales (#156534)
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
2026-09-23 13:44:38 +00:00
openclaw-mantis[bot]
e0da159a52
chore(ui): refresh control ui locales (#156404)
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
2026-09-23 13:44:12 +00:00
Peter Steinberger
31dd34ff19
fix(macos): report named-profile startup failures without a listener (#156441)
A named profile's readiness failure required port-ownership proof, which always fails without a listener PID, so the real startup error was replaced by a phantom port conflict. Check ownership only when something is listening; ready results keep strict ownership proof.

Co-authored-by: Peter Steinberger <steipete@gmail.com>
2026-09-23 06:41:18 -07:00
Peter Steinberger
c1ead4b71e
fix(setup): keep activation alive for the provider device-code window (#156440)
* fix(setup): keep activation alive for the provider device-code window

Setup activation of a detected route hosts the provider's device-code sign-in when no saved profile or key exists, but its wizard session expired after 8 minutes while the code is valid for 15. Share the provider-auth session budget so the sign-in can finish.

* test(setup): reject aborted activation waits with an Error

Co-authored-by: Peter Steinberger <steipete@gmail.com>
2026-09-23 06:34:57 -07:00
Peter Steinberger
d0de886394
test: advance admitted Telegram media retry waits (#156413) 2026-09-23 06:23:22 -07:00
Peter Steinberger
4205a47a63
fix(telegram): defer unused image discovery and settle test updates (#156482)
* test(telegram): settle updates before fixture cleanup

* perf(media): defer registry discovery for explicit image models
2026-09-23 13:21:14 +00:00
Jason (Json)
4c468412df
fix(update): inspect plugin migration data before Doctor rehearsal (#156326)
* feat(plugins): report undeclared migration recovery resources

Preserve read-only resource declarations through the Doctor contract and SDK plan adapter. Warn for legacy migrations without declarations, and reject invalid inventories without running migration work.

Co-authored-by: Peter Steinberger <steipete@gmail.com>

* test(plugins): qualify migration warnings through bundled public artifacts

* fix(update): inspect plugin resources in guarded Doctor rehearsal

Consume declared paths and typed legacy warnings from the existing migration inventory. Retain original parent and physical copied-root custody through admission and cleanup, and join readonly inventories before releasing copies.

Co-authored-by: Peter Steinberger <steipete@gmail.com>

* fix(update): preserve producer code links in guarded rehearsal

Bind validated plugin code links across snapshot workers, retain readonly Workshop history, preserve copied database bytes, and reject raw parent traversal before locality normalization. Keep declared migration data strict and recheck physical custody at admission.

Co-authored-by: Peter Steinberger <steipete@gmail.com>

* test(update): scope physical traversal witness to POSIX

Node normalizes Win32 parent traversal before realpath. Keep only that physical witness platform-specific while retaining every migration-refusal and external-byte assertion on all platforms.

---------

Co-authored-by: Peter Steinberger <steipete@gmail.com>
2026-09-23 07:19:49 -06:00
Peter Steinberger
7dba140902
fix: allow group defaults before the first Git commit (#156516) 2026-09-23 13:18:41 +00:00
Serg
4f43e47518
fix(xai): give new Grok releases thinking levels and image input (#156397)
Recognize plain, latest-alias, and dated Grok release IDs using xAI's documented capability ranges, so newer releases receive supported thinking levels and image input without another exact-ID update. Keep Grok 4.20 and variant suffixes on their existing conservative paths, and share the release rule with Responses tool defaults.

Validation: 56 focused regression tests passed on the refreshed head; independent review found no P0/P1 issues.

Co-authored-by: Takhoffman <781889+Takhoffman@users.noreply.github.com>
Co-authored-by: Tak Hoffman <781889+Takhoffman@users.noreply.github.com>
2026-09-23 13:18:28 +00:00
Peter Steinberger
63e4773938
test(workshop): verify reviews through deferred tools (#154915) 2026-09-23 06:17:04 -07:00
Peter Steinberger
5a15cf46eb
test: use native equality for Matrix archive bytes (#156490) 2026-09-23 06:12:01 -07:00
Peter Steinberger
0aa3c84d27
test: consolidate Doctor lint selection coverage (#156488) 2026-09-23 06:11:36 -07:00
Jason (Json)
d7b45d7e76
fix(worktrees): retire redundant removed snapshots explicitly (#156464)
* fix(worktrees): retire redundant removed snapshots explicitly

* fix(worktrees): preserve current recovery owners during explicit retirement

* test(session-share): run integration catalog service lifecycle
2026-09-23 13:08:27 +00:00
Shakker
d17bc2129b
fix: respect role model restrictions in model pickers (#154839)
Project permitted model catalogs and defaults through the shared role-policy owner. Retire stale choices on policy changes while preserving session history and saved model preferences.
2026-09-23 14:04:13 +01:00
siri1410
b6b45244ab
fix(secrets): preserve protected egress for non-ASCII attachment filenames (#156166)
Closes #154685

## What Problem This Solves

Fixes: downloading a forwarded attachment with a non-ASCII filename stops the secret egress proxy Worker and leaves protected egress unavailable until the Gateway restarts. Before #156354 moved forwarding off the main thread, the same defect crashed the Gateway itself.

## User Impact

User impact: attachments with non-ASCII filenames download with their original name preserved in `filename*`, and one bad upstream response head no longer stops protected egress. A rejected ordinary response gets a request-local 502 (or a closed connection for an already-bodyless response), while the proxy keeps serving.

## Why This Change Was Made

The secret egress proxy passes upstream response headers straight into `writeHead`. Node exposes received header bytes as latin1 strings and forwards them byte-for-byte, with one exception. Once `Content-Length` is stored, Node revalidates `Content-Disposition` as UTF-8, and a raw UTF-8 filename throws `ERR_INVALID_CHAR`. The throw happens in the upstream response callback, outside any `try`. On current main it terminates the proxy Worker: the Gateway remains live, but protected egress stops until restart.

This change:

- **Rewrites only `Content-Disposition`, and only when it contains non-ASCII bytes.** The legacy `filename=` becomes an ASCII fallback. An upstream `filename*` is kept verbatim; otherwise one is generated as `filename*=UTF-8''…` from the decoded name, so the original name survives. Other ASCII parameters (for example `size=`) are kept.
- **Leaves every other header untouched.** I checked this on Node 26.8.1: after `Content-Length`, the same raw UTF-8 bytes in `ETag`, `Link`, `Content-Type`, or a custom header forward fine, and only `Content-Disposition` throws. So opaque values such as `ETag` still work with `If-Match`/`If-None-Match`.
- **Contains any remaining `writeHead` failure to the one request.** The upstream response is destroyed and the client gets a 502. For 1xx/204/304 heads, where Node may already have marked the response bodyless, the socket is closed instead of framing a 502 body.
- Applies the same `Content-Disposition` handling to the forwarded `101` upgrade head.

There is no global `ServerResponse` patch, and request-side behavior is unchanged. #155121 took a similar approach and was closed by its author while awaiting review; this is an independent implementation that doesn't touch Node's private `_hasBody` flag.

Security note: this touches `src/secrets/egress-proxy/**`, so it needs `/allow-security-sensitive-change` from a maintainer.

## Evidence

**Real proxy and real upstream transport.** New test in `proxy-server.test.ts`: `forwards a CJK attachment filename from a real upstream and keeps serving`. It starts the actual `startSecretEgressProxyServer`, sends requests through an authenticated CONNECT tunnel with TLS to the proxy, and forwards them over real HTTPS to a real TLS upstream. The upstream writes raw bytes to its socket, because Node's own server cannot emit this header order. In sequence:

1. `Content-Length: 4` followed by `Content-Disposition: attachment; filename="附件_2026-09-21.log"` as raw UTF-8. The client receives `200`, body `file`, and `content-disposition: attachment; filename="___2026-09-21.log"; filename*=UTF-8''%E9%99%84%E4%BB%B6_2026-09-21.log`.
2. A head that Node rejects mid-write (`Trailer` on a non-chunked response). The client receives `502 Secret egress proxy could not forward the upstream response.`
3. A normal request on the same proxy returns `200 ok`, so the proxy is still serving. No `uncaughtException` is raised at any point.

On `main`'s `proxy-forward.ts`, step 1 gets no HTTP response at all and the test fails.

**Focused tests.**
- `proxy-forward.test.ts` loopback tests, with the upstream response emitted on a later tick as the real client does:
  - A CJK filename after `Content-Length`. Without the fix this fails with the production error, `TypeError: Invalid character in header content ["content-disposition"]`.
  - A real mid-`writeHead` rejection returns a `502` on a 200 head. On a 304 head the connection closes with no status line; with an unconditional 502 instead, the client would get a `502` head and then an aborted body, and the test catches that.
- `response-headers.test.ts` covers:
  - non-`Content-Disposition` headers, including a non-ASCII `ETag`, forwarded unchanged;
  - an upstream `filename*` and `size=` kept;
  - received UTF-8 bytes vs. decoded characters, and quoted vs. unquoted filenames;
  - non-UTF-8 Latin-1 input;
  - unparseable values.

```
✓ proxy-server.test.ts > secret egress proxy > forwards a CJK attachment filename from a real upstream and keeps serving 240ms
✓ proxy-forward.test.ts > secret egress forwarded response heads > forwards a CJK attachment filename that follows Content-Length 2ms
✓ proxy-forward.test.ts > secret egress forwarded response heads > answers 502 instead of crashing when the forwarded head is rejected 1ms
✓ proxy-forward.test.ts > secret egress forwarded response heads > answers 502 after Node rejects a forwarded head mid-write 1ms
✓ proxy-forward.test.ts > secret egress forwarded response heads > closes a rejected bodyless head instead of framing a 502 body 1ms
✓ response-headers.test.ts > toForwardableResponseHeaders > forwards non-ASCII bytes in other headers unchanged 0ms
✓ response-headers.test.ts > toForwardableResponseHeaders > keeps an upstream filename* and other ASCII parameters 1ms
… (6 more response-headers cases)
Test Files  6 passed (6)
Tests  115 passed (115)
Duration  6.70s
```

**Checks.**
- `pnpm check:changed` passed on the first revision. For this revision:
  - `pnpm tsgo:core` and the `agents-tools` core test shard typecheck (which includes `src/secrets/**` tests) are clean.
  - The gate's `scripts/run-oxlint.mjs --tsconfig config/tsconfig/oxlint.core.json` and `oxfmt --check` are clean.
- CI note: the two failing Node shards on the first revision were unrelated to this change. They were `src/gateway/session-lifecycle-state.persistence.test.ts` and `src/gateway/server.sessions.abort-authorization.test.ts` ("Matcher did not succeed in time"), and both pass locally on this branch.

## Maintainer verification

At head `2ec35def5c3fc9cb3ca4c2c9241491880b01ed49`, all three focused files passed (58 tests). Each file was run separately with `OPENCLAW_VITEST_MAX_WORKERS=1 node scripts/run-vitest.mjs src/secrets/egress-proxy/<file>`; measured command wall time, including startup:

| File | Wall time |
| --- | ---: |
| `proxy-server.test.ts` | 7.14s |
| `proxy-forward.test.ts` | 1.39s |
| `response-headers.test.ts` | 1.32s |

Negative control: replacing only `proxy-forward.ts` with merge-base `86d353b748` made the real authenticated CONNECT/TLS CJK-attachment test fail: no HTTP status was received (`NaN`, expected `200`). Restoring the candidate made the complete server suite pass, including the attachment, rejected-head 502, subsequent 200, and secret destination/authentication checks.

Available CI timing: initial `checks-node-compact-large-10` job ran 05:12:53–05:19:20 UTC (6m27s), with its failing owner-claim shard reporting 327.06s. The authorized failed-job rerun still failed that same unrelated owner-claim case. Current main was merged without rebasing the contributor commits; new CI is required.

### Built Gateway end-to-end proof

At candidate `1261c0851e317ca70a2286f398aa05e1d01e78fe` (main `0864fb6f86` merged, including #156354), `pnpm openclaw gateway run` rebuilt the CLI and started an isolated Gateway with `secrets.egressProxy.enabled=true`. A normal `pnpm openclaw agent` turn used the supported OpenClaw runtime and a synthetic local Chat Completions provider to issue an actual Gateway-hosted `exec` command. Curl used the process-owned authenticated proxy and TLS to download a raw-byte HTTPS attachment response with `Content-Length` preceding the UTF-8 `Content-Disposition`. Both sides used fresh isolated state.

- Candidate: the exec tool result contained `200 OK`, body `file`, ASCII `filename="___2026-09-21.log"` plus RFC 6266 `filename*=UTF-8''%E9%99%84%E4%BB%B6_2026-09-21.log`. A second request returned `200 OK` / `ok`, and the agent turn completed through the still-running Gateway.
- Current-main negative control: rebuilt the same checkout with only `proxy-forward.ts` replaced by its `0864fb6f86` version (the only changed production caller; the candidate helper was then unused). The identical agent flow received curl error 52 / empty reply for the attachment, then curl error 7 / connection refused for the next request. Gateway logs reported `secret egress proxy Worker stopped; restart the Gateway to restore protected egress`. `/healthz` still returned `{"ok":true,"status":"live"}`: #156354 contains the process crash, but does not fix attachment delivery or the protected-egress outage. Candidate source was restored afterward.
- The HTTPS fixture used localhost port 19443 because binding privileged port 443 was denied, and a certificate signed by the isolated Gateway's process CA. This proves the built Gateway/proxy/exec path, not Feishu delivery or general non-443 compatibility. No product callback or transport was mocked; only the model and upstream service responses were synthetic.

Co-authored-by: Ayaan Zaidi <hi@obviy.us>
2026-09-23 18:29:17 +05:30
Peter Steinberger
eeb816a418
chore(ui): retain Settings entrance failure evidence (#156500) 2026-09-23 12:57:46 +00:00
RoboClaw
bdd47f252a
feat: edit personal instructions on multi-user gateways (#155256)
Enable authenticated users to edit their personal instructions from Profile or chat on multi-user Gateways. Keep single-user Gateways on the workspace-root USER.md.

Preserve requester-bound authority, safe local writes, conflict checks, reconnect drafts, normal Profile spacing, and the startup performance budget.

Co-authored-by: VACInc <3279061+VACInc@users.noreply.github.com>
2026-09-23 05:55:12 -07:00
Peter Steinberger
d61bf615bb
perf(infra): keep hot Workers warm and expose churn
## What Problem This Solves

Intermittent tasks repeatedly recreate healthy Workers after the existing 30–60 second idle timeout, paying isolate startup costs and leaving resident memory behind after teardown. Existing metrics cannot distinguish which scripts are starting and exiting.

## User Impact

Hot task pools retain one Worker for five minutes of inactivity after an idle-retired Worker is promptly needed again. Node Code Mode retains its existing one-worker cache for five minutes. Operators can graph `openclaw_worker_started_total{script}` and `openclaw_worker_retired_total{script,reason}` through the existing diagnostics heartbeat.

## Why This Change Was Made

The existing retirement owner now owns idle timers and the single warm slot. Other slots keep their ordinary timeout; native cleanup, cancellation, pressure retirement, and rotation retain their existing ownership barriers. Existing idle GC collects released payloads in place. Code Mode still checks the runtime entry and heap limit before reuse.

The existing Worker resource registry owns cumulative lifecycle counts, including bundled direct Worker constructors. Counts advance on successful creation and confirmed native exit, survive exporter restart, and use bounded labels. Eval/unknown Workers remain `other`; nested Workers and V8-native threads remain outside the parent registry. Memory projection moves into a focused Prometheus module to keep the existing service below its growth ratchet.

No new configuration, database changes, dependencies, or managed-service environment changes. Updating installs the new runtime behavior through the normal restart; no migration or operator-state mutation is needed.

## Evidence

Real production-owner calls, ten verified synthetic tasks at 70-second simulated idle intervals (630 seconds observed), with real Worker creation and native teardown:

| Script | Starts before → after | Starts/min before → after | Task-run wall before → after |
| --- | ---: | ---: | ---: |
| `cron-stream-matcher.worker.js` | 10 → 2 | 0.952 → 0.190 | 485 → 111 ms |
| `git-operation.worker.js` | 10 → 2 | 0.952 → 0.190 | 3,076 → 824 ms |
| `code-mode-node.worker.js` | 10 → 1 | 0.952 → 0.095 | 858 → 187 ms |

The original source fails both new warm-window regressions; the pool creates a third Worker and Code Mode creates a second pool. In-place GC proof releases a 138 MiB payload to a 7.4 MiB heap while reusing the Worker.

### Allocator experiment and limits

Requested the largest advertised Blacksmith Linux class, 32 vCPUs, through a temporary proof-only workflow branch. The resulting guest exposes **8 online CPUs**, so this does **not** fulfill ≥32-vCPU validation or prove production memory behavior. Node 24.19.0, glibc 2.39, 32 concurrent Workers, 1,000 identical tasks touching 8 MiB each:

| Fresh process | Worker starts | Wall | RSS minus aggregate heapTotal increase after all Workers exit |
| --- | ---: | ---: | ---: |
| Create/terminate each task | 1,000 | 8.11 s | +639.59 MiB |
| Same churn, `MALLOC_ARENA_MAX=2` | 1,000 | 7.54 s | +423.86 MiB |
| Reuse 32 Workers | 32 | 3.08 s | +512.05 MiB |

The arena control reduces final residue 33.7%; reuse reduces it 19.9%. The churn proxy's post-100 slope is 102.63 KiB/task, versus 19.26 KiB/task with the arena control. A comparable retained-worker proxy slope is **not interpretable**: V8 heap capacity shrinks 992 MiB while RSS remains about 716.5 MiB, producing negative subtraction values and a false slope. Retained raw RSS grows only 0.152 MiB from tasks 400–1,000. These short exploratory results do not establish a many-core native-memory slope or allocator-contention tradeoff; consequently this PR does not set `MALLOC_ARENA_MAX`. Live deployment/production attribution remains a follow-up.

### Validation

Provider: `blacksmith-testbox`; profile: `openclaw-check`; lease: `tbx_01m36makaczyy9pds4yt5t2pem`; [workflow run](https://github.com/openclaw/openclaw/actions/runs/35834466664). One active remote command at a time, normal synchronization throughout. All test/typecheck/build work ran remotely.

- `node scripts/check-changed.mjs --base HEAD -- <all 25 changed paths>`: passed in 43m18s, including all 25 core-test typecheck graphs, core/extension production and test types, lint, documentation, SDK boundaries, dead exports, and import cycles.
- `pnpm test <file> --maxWorkers=1` for every changed test file, then `node scripts/run-vitest.mjs` for the five affected lifecycle/transport siblings: passed; 153 tests across 11 files. Final combined remote command: 59.2s. Original-source targeted runs failed both new regressions for the intended extra Worker/pool creation; candidate runs pass.
- The production-owner before/after benchmark command passed, including synthetic output assertions and native-exit cleanup for all three scripts. The allocator benchmark completed all three fresh-process cases with identical checksums.
- `pnpm build` followed by `OPENCLAW_LOCAL_CHECK=0 node --import tsx scripts/profile-extension-memory.mts --extension telegram --skip-combined --concurrency 1`: passed in 3m54s combined. All 95 plugin distributions built; bootstrap and 53 native control-plane module checks passed. Telegram isolated import exited cleanly, 160.09 MiB peak RSS (113.85 MiB above the profiler's empty-process baseline).
- Independent Codex review: scoped-clean, no actionable P0–P2 findings. `git diff --check`: passed. Net production growth: 146 lines; no test-only production seam.

Measured single-worker test command cost (seconds, including setup):

| Changed file | Wall |
| --- | ---: |
| `src/infra/worker-task-pool.test.ts` | 20.33 |
| `src/infra/worker-cpu.test.ts` | 7.48 |
| `src/logging/diagnostic-memory.test.ts` | 6.58 |
| `src/agents/code-mode-node.lifecycle.test.ts` | 5.21 |
| `extensions/diagnostics-prometheus/src/service.event-loop.test.ts` | 1.81 |
| `extensions/telegram/src/telegram-ingress-worker.test.ts` | 5.18 |

The two added behavior regressions use a fake idle clock; the real-worker pool case took 45ms and the Code Mode lifecycle case 35ms. CI timing will be updated after the head run exists. Crabbox reported an external runner-portal sync timeout after successful changed-check, final-test, and build commands; their underlying commands and reported run status succeeded.

Setup failures were diagnosed rather than counted as product failures: the transport checkout initially lacked `tsx` (fixed with remote frozen install), has no cgroup-v2 `cpu.max`, and lacks `origin/main` (changed checks use the actual original `HEAD` base plus exact changed paths). The initial benchmark fixture awaited an unrelated/unreferenced Worker exit; the corrected fixture uses the pool's resource-release boundary. The first changed gate identified line-cap growth; moving idle timing into its owner and extracting memory projection resolved it without an exception.
2026-09-23 12:53:07 +00:00
Peter Steinberger
bb8aea886b
test(ui): construct one Vite configuration per case (#155044) 2026-09-23 05:51:17 -07:00
Peter Steinberger
58cc1a65cc
fix(ui): show automation sources in task transcripts (#156487)
* fix(ui): show automation sources in task transcripts

* test(ui): load attribution stylesheet in session link fixture
2026-09-23 12:49:35 +00:00
Peter Steinberger
b1f980feb5
refactor(auto-reply): deslop auto-reply (#156377)
* refactor(auto-reply): deslop auto-reply

* refactor(auto-reply): decouple command metadata input type
2026-09-23 05:48:44 -07:00
Peter Steinberger
4663b6e404
fix(update): preserve original update failure diagnostics (#156189)
* fix(update): preserve whole npm failure lines

Failure records keep whole leading npm error lines and explicitly report oversized omissions within the existing diagnostic limits. Refs #156112.

* fix(update): retain the original baseline scan failure

Record baseline-scan-failed with the initiating error before identity fallback can replace it. Keep recovered timeouts advisory and preserve non-baseline classifications.

Use canonical schema inlining in the self-replay fixture; the unchanged base reproduced its missing SQL asset failure.

Release-note context: failure records identify the original package baseline scan cause.
Refs #156112. Thanks @shadesurgeon for the report.

* fix(update): mark truncated npm diagnostic lines

* docs(update): fix code span in truncation note
2026-09-23 12:43:31 +00:00
Peter Steinberger
6f942ef434
refactor(tests): reuse single-turn agent loop streams (#156414) 2026-09-23 05:43:05 -07:00
Vincent Koc
5706a06263
feat(ui): show latest commentary while agents work (#156340)
* feat(ui): show latest commentary while agents work

* test(ui): preserve commentary fixture tuple shape

* test(ui): retain historical commentary interaction coverage

* test(ui): preserve commentary recovery and scroll coverage

* fix(web): keep retained file drafts owned by active edits

* refactor(web): consolidate file line navigation guards
2026-09-23 20:39:43 +08:00
Peter Steinberger
ab16044564
test(gateway): await dispatch error settlement before late abort (#156460) 2026-09-23 12:39:13 +00:00
Vincent Koc
904f9c03a3
feat(diagnostics): trace completed harness commentary (#156343)
* feat(diagnostics): trace completed harness commentary

* fix(diagnostics): keep commentary event type internal

* fix(diagnostics): account for commentary in stability projection

* refactor(diagnostics): separate event fields and snapshot queries

* test(ui): provide scroll pane identity in ownership fixture
2026-09-23 20:33:15 +08:00
Vincent Koc
e77e2791e0
fix: show current activity for directly loaded sessions (#156462)
Co-authored-by: Vincent Koc <vincentkoc@ieee.org>
2026-09-23 12:31:57 +00:00