Commit graph

83 commits

Author SHA1 Message Date
Dallin Romney
2cb02d3fcb
refactor(paths): centralize home directory resolution (#145866)
* refactor(paths): centralize home directory resolution

* fix: preserve home expansion compatibility
2026-09-13 18:58:17 -07:00
Peter Steinberger
b1b314e4bc
perf: reuse one-shot digests and batch file hash reads (#146581) 2026-09-12 17:59:28 -07:00
Peter Steinberger
d13f07b1c2
chore(deps): advance cooled dependencies and major upgrades (#146258)
* chore(deps): advance cooled dependencies and major upgrades

* test(logging): migrate failed-sink regression to tslog 5

* test: retain dependency upgrade coverage within lint limits

* fix(deps): preserve compiler launches, Matrix sync and chat metadata

Keep copied script harnesses independent of declaration modules and preserve
Windows executable prefixes after admission. Audit the Matrix sync guard for
42.3, align CI toolchain/cache pins, and refresh session facts after accepted
model-catalog invalidation without relying on picker timing.

* fix(ui): preserve scoped session reconciliation after catalog refresh
2026-09-12 13:28:12 -07:00
Peter Steinberger
1e6983ff43
perf(gateway): avoid unused projections and temporary copies (#146344) 2026-09-12 13:21:39 -07:00
RoboClaw
ccb230eae4
fix(mcp): valid results fail with nested union schema references (#145533)
Use the shared JSON Schema normalization owner for MCP structured-result validation instead of a competing traversal that moves nested resource definitions. Preserve original-schema diagnostics, draft-specific validation and annotation-only format behavior. Production +45/-99, net 54 lines removed.

With TypeBox 1.3.26, the parent independently exercised the real stdio/session runtime before and after the change: baseline rejected a valid leaf with a false-schema error while rejecting the invalid leaf; the candidate accepts the valid leaf and still rejects invalid output. Both production files were restored and source hashes verified.

The 15-case runtime selection, 97 sibling cases, complete changed checks, normal build and fresh P0–P2 review passed. The two parent stdio cases overlap the runtime selection. Exact-head CI passed and the current-head Claw review reported no blocking findings; the earlier head’s hosted failure remains unattributed. No HTTP/provider requests, dependency edits or guard overrides were used for proof.

Co-authored-by: steipete <58493+steipete@users.noreply.github.com>
2026-09-12 00:03:06 -07:00
Peter Steinberger
c7d356e2da
test: table UTF-16 slicing and truncation cases (#145570) 2026-09-11 21:04:51 -07:00
Peter Steinberger
6f1185305d
chore(deps): advance cooled TypeBox, AWS and Copilot SDKs (#145381)
* chore(deps): advance cooled TypeBox, AWS and Copilot SDKs

Update six direct package targets after the frozen 2026-09-04T22:10:05Z seven-day cutoff. All 15 new resolved artifacts match registry integrity and publication-age requirements.

Keep TypeBox pins and plugin examples aligned, and describe Copilot SDK native runtime packaging accurately. Regenerate the Workboard asset references with the updated graph.

Validation: independent P0-P2 review clean; live S3-compatible signed upload/readback/delete passed. Remaining full checks and Copilot live proof run in an isolated Testbox after local dependency-store loss.

* test(copilot): handle the SDK session detach handshake

Keep the held cleanup boundary and all cancellation and host-lifetime assertions while answering session.detach with success. The previous fake session.destroy handler left the real SDK test waiting until its deadline.

* fix(ui): clean up environment picker validation gates
2026-09-11 20:15:55 -07:00
Peter Steinberger
cd6fe8252e
chore(deps): refresh seven-day-cooled npm dependencies (#145124)
* chore(deps): refresh seven-day-cooled npm dependencies

Update 64 direct dependency targets and six pinned transitive targets using
registry publications on or before 2026-09-04T16:49:00Z. Verify publication
timestamps and registry integrity for all 164 newly resolved versions.

Deduplicate compatible resolutions so CodeMirror shares one state instance.
Retire expired cooldown exclusions and Mailparser's redundant security
scopes, retaining Mailauth's fixes at its new parent version. Synchronize
Claude ACP fallback assertions, plugin examples, and generated Workboard
asset metadata. Preserve patched packages and intentional migration holds.

Build, repository-wide formatting, 627 SDK/schema tests, 3613 plugin tests,
142 native/terminal/ACP tests, and independent P0-P2 review passed. Final
changed checks, browser editor proof, and exact-head CI are tracked in the PR.

Closes #145096

* test(tui): complete updated overlay handle fixtures

* fix(diffs): retain read-only types with updated renderer

* fix(validation): retain compact TypeBox rejection diagnostics

Normalize complete boolean child groups immediately before their matching
additional-property aggregate at the shared JSON-schema owner. Preserve
unrelated failures, literal-path collisions, typed/nested property errors,
truncated lists, and original error object identity. Both normalized-error
producers use the helper; validation decisions remain unchanged.

Use one oversized property in the Codex truncation regression so the test
continues exercising its original bound under TypeBox's new error layout.
Keep all rejection and non-execution assertions. Remove one unnecessary
validation-error cast and shrink its assertion ratchet accordingly.

The original Gateway assertions, 342 owner and sibling tests, native type
checks, targeted lint, and independent P0-P2 review pass. A real collision
regression failed the initial implementation and passes with ordered groups.
2026-09-11 14:29:57 -07:00
Peter Steinberger
084c5125f7
refactor: reduce duplicate queue work in error handling (#141530) 2026-09-07 13:30:04 -07:00
Peter Steinberger
13c7d0b0b6
docs(packages): remove repetitive module captions (#141500) 2026-09-07 13:01:23 -07:00
Peter Steinberger
5d4494faa6
refactor(crypto): share SHA-256 identifier helpers (#140887)
* refactor(crypto): share SHA-256 identifier helpers

Preserve raw digest slicing and logging label normalization in a Node-only normalization-core subpath. Keep browser exports and public logging contracts unchanged.

Proof: 453 old/new comparisons and browser-root compilation passed; independent review clean. Consumer and package gates continue in the PR.

* fix(build): resolve shared crypto source during type checks

NodeNext requires the exact TypeScript source mapping before package artifacts exist. A resolver probe fails without this mapping and resolves the canonical owner with it; independent review is clean.
2026-09-07 07:58:12 -07:00
Peter Steinberger
a7e0881c79
refactor(http): share bounded response byte consumption (#140654)
Preserve caller byte caps, decoder modes, abort errors, and cancellation lifetimes. Keep distinct operation and per-read timeout policies.

Proof: focused response and Dashboard suites (103 tests), changed checks, and 1232 differential reader cases passed.
2026-09-07 05:06:47 -07:00
Peter Steinberger
abc35e08c1
refactor(format): share compact token unit ladder (#140769)
Keep caller validation, flooring, precision, suffixes, and rollover policy unchanged across status, subagent, announce, and Dashboard displays.

Proof: focused formatter suites passed 164 tests; 96,660 original/candidate parity cases passed. Independent review found no actionable P0-P2 findings.
2026-09-06 21:55:21 -07:00
Peter Steinberger
b7de4987a7
refactor(core): reuse grouped phone metadata (#140514) 2026-09-06 16:38:08 -07:00
weiqinl
0d857862c3
fix: allow queued resets after restart recovery tombstones a chat (#137235)
Treat the stable restart-recovery tombstone rejection as terminal in shared ingress so an authorized queued reset can proceed. Retain the failed event for inspection and carry reset authorization through the authoritative lifecycle reread.

Real Telegram Test Server baseline and candidate evidence covers recovery, continuation, restarts, both reset commands, and authorization. Focused source checks protect nested errors, retries, lane ordering, durable claims, and reset admission.

Closes #137214

Co-authored-by: Ayaan Zaidi <hi@obviy.us>
2026-09-07 04:03:35 +05:30
Peter Steinberger
62fca0de77
fix(errors): preserve diagnostics and successful tool results (#140362)
* fix(errors): preserve cross-realm and async disposal failures

* test(restart): preserve cleanup across diagnostic failures

* fix(codex): preserve tool results when observer diagnostics fail
2026-09-06 12:35:01 -07:00
Eden
7683bc3cf4
perf: bound code-point prefix allocations (#136293)
Bound discarded-text allocations in prompt-cache keys, compaction instructions,
widget titles, typing previews, reply identifiers and Slack fallback chunks.
Reuse one pure code-point prefix helper across eight consumers while preserving
existing limits and output bytes. Keep newer streaming and one-pass owners intact.

Ordinary Unicode parity and isolated consumer measurements support the change.
The separate compaction seconds/RSS pair showed peak RSS increased by 244KiB;
no whole-process memory reduction is claimed. Preserve contributor ancestry
and carry the canonical CI repairs without expanding the performance scope.

Co-authored-by: xuyuanhao <aa9736195201@gmail.com>
Co-authored-by: 許元豪 <146086744+edenfunf@users.noreply.github.com>
Co-authored-by: Peter Steinberger <steipete@gmail.com>
2026-09-06 12:18:50 -07:00
ooiuuii
b791fce46b
fix(cli): ignore quoted JSON examples in mixed output (#122373)
* fix(cli): restore quote-aware mixed output parsing

Respect quoted CLI prose before selecting balanced JSON fragments while
retaining delimiter-first extraction for diagnostics and tool repair.
Consume fragments directly and remove the one-use candidate wrapper.

CLI backends retain real responses and session metadata around quoted
examples and brace-bearing banners. Ambiguous unmatched prose keeps its
visible raw-output fallback; JSONL quote state remains line-local.

Refs #122353.

Co-authored-by: luyifan <al3060388206@gmail.com>

* refactor(cli): reuse canonical JSON record parsing

Preserve whole-record precedence and quote-aware CLI fragment selection
while removing duplicate JSON parsing and non-array validation. Keep the
shared delimiter-first default used by diagnostics and tool repair.

Co-authored-by: luyifan <al3060388206@gmail.com>

---------

Co-authored-by: Peter Steinberger <steipete@gmail.com>
Co-authored-by: luyifan <al3060388206@gmail.com>
2026-09-06 11:39:38 -07:00
Peter Steinberger
512d8f2ee5
chore(deps): refresh seven-day-cooled dependencies (#138199)
* chore(deps): refresh seven-day-cooled dependencies

* fix: resolve dependency refresh CI blockers

Recheck caller cancellation after the OpenAI SSE iterator ends and before
Chat Completions can promote provisional tool calls. OpenAI 7.8 may end
an aborted iterator normally; preserve the shared transport's abort contract.

Wait for the fake WebSocket receive callback before emitting replies in
three Watch journal fixtures, retaining their existing timeout and assertions.
Point the QA release-policy catalog at the repository-owned plugin guide
after main removed its duplicate ClawHub publishing page.

The existing OpenAI regression fails before the owner fix and passes after,
with 243 owner/sibling tests and both transports exercised over real loopback
HTTP. The 58-test QA catalog suite, changed gates, AI package build, and
independent scoped P0 review pass. Fresh exact-head hosted CI, including
iOS lifecycle and production advisory checks, remains required before merge.

* test(mattermost): control loopback timeout deadlines

* test(ai): cover cancellation at normal stream completion

Prove the shared Chat Completions parser rejects an abort immediately before normal iterator return and never finalizes the provisional tool call. The regression fails with only the post-loop guard removed; 312 owner and sibling tests and the changed-file gate pass with the guard intact.

* fix(ui): preserve focused popovers during sidebar updates

Keep community invitation geometry deferred while a DOM-owned popover item has keyboard focus, even when Chromium reports no sidebar focus-within. Exercise background presence updates after real menu focus.

Distinguish independent Swarm child hydration from canonical roster refreshes in the held unread acknowledgement test. Preserve immediate badge clearing and prove an extra canonical refresh still fails the assertion.

Validation: 12 Control UI E2E cases and 431 related UI tests pass; the focused invitation case fails on the previous production condition. Independent Codex P0 review is scoped-clean.

* chore(deps): upgrade CUA and preserve published docs anchors

Upgrade CUA 0.22.0 to 0.22.2 with its coordinated accepted native
artifact records, and slugify 2.2.0 to 2.2.1 under the frozen seven-day
cutoff. Preserve Mint's published heading and component anchors before
counter allocation.

Synchronize the docs publisher's independent slugify manifest and npm
lock atomically with its parser; reject unrelated dependency drift after
rebasing the publish commit. Preserve every existing product security
exception, patch, toolchain document, and public configuration contract.

209 selected owner tests, source/publisher anchor corpus comparisons,
full changed checks and build, the exact dependency age/integrity audit,
and independent managed P0 review passed. Native CUA execution and
fresh exact-head CI remain required before landing PR #138199.

* chore(deps): refresh newly cooled September 5 dependencies

Advance the frozen seven-day selection to 2026-08-29T18:40:39Z. Update AWS, ACP, TUI, Discord types, Matrix WASM and duration formatting; align standalone broker Node types. Preserve main security exceptions and hold incompatible direct Zod upgrades.

* fix: align dependency refresh CI fixtures

Model the atomic publisher manifest/lock handoff in the process-fault fixture and execute its real validator before push. Preserve strict command matching, drain ordering and terminal rejection semantics. Inline the single-use reasoning-effort resolver to keep the shared stream below its existing line limit without changing cancellation or reasoning behavior.

* test(ui): await settled skill-menu geometry

Wait for the existing semantic and animation readiness boundary before comparing list and action widths. Preserve exact width tolerance, viewport bounds and read-only pin assertions; do not fast-forward animation or alter production styling.

* test(ui): bind live browser disclosures to their owners

Adapt the run/tool identity repair from 90c51add82 while preserving the existing 15-second overall observation budget. Remove live page-wide positional disclosure polling; keep history and inert-route assertions, and capture optional synthetic proof.
2026-09-05 22:59:41 -07:00
Peter Steinberger
ec40f39a8e
fix(sdk): report replaced runs as cancelled (#139125)
* fix(sdk): share canonical run terminal interpretation

* test(ui): wait for skill deletion actionability

* test(macos): retain health ownership coverage after settings move

Port the still-active HealthStore and ControlChannel cases out of the retired settings tests, retaining operation-boundary readiness and cleanup. Preserve the onboarding callback fix already on main and align the migrated timeout type imports with their existing module headers.

SDK integration proof: 370 focused Node tests, seven browser cases, native test compilation, full build and current Gateway/packed-consumer boundaries pass. Fresh managed review completed; duplicate unsupported wait-envelope findings were rejected after current and stable producer inspection.

* test(agents): drop retired transient-retry expectation

The assistant failure handler no longer receives transient retry control after recovery ownership moved. Keep the storage outcome, empty-error retry count, and credential-side-effect assertions without inspecting the removed dependency.

The exact failed CI snapshot reproduced TS2339 and an undefined-spy failure. All 23 existing tests, the same owning type graph, focused lint/format, and independent review pass after this one-line deletion.
2026-09-05 17:27:10 -07:00
Peter Steinberger
1f3b8063e4
perf(agents): share allocation-light weighted text counting (#138646) 2026-09-04 17:08:08 -07:00
Peter Steinberger
94e306a9d0
perf(normalization): sort the owned unique string array in place (#138305) 2026-09-04 08:47:44 -07:00
Peter Steinberger
270059bbf9
fix(agents): completed CLI actions replay when session saving fails (#137821)
* fix(agents): preserve results when finalization fails

Keep captured output, usage, and delivery evidence when native CLI session
continuity cannot be saved or remote computer cleanup fails. Record the
failure at its owner, classify replacement results independently, and prevent
another model attempt from repeating effects. Preserve typed frozen errors
through cleanup, rejecting credential cleanup, and the canonical terminal
outcome rather than reporting a false winning model.

This is the result-finalization prerequisite split from #136049 (paired Mac
cleanup and error preservation), related to #135944 (cleanup hides its cause).
It does not include native input lifecycle changes or complete that issue.

Validation: 565 selected tests passed; the final test-only cleanup was covered
by 77 rerun cases. Completed changed-check component plan, production build,
and fresh isolated Codex autoreview (P0 scoped-clean). Live proof follows on
this exact source before landing.

Production +104; tests +621; docs +2. The added production surface owns retained
results and frozen-error no-replay evidence; terminal projection is simplified.

* test(agents): pin cleanup stops through model fallback

Exercise the full outer fallback loop for typed overloads, HTTP 507 errors,
and wrapped cleanup failures. Ordinary provider failures may retry; errors
recorded by failed cleanup must retain identity and prohibit replay.

Extend the existing remote-exec table rather than add a new test seam.
The current-main owner negative control failed both cleanup variants while
non-cleanup controls passed; restoring the repaired owner passed all cases.

Integration for https://github.com/openclaw/openclaw/pull/137821.
Production patch rebased unchanged onto current main; this commit is tests
only (+32 net). Focused composition tests, changed checks, and build passed.
2026-09-03 22:56:57 -07:00
Peter Steinberger
f52713cd64
chore(deps): refresh seven-day eligible packages (#135177)
* chore(deps): refresh cooled packages and trusted Codex

Refresh 17 direct targets and owner-constrained transitive families using the
fixed 2026-08-25T02:09:07Z cutoff. Preserve the seven-day policy, trusted Codex
family and exact grammY exceptions. Pair native digests, runtime constants,
current-version documentation, UI boot manifest and Vercel lock fingerprint.
Apply only the approved TypeBox/Codex override bumps and remove the obsolete
Mailparser HTML-converter override now owned directly by Mailparser 3.9.16.

Consumer validation exposed an empty-reply outcome bug: final payload filtering
could report failure without notifying dispatch, hiding the diagnostic from
Gateway clients. Record failed outcomes at both existing payload failure
producers. Preserve deliberate silence, continuations and committed delivery.
Regression coverage checks directive-only output, real Gateway/TUI errors and
successful subsequent turns. Align the reset assertion with its existing
clear-context boundary; no reset behavior or schema changes.

Clarify release-only changelog edits in contributor guidance. No changelog,
OpenClaw release version, new configuration, or protocol version changes.

Proof: full builds, 993 dependency-owner tests, 111 reply/Gateway tests, 28
original-order PTY cases, full static/package checks and 29 repair checks;
exact-tag Codex protocol gate; real native SDK/Codex/CUA probes and inspected
synthetic Control UI before/after screenshots/video. The loaded-host PTY retry,
three preexisting exploratory library defects and partial advisory coverage
remain documented, not presented as a clean upstream security sweep.

* fix(codex): align catalog and fixture runtime versions

* test: expose Windows gateway cleanup failures

Preserve original cron assertion errors and report bounded taskkill/process/pipe diagnostics without changing shutdown policy or deadlines. Correct the Codex model/list cache-side-effect wording. The original Windows failure still requires diagnosis from native CI evidence.

* test: align native validation and late-filter proof

Keep Windows projects serial within each machine while retaining both matrix jobs and all assertions. Use valid Responses events and a late-filtered heartbeat marker to protect the terminal-failure callback after upstream streaming changes. Assert the exact error and successful next turn, not erasure of earlier PTY stream history.

* test(ci): align Windows guard with serial projects

* chore(deps): regenerate boot groups after integration

* test: remove empty package manifest suites

Delete eight plugin-only registrations that declare no assertions. The existing manifest helper registers only dependency-ownership and host-floor checks, so these rows fail Vitest collection while protecting no contract. Preserve all 501 manifest/dependency assertions; the original scoped command now passes. Independent Codex review found no actionable P0 issues.

* fix(talk): retain playback ownership until the player drains

Integrate the focused playback-owner repair from
59c2767beed12c101dc52225dc20d1e5692af3e6 in #136049 to unblock the native
CI failure in dependency refresh #135177. Remove estimated-duration
completion; only the generation-checked PCM player result completes normal
output. Explicit cancellation, clear, replacement and teardown retain their
existing ownership.

Keep the turn, playback marks and microphone echo suppression while queued
audio remains pending. The deterministic regression uses the existing
microphone timestamp seam; the original stale-player failure case is
unchanged. Document actual-drain behavior for Apple clients.

Also retain the canonical boot-generator refresh after the required conflict
rebase: main's gateway-suspend schema and download helper join the captured
shared boot group. No manual budget change or new dependency selection.

Local focused Swift proof passed 35 tests in four suites with synthetic
transport/capture/player boundaries. Full candidate isolated P0 review is
scoped-clean. Hosted toolchain parity and remaining landing gates remain
required; no merge-recovery or publication bypass.
2026-09-02 18:43:03 -07:00
Peter Steinberger
6e11486c48
perf(ui): avoid redundant work when loading chat history (#136396)
Reduce redundant history work by keying completed frames on their rendered outcome and nested content, reusing the prepared message-to-row index, and rejecting impossible JSON-object probes before parsing. Remove the duplicate yield-output parser and unused render-dependency slot.

Preserve the initial 800-message load, 1,000-message older pages, saved reader position, disclosures, and final-response actions. Sixteen source-pinned Gateway/Chromium trials measured 11%/15% lower initial/older maximum frame-gap medians with zero anchor drift; saved-return timings remain mixed.

Consolidate prepared CI comparison refs at authenticated checkout while preserving identity checks and credential lifetime. Shared UI fixture cleanup joins admitted Gateway work before shutdown and works in afterAll. The hovercard test controls its existing dismissal clock after preview loading and retains native pointer coverage.

Validation: 854 cases on the measured source; 23 Gateway/publication cases on unchanged affected paths; 15 complete-file Linux browser cases on the final tree; canonical static gate and managed P0 review. Final exact-head CI: 33700312894 attempt 2 passed 159 jobs with 10 skips, including the required aggregate. Three initially unassigned Mac jobs completed through the documented failed-job retry.
2026-09-02 18:12:27 -07:00
Peter Steinberger
0ebdbe6155
perf(normalization): reuse prepared string entries (#136418) 2026-09-02 10:18:44 -07:00
Peter Steinberger
d157f910a2
fix(dashboard): authenticate GitHub Actions data reads (#135740)
* fix(dashboard): authenticate GitHub Actions data reads

Add the repository-scoped github.actions.runs host binding using the
owning agent's existing GitHub identity, without exposing credentials to
sandboxed widgets or silently widening network grants.

Revalidate current widget, Gateway, and identity authority across reads;
bound and cache projected results. Teach widget authors the supported
bindings and report actual pin permission outcomes with recovery guidance.

Reuse shared identity preparation while preserving publication snapshots
and correctly classifying missing credentials as an identity failure.

Refs #135515

* refactor(validation): share ASCII control checks across owners

* fix(dashboard): validate pin identity and isolate shared readers

Verify the owning agent's usable GitHub identity before saving an
Actions-backed widget or requesting its capability approval. Preserve
existing content on failure and recheck caller, session, agent, and
credential authority before persistence. Keep author guidance conditional
without adding tool-construction authentication probes.

Let shared transport own its bounded credential-scoped cache result while
each widget independently authorizes delivery. A removed initiator no
longer fails a valid follower or evicts work before another reader settles.
Keep the existing approval policy in its own module without changing it.

Verified with 205 focused tests, complete changed-file checks, a full
build, isolated Codex review, and real Gateway/GitHub before-after proof:
anonymous quota 403, authenticated data and refresh, repository denial,
missing-identity rejection before mutation, and all three original CI cards.

Refs #135515 (authenticated dashboard reads), #135740 (implementation).
2026-09-01 20:42:55 -07:00
Peter Steinberger
3ca0d09d1b
fix(ai): diagnose pre-output provider transport failures (#135105)
Record transient failures at the shared provider error boundary so settled-tool
runs can use the existing bounded same-model continuation without replaying tools.
Route SDK catch paths through that owner and classify before partial-call cleanup.

Move the canonical network classifier below the core/AI boundary and preserve
native aggregate causes in safe error projection. Keep cancellation, mid-stream
failures, retry budgets, and operator configuration unchanged.
2026-09-01 03:46:18 -07:00
Peter Steinberger
1f71c763ea
chore(deps): refresh eligible seven-day npm dependencies (#133772)
* chore(deps): refresh eligible seven-day npm dependencies

* docs(plugins): align embedded TypeBox dependency pins

* test(deps): align evidence and Escape ownership

* fix(ci): repair native PID imports and cancellation assertions

* test(ui): make effort Escape ownership explicit

* fix(agents): keep error presentation on prepared policy

* fix(agents): preserve loaded provider policy in error presentation

* fix(agents): carry prepared provider owners into lifecycle errors

Preserve endpoint-owned recovery guidance for custom provider routes in terminal events and callbacks. Reuse the prepared model handle and full-signal classifier, with a real Agent/AgentSession boundary regression.

* fix(agents): reconcile explicit diagnostic ownership and structured errors

Keep presentation on explicit prepared owners, preserve full assistant error facts ahead of generic request wrappers, and retain raw-schema diagnostics. Carry prepared owners into terminal observations and prove source/compiled scope boundaries. Complete the shared attempt fixture with the real model-handle getter.

* fix(agents): carry full classified facts into safe failure copy

Share explicit-owner assistant classification between direct formatting and the user-facing wrapper. Preserve structured codes, types and body evidence in safe provider/model/status copy, including message-less failures, while retaining raw-schema diagnostics and ownerless policy boundaries.

* fix(ui): keep Home work context lazy and current

Let the existing deferred assistant panel prepare page work context once,
using the shell's validated route facts. Keep explicit agent ownership
through global/main aliases and refresh the quoted reference when session,
agent or Gateway snapshots change.

Reuse the frozen refinement from PR #134059:
6e2a8f9550e6e6da957b0352fa674024f198118e.
Add source-bound roster-refresh/send proof, extend the existing owner
fixture for snapshot updates and cleanup, and regenerate the boot manifest
for the pinned dependency graph. Startup gzip is 347243 B under the
unchanged 347353 B gate. The UI repair removes two production lines net.

* test(agents): align generation fixtures with prepared metadata

Use the captured main generation-scope contract in lifecycle and
source/compiled provider-owner fixtures. Remove its retired config input
while preserving provider selection and empty-generation fencing.

Integrate captured main 10564e2 with the Home context refinement from
PR #134059 and the assistant dock cleanup from PR #134435. Preserve the
existing contributor credit and canonical catalog owner already on main.

The integrated candidate passes the normal full build, scoped checks,
679 original-order model cases, 400 backend owner cases, both catalog
E2Es, 290 UI cases and 18 browser cases. Final grouped startup gzip is
347299 B under the unchanged 347353 B enforcement limit.

* refactor(ui): keep submission projection in lazy chat owner

Keep the app store responsible for bounded retained bytes and client lifetime.
Move receipt adaptation and display retirement into the existing history
projection owner, shared by both lazy chat consumers. Preserve missing-store
behavior and the retained-prompt, attachment and reconnect contracts.

Continue the retained-submission owner from PR #134059
(0f3e17e56b).

Validation: 808 owner tests, 21 Chromium cases, grouped-bundle Home and
retired-prompt proof, changed checks and fresh full-candidate autoreview.
Startup gzip: 347588 -> 347320 bytes; unchanged limit 347353.
2026-08-31 16:48:58 -07:00
Peter Steinberger
8d9282414c
test: consolidate CJK estimator coverage at its owner (#133701) 2026-08-30 19:41:26 -07:00
Peter Steinberger
6632c51526
perf(normalization): reuse schema branches and JSON traversal state (#133710)
* perf(normalization): reuse schema branches and JSON traversal state

* refactor(normalization): keep the JSON scan cursor local
2026-08-30 19:10:08 -07:00
Peter Steinberger
5d6918244e
perf(cli): reclaim startup-memory headroom on status and validation paths (#131428)
* perf(cli): defer status usage runtime imports

Keep credential resolution and provider usage dependencies behind the
optional status --usage boundary. Default status no longer loads the
model/auth graph merely to collect runtime diagnostics.

Preserve agent selection, credential routing, and usage merging while
removing redundant private loader wrappers. Extend the existing cold-import
suite with a reproduction that fails when default status imports auth.

Six local Node 24 samples reduce status JSON median peak RSS by 19.29 MiB
before the separate schema import cleanup, with a 1.77 MiB sample spread.

* perf(cli): trim schema validation startup imports

Use TypeBox's schema compiler and guard APIs for config and plugin manifest
validation instead of loading unused value codecs and transformations.
Preserve checks, format scoping, defaults, and structured validation errors.

Remove 145 TypeBox modules from the status and plugin inventory import
graphs. Extend the existing validation suite to reject eager value/compiler
imports, and shrink the assertion baseline after removing an unnecessary
error-array cast.

* test(cli): cover failed status usage imports

Extend the cold-import regression to prove requested usage and formatting preserve lazy-import failures while default status remains available before and after failure. Keep production behavior unchanged.
2026-08-27 19:48:12 -07:00
Peter Steinberger
395e5db41b
chore(deps): refresh dependencies after seven-day cooldown (#130296)
* chore(deps): refresh cooled npm and plugin dependencies

* chore(deps): refresh cooled build and workflow tooling

* chore(deps): retain formatter compatibility

* chore(deps): retain lint compatibility
2026-08-26 16:13:18 -07:00
Peter Steinberger
5c8dbdcabf
refactor(normalization): consolidate seconds conversion (#130069)
Amp-Thread-ID: https://ampcode.com/threads/T-01a03a19-be71-77a7-a886-4c688012c709

Co-authored-by: Amp <amp@ampcode.com>
2026-08-26 05:05:42 -07:00
Peter Steinberger
4dc7bb7411
chore(deps): refresh dependencies after seven-day cooldown (#129187)
* chore(deps): refresh dependencies after cooldown

* fix(gateway): emit append-only Responses content events

* chore(deps): retain unverified Sherpa runtime
2026-08-25 05:00:46 -07:00
Peter Steinberger
234df15a6d
chore: refresh dependencies after seven-day cooldown (#128414)
* build(deps): refresh dependencies after cooldown

Apply dependency, toolchain, action, image, and exact tool updates released by the inclusive 2026-08-16 seven-day cutoff. Adapt owner boundaries for the resulting CUA, logging, Teams, Markdown, native, and test-harness contract changes while retaining versions blocked by upstream compatibility constraints.

* fix(ui): align markdown renderer env typing

* fix(deps): align postcss and mistral peer contracts

* fix(deps): repair refreshed dependency contracts

* fix(deps): retain tslog startup budget

* fix(ci): verify Android tools with SHA-256

* fix(ci): fence Android SDK cache version
2026-08-24 03:01:54 -07:00
Jesse Merhi
0e8faacd71
fix(scripts): build heap ignores its systemd memory budget and takes the full default (#123979)
* fix(scripts): size the tsdown heap from the build's own cgroup budget

The build heap probe only read the cgroup root (/sys/fs/cgroup/memory.max and
the v1 equivalent). Those files exist only when the process runs in a
namespaced container cgroup; under systemd the budget lives on the process's
own slice, and the v2 root carries no limit at all. So every systemd-managed
build found no limit, fell back to /proc/meminfo MemTotal, and took the full
12288 MB default heap regardless of its actual budget.

Observed on a 15.4 GiB host: openclaw-main-update.service ran tsdown with
NODE_OPTIONS=--max-old-space-size=12288 while its user@999.service slice was
bounded at 5 GiB, reaching 3.2 GB RSS and 6.25 GB peak before the host began
OOM-killing unrelated services.

Resolve the limit from /proc/self/cgroup and walk that chain instead, reading
memory.high alongside memory.max (memory.high throttles reclaim rather than
failing allocation, so a heap above it stalls the build instead of OOM-ing),
and take the tightest bound found. Root paths stay as the container fallback,
and an explicitly injected path list still disables detection.

Co-authored-by: jesse-merhi <79823012+jesse-merhi@users.noreply.github.com>

* fix(scripts): resolve the build heap budget from the v1 memory controller too

The slice walk only accepted the unified 0:: record, so a legacy or hybrid
systemd host fell back to the root probe and kept taking host memory. One
resolver now walks both hierarchies leaf-to-root, which makes the static root
list its own depth-0 case and removes it.

Co-authored-by: jesse-merhi <79823012+jesse-merhi@users.noreply.github.com>

* fix(scripts): read cgroup controller mounts instead of assuming their paths

v1 controllers can be co-mounted at the cgroup root, where memory.limit_in_bytes
sits under the slice with no per-controller directory, so the hardcoded
/sys/fs/cgroup/memory probe missed the budget and the build took the full
12288MB default. Mount points now come from mountinfo.

Co-authored-by: jesse-merhi <79823012+jesse-merhi@users.noreply.github.com>

* fix(scripts): translate cgroup records through the mount root

mountinfo field 4 is the subtree a cgroupfs mount exposes. Under a container
mount the /proc/self/cgroup record stays host-absolute, so walking it verbatim
probed paths below the visible mount and the build fell back to host memory.
Records now translate through the mount root before the walk.

Co-authored-by: jesse-merhi <79823012+jesse-merhi@users.noreply.github.com>

* fix(scripts): skip cgroup mounts that cannot represent this process

Falling back to the mount root for a record outside the mount's subtree sized
the build from an unrelated cgroup: an inherited namespace clamped the heap to
the 2048MB floor from a foreign 1GiB limit. Non-representable mounts are now
skipped, and the blind root probe only runs when no memory record exists.

Co-authored-by: jesse-merhi <79823012+jesse-merhi@users.noreply.github.com>

* fix(scripts): keep every cgroup mount view, not just the last one seen

One hierarchy can be visible through several mounts and only some expose a
subtree containing this process. Retaining only the last view dropped the
budget whenever a non-representable bind view came later, sending the build
back to host MemTotal.

Co-authored-by: jesse-merhi <79823012+jesse-merhi@users.noreply.github.com>

* fix(scripts): decode octal-escaped mountinfo paths before matching cgroups

ClawSweeper P2 on 7e64ad61f70: the cgroup resolver compared mountinfo's mount
root and mount point verbatim. The kernel escapes space, tab, newline, and
backslash in those two fields, so any cgroup mounted under such a path never
matched, the bounded slice was missed, and heap sizing silently fell back to
host memory.

Decode both fields before matching. The decoder lives in scripts/lib beside the
other shared script helpers rather than inline, so the scripts program has one
copy rather than a new ad hoc one.

Regression test fails pre-fix: a v2 mount at "/sys/fs/cgroup\040dir" with a
5 GiB memory.high yields --max-old-space-size=12288 (host fallback) before the
fix and 4352 after.

Follow-up, deliberately not bundled here: src/infra/sqlite-wal.ts,
src/commands/doctor-state-integrity.ts, and src/plugins/bundled-source-overlays.ts
each carry their own private copy of this same decoder. Consolidating all four
into @openclaw/normalization-core is the right end state, but it touches a
shared package plus three core modules and belongs in its own reviewable change.

Co-authored-by: jesse-merhi <79823012+jesse-merhi@users.noreply.github.com>

* fix(scripts): resolve cgroup-namespace-relative records to their mount

ClawSweeper P1 on d6fe49dd3f4: inside a cgroup namespace /proc/self/cgroup
reports the namespace root ("0::/") while mountinfo field 4 stays the host
subtree the cgroupfs was mounted from ("/docker/<id>"). relativeCgroupPath then
found no prefix match and returned null; because a memory record had already
been seen, the root probe was skipped and the build fell back to host MemTotal.
A constrained container therefore missed its own budget entirely.

That namespace root is exactly what the mount exposes at its mount point, so it
resolves to "/" rather than failing closed.

Regression test fails pre-fix: a "0::/" record against a /docker/2f1a9c mount
root with a 5 GiB memory.max yields --max-old-space-size=12288 before the fix
and 4352 after.

Co-authored-by: jesse-merhi <79823012+jesse-merhi@users.noreply.github.com>

* fix(scripts): reject inherited cgroup mount views instead of guessing

ClawSweeper P1 on b4d200c5d20: the previous commit resolved a namespace-relative
record against any mount root, including the inherited views cgroup_namespaces(7)
documents, whose field-4 root reads "/..". Which cgroup such a view exposes is not
derivable from mountinfo, so probing it can size the build from an unrelated
cgroup's limit.

Reject non-canonical mount roots outright. An undecidable view now falls back to
host sizing, which is current main's behavior, rather than silently adopting the
wrong budget.

Regression test covers the "/.." inherited mount: it must yield host MemTotal
sizing, not the 5 GiB limit sitting behind that mount.

Co-authored-by: jesse-merhi <79823012+jesse-merhi@users.noreply.github.com>

* fix(scripts): fail closed on namespace-root records against non-root mounts

ClawSweeper P1 on 731d3bbc8e4: a "0::/" record does not prove that a mount
rooted at some other subtree exposes this process's cgroup. Resolving that pair
could cap the build heap from an unrelated cgroup's limit.

Return no mapping for it. An undecidable pair now falls back to host sizing,
which is current main's behavior, so the failure mode is a missed optimisation
rather than a wrong budget. The "/.." inherited-mount rejection stays; this
covers the broader ambiguous mapping it did not.

The namespace-relative test is repointed accordingly: an unrelated mounted
subtree must yield host sizing, not that subtree's limit.

Net production change: none (4 lines swapped).

Co-authored-by: jesse-merhi <79823012+jesse-merhi@users.noreply.github.com>

* fix(build): cap tsdown heap to the real budget and refuse hosts that cannot build

The 2048MB floor was applied on top of a discovered cgroup limit, so a small
container was handed a heap larger than it could honour. Measured in real
cgroups, that does not OOM-kill, it thrashes: a 1500MiB container sat pinned at
its ceiling for 10 minutes with oom_kill at 0, never finished the second of
eleven invocations, and starved every other process on the host.

Cap to the discovered budget, then refuse up front when that budget cannot hold
the build. The threshold is the whole-build peak, not a single pass: a full
eleven-invocation build peaks at 4730MiB, so a 5GiB slice completes while 4GiB
and 2816MiB slices are both killed partway through the third invocation.

The refusal runs before any output is cleaned, so a host that cannot rebuild
does not also lose the build it has.

* fix(build): harden tsdown heap admission

* fix(build): guard the default tsdown plan

* fix(build): preserve runtime-only Docker builds

* fix(build): admit only declaration cache misses

* fix(build): scope heap admission to real budgets

* fix(build): guard direct unified declarations

* fix(build): guard the canonical tsdown config

* fix(build): satisfy cache planning lint

* fix(gateway): release empty orphan leases

* fix(build): cap cgroup budget by host memory

* fix(build): serialize the canonical tsdown config

* test(build): freeze host memory fixtures

* fix(build): honor cgroup v1 soft limits

* fix(build): respect cgroup v1 hierarchy mode

* fix(build): admit unified runtime plans

* fix(build): admit every unified runtime path

* fix(build): collect repeated tsdown filters

* fix(build): ignore cgroup v1 soft limits

* fix(build): use explicit heap override as opt-in

* refactor(build): simplify memory admission

* fix(build): harden constrained build recovery

* fix(ci): prebuild runtime before real CLI shards

* fix(build): honor runtime-only runner environment

* fix(ci): satisfy tooling shard lint
2026-08-24 14:18:48 +10:00
ruel225
587c7524e5
fix(normalization-core): preserve non-Error object causes with extra keys in formatErrorMessage (#126654)
* fix(normalization-core): preserve non-Error object causes with extra keys

The cause-chain branch of formatErrorMessage called only
formatStatusAndCode(cause) with no stringifyUnknown fallback, while the
top-level branch used formatStatusAndCode(value) ?? stringifyUnknown(value).
formatStatusAndCode returns undefined for any object whose keys are not
exactly status/code, so a non-Error object cause carrying extra keys (e.g.
{ statusCode: 429 } or { status: 503, code: "UNAVAILABLE", requestId: "abc" })
was silently dropped — appendCauseMessage(undefined) no-op'd and the loop
broke, losing the diagnostic/retryable detail.

Mirror the top-level branch: appendCauseMessage(formatStatusAndCode(cause) ??
stringifyUnknown(cause)). Behavior-neutral for causes that already render;
restores the dropped detail for the asymmetric case. stringifyUnknown is a
local helper in the same file.

Closes #126652

Co-Authored-By: Claude <noreply@anthropic.com>

* test(normalization-core): assert structured cause metadata

---------

Co-authored-by: ruel225 <ruel225@users.noreply.github.com>
Co-authored-by: Claude <noreply@anthropic.com>
Co-authored-by: Altay <altay@hey.com>
2026-08-21 16:09:25 +03:00
Peter Steinberger
30337c8962
fix(qa-lab): include plural execution channels (#126140)
* fix(qa-lab): include plural execution channels

* chore(qa-lab): register browser error adapter

* fix(qa-lab): keep browser errors within boundaries

* fix(qa-lab): redact browser error credentials

* fix(sessions): drain sqlite writers during test cleanup

* fix(sessions): scope sqlite test handle cleanup

* test(codex): dedupe run attempt tools shard

* test(codex): converge run attempt tools shard
2026-08-19 01:17:22 -07:00
Peter Steinberger
5db8368f35
fix(errors): drop cause text the message already states (#125687)
* fix(errors): drop cause text the message already states

Cause-chain dedupe compared whole strings, so a wrapper that embeds its cause
verbatim printed it twice, and an errno detail was followed by its own bare
code. Skip any cause segment already contained in the accumulated message.

* test(backup): stop pinning the duplicated errno suffix

Both debug-view assertions required the bare code to follow the errno detail
that already names it. Assert the detail itself instead, so they pin the
message content rather than the redundancy.

* fix(errors): narrow cause dedupe to wrapper-embedded messages

Suppressing every contained segment also dropped trailing bare codes, which
cron, fs, and backup tests pin deliberately as this formatter's convention.
Restrict containment to cause messages so the wrapper duplication is fixed
without changing the code suffix, and restore the backup assertions.
2026-08-18 00:42:27 -07:00
Peter Steinberger
9a947e4bb8
fix(sessions): flatten Markdown in session list previews (#125113)
* fix(sessions): flatten Markdown in session list previews

Session-list previews were extracted verbatim from the last transcript
message, so raw Markdown leaked into every surface that renders the
subtitle as plain text — Control UI sidebar, TUI picker, native session
lists, and the sessions_list tool. A finished session read as
"Landed [PR #124879](https://github.com/...)".

Flatten lastMessagePreview at the Gateway producer, inside the
watermark-validated title-field cache, so the cost is amortized and no
consumer re-implements stripping. The flattener is the regex chain that
already existed privately in the Control UI narration line; it moves to
@openclaw/normalization-core/markdown-plain-text and both surfaces now
share one implementation. Session titles keep their own normalization
and are unchanged.

Also let an unread final observer digest outrank the raw last reply in
the sidebar subtitle, so the utility model's headline wins the slot it
was written for. The finalDigestUnread gate is untouched, so an
already-read digest still falls back to the flattened preview.

* fix(ci): register Markdown preview module

* fix(sessions): preserve literal preview punctuation

* fix(sessions): flatten previews before truncation

* fix(test): adapt session projection callback
2026-08-17 00:40:16 -07:00
Peter Steinberger
aeff737da9
fix(agents): prevent invalid names from targeting the default agent (#124670)
* fix(agents): reject unrepresentable agent ids

* refactor(system-agent): split model selection setup

* chore: shrink assertion safety baseline

* docs: record strict agent id validation proof

* style: format strict agent id report

* chore: drop stray unrelated report artifact

* chore: restore REPORT.md to main state
2026-08-16 10:37:06 -07:00
Peter Steinberger
4b8ca98a94
fix(gateway): omit internal class names from RPC failures (#124329)
* fix(gateway): drop internal error class names from operator output

* test(cli): drop stale Error: prefixes from capability expectations

* test(cli): align invalid-port output with canonical formatting

* fix(errors): preserve primary structured codes

* style(cli): format invalid-port expectation

* test(cli): align shared error rendering expectations

* fix(gateway): scope error code rendering to agent failures

* test(errors): redact opt-in structured codes

* refactor(errors): keep canonical formatter callback-safe

* test(audit): drop stale Error: prefixes after formatter cleanup

PR #124336 landed audit's gateway-error rendering while this branch was
in flight, so its two new assertions were written against the prefixed
output. The canonical formatter no longer emits the generic Error:
prefix, so the expected strings are updated to match.
2026-08-15 19:47:42 -07:00
Peter Steinberger
703c681aac
test: remove latest duplicate coverage (#123551) 2026-08-14 01:39:19 -07:00
Peter Steinberger
9da43d67e1
refactor: remove residual normalization adapters (#122771) 2026-08-12 11:51:45 -07:00
Peter Steinberger
c23d66e3b5
refactor: consolidate coercion ownership (#122692)
* refactor: consolidate coercion ownership

* test: align shard check with weighted planning

* chore: refresh plugin SDK API baseline
2026-08-12 09:25:28 -07:00
Peter Steinberger
b080dd1e76
refactor: consolidate coercion contracts (#122458)
* refactor: consolidate coercion contracts

Centralize exact string, record, numeric, date, Boolean, argument, and structured-error coercions while preserving call-site semantics.

Migrate canonical-name collisions and deprecated internal SDK bypasses, deleting 55 net production/tooling lines. Expand declaration ownership enforcement to 101 allowed helpers and add a narrow export-completeness audit.

* fix: preserve standalone script coercions

Keep copied Control UI tooling self-contained and retain the trusted release harness module-relative source seam when the harness runs against an old target cwd.
2026-08-11 23:26:37 -07:00
Peter Steinberger
964c8c84c1
refactor: consolidate coercion ownership (#122299)
* refactor: consolidate coercion ownership

Centralize four canonical coercion helpers, migrate exact core and plugin duplicates through narrow Plugin SDK facades, and enforce declaration and plugin-normalization ownership boundaries.

The sweep adds eight focused SDK exports while deleting more production and tooling code than it adds. User-visible behavior is unchanged except for safer equivalent object and UI parsing at existing boundaries.

* fix: guard integer option ownership

Register resolveIntegerOption with the canonical function owner and extend the declaration-guard fixture so future local duplicates fail validation.

* fix: keep integer helpers on numeric facade

Remove the unshipped duplicate string-coerce exports and route every affected plugin consumer through the existing number-runtime contract.

* fix: point numeric coercion to number runtime

Make boundary and declaration diagnostics recommend the canonical numeric facade, with failing-before coverage for both guidance paths.
2026-08-11 17:14:53 -07:00
Peter Steinberger
cad77fb39c
refactor: consolidate remaining coercion helpers (#122020) 2026-08-11 10:22:01 -07:00
Peter Steinberger
fa03d9b913
refactor: consolidate coercion helpers (#121366)
* refactor: consolidate coercion helpers

* fix: remove duplicate coercion imports

* fix: preserve serialized coercion guard

* chore: ratchet coercion helper carve-outs

* fix(test): keep gauntlet subprocess startup lean

* fix: preserve imported session timestamp semantics

* fix: preserve catalog timestamp string semantics

* chore: align plugin SDK surface ratchet

* fix: preserve trajectory and SDK string contracts

* fix(test): preserve QA record assertion semantics

* fix: complete standalone record guard rename

* refactor(cron): use canonical string coercion

* fix(acpx): preserve Pi timestamp parsing

* test(channels): adapt custody test harnesses

* test(telegram): classify media harness as test support

* test(acpx): split timestamp contract coverage

* test(channels): support generated custody contracts

* chore: ban the full coercion helper name set

Extends the declaration guard to all eleven consolidated helper names and
renames the cron schedule-identity readNumber wrapper to readScheduleInteger
so the banned generic name cannot regrow.

* fix(scripts): repair release-validation guard drift and lint cause

Restores the renamed isJsonRecord guard in assertTrustedWorkflowHarness after
main added isRecord call sites in parallel, and attaches the caught YAML error
as the thrown error cause (preserve-caught-error was red on main).

* fix: preserve Claude timestamp string semantics

* fix: preserve persisted timestamp string semantics

* fix: preserve date-first timestamp contracts

* fix(openai): harden delegation failure formatting

* chore: close coercion helper guard gaps

* test(openai): model non-error delegation rejection

* chore: refresh plugin SDK API contract

* fix(tasks): use canonical string field reader

* fix(ai): use canonical provider error field coercion

* fix(browser): migrate native bootstrap coercion

* docs(plugin-sdk): clarify text record export compatibility

* fix(gateway): normalize approval execution identity

* test(outbound): isolate message action poll harness
2026-08-11 00:02:18 -07:00