Commit graph

705 commits

Author SHA1 Message Date
Dallin Romney
869c182953
fix(maturity): land Windows, ChromeOS, and durable-work corrections (#142066)
* fix(maturity): include ChromeOS in Linux cohort (#142067)

* fix(maturity): recognize the Windows App and Node

* fix(maturity): cover durable work orchestration (#142068)
2026-09-08 15:39:21 -07:00
Dallin Romney
ba9b1e6749
fix(maturity): align current terminology contracts (#142065) 2026-09-08 13:29:18 -07:00
Teddy Tennant
996c928205
fix(auto-reply): show usage for invalid approval decisions (#137877)
## What Problem This Solves

`/approve abc constructor` and `/approve abc __proto__` returned a failed-submission error instead of usage. The parser read inherited object properties as decisions. It also misread `constructor` when used as an approval ID before a valid decision.

## Why This Change Was Made

Both supported argument layouts now check that a decision is a declared alias before reading its value. The ten aliases and downstream authorization and resolver behavior remain unchanged.

Downstream validation already rejected these values before making a Gateway request. This change gives the caller the normal usage reply at parsing.

## User Impact

Unknown decisions return the usage reply. Valid decisions still resolve pending exec and plugin approvals. An approval ID such as `constructor` also works when followed by a declared decision.

## Evidence

Verified candidate `b03792933d42978e969d9da25748f15fa0f79326` against pinned main `d3a2fb0296`, using separate installs and state on one isolated Linux Testbox.

- Real source Gateway and QA channel: baseline returns the invalid-decision submission error; candidate returns usage. The expanded matrix records 54 cases per pin, including every alias in both layouts, real pending exec/plugin records, duplicates, conflicts, unknown and expired IDs, whitespace, and prototype-like IDs.
- Real Telegram Test Server user and bot: the same before/after behavior, with one lease shared by both pins within each pair. Fast approval commands make zero provider requests. Owner records and resolution events confirm valid decisions take effect once and invalid inputs leave pending approvals unchanged.
- Separate controls cover the named account, unauthorized sender, disabled approval capability, disabled text/native commands, foreign-bot commands, and a real Gateway client missing approval scope. Disabled text handling follows normal deterministic model dispatch without resolving the pending approval.
- Fresh-pending syntax controls preserve whitespace and extra-token behavior. The earlier timed duplicate probe is retained separately because it crossed the existing resolved-entry grace window.
- All 71 focused approval/parser tests and 134 QA catalog tests passed. Changed-file formatting and lint passed. [Exact-head CI](https://github.com/openclaw/openclaw/actions/runs/33940306696) is green.

`toString` and `valueOf` were already rejected after lowercasing; they are preservation controls. Normal channel delivery may update message state. The approval invariant concerns decision and grant effects; audit-row absence is not treated as proof of absence. No execution-identity diagnostics or production approvals were enabled for this proof.

AI-assisted.

Co-authored-by: Ayaan Zaidi <hi@obviy.us>
2026-09-09 00:49:19 +05:30
Vincent Koc
f73b5c009e
docs(providers): split the OpenAI provider page by reader job (#141224)
* docs(providers): split the OpenAI provider page by reader job

docs/providers/openai.md was 85,081 characters and mixed explanation,
how-to, reference tables, and a verification log in one page. It is now a
short index over eight pages, one per reader job:

- /providers/openai/setup - API-key and Codex subscription paths, route
  summaries, OAuth recovery, and the long-context opt-in
- /providers/openai/models - quick choice, GPT-6 Astra, GPT-5.6 tiers
- /providers/openai/runtimes - naming map, implicit runtime policy, native
  Codex app-server auth
- /providers/openai/coverage-and-cost - capability matrix, memory
  embeddings, usage and cost tracking
- /providers/openai/image-and-video - image and video generation
- /providers/openai/voice-and-speech - TTS, transcription, realtime voice
- /providers/openai/azure - Azure OpenAI endpoints
- /providers/openai/advanced - prompt contribution, transport, Fast mode,
  compaction, strict-agentic mode, route compat

Anchor strategy: every anchor the single-page version published stays alive
on the index as an authored <a id="..." /> stub in a "Where each section
moved" list, so deep links such as /providers/openai#implicit-agent-runtime
keep resolving. Per-anchor redirect routes are not possible here because a
redirect source cannot contain a fragment. All 66 ids were enumerated with
parseDocsDocument from scripts/lib/docs-markdown.mjs, including the four
punctuated headings that publish both an encoded and a cleaned id
(async-tools%2C-steering%2C-and-reasoning-changes and
async-tools-steering-and-reasoning-changes; gpt-5.6-limited-preview and
gpt-5-6-limited-preview; route-summary-1/-2; config-example-1/-2). 65 ids
are stubbed on the index and published by exactly one child; the 66th,
"related", is still published by the index's own Related heading, so it is
not stubbed. parseDocsDocument reports zero duplicate authored/canonical
IDs on the index and on every child.

Losslessness: child bodies concatenate back to the original section bodies
byte for byte. Moved content before -> after: 9,538 -> 9,543 words,
83,074 -> 83,275 chars, 98 -> 98 code-fence markers, 34 -> 36 links,
121 -> 121 table rows (107 -> 107 data rows, identical as a multiset).
The deltas are exactly the four cross-references that the split orphaned
and that had to become real links ("described above", "accordion below",
and two same-page (#anchor) targets). Model refs (70), gpt/sora ids (98),
context and token numbers (30), prices (4), and pricing multipliers (7)
are byte-identical multisets before and after; no model, pricing, or
context-window figure was edited.

Also repoints the qa-lab realtime Talk scenario docsRefs from
docs/providers/openai.md to docs/providers/openai/voice-and-speech.md in
both qa/scenarios/media/realtime-talk-live.yaml and
test/e2e/qa-lab/media/realtime-talk-live.ts, so the reference follows the
content instead of the old path, and rewrites 15 in-repo deep links to
their new destination pages.

Closes audit findings: r3-0649, r3-0650

* docs(providers): restore two cross-page links reverted by the rebase

ClawSweeper found both. The split had already repaired these orphaned
directional references, but re-extracting child sections from main
during the rebase restored main's pre-split wording and reverted them.

- azure.md pointed at the Server-side compaction accordion "below";
  it now lives on advanced.md
- advanced.md said long-context multipliers stack "as described above";
  that explanation moved to setup.md

Unlike a broken link, an orphaned "above" or "below" passes every
validator, so nothing caught these mechanically.

docs-link-audit --anchors: 0 broken links.

* docs(tools): restore the Control UI Fast applicability paragraph

ClawSweeper caught this. My conflict resolution on docs/tools/thinking.md
took this branch's side wholesale to keep a repointed link, which also
deleted a paragraph main had added: the Control UI disables Fast choices
confirmed to have no effect, saved preferences stay visible and
clearable, and unknown applicability retains existing behavior. That
behaviour still exists in ui/src/lib/chat/model-select-state.ts and the
new provider pages do not carry the explanation.

Restored main's paragraph alongside the repointed link, which is what
the resolution should have done in the first place. Verified no other
line present on main is missing from this file.

docs-link-audit --anchors: 0 broken links.
2026-09-09 02:46:25 +08:00
openclaw-mantis[bot]
9c3e1c1382
docs: update maturity scorecard (#142114)
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
2026-09-08 11:06:56 -07:00
Dallin Romney
2bb684a68a
fix(maturity): make iMessage the active channel identity (#141963) 2026-09-08 01:49:35 -07:00
Dallin Romney
1d87919ec5
test(qa): align plugin lifecycle maturity proofs (#141964) 2026-09-08 01:11:48 -07:00
Dallin Romney
937aa4d1df
test(qa): bind Claude compaction fixture to active turn (#141978) 2026-09-08 01:11:33 -07:00
Dallin Romney
344c3a9c53
fix(maturity): align Linux companion scorecard with shipped app (#141955) 2026-09-08 01:10:42 -07:00
Peter Steinberger
941d9d752c
fix(telegram): preserve styled links inside rich message containers (#139976)
* fix(telegram): preserve authored HTML through rich composition

Record complete authored HTML lexemes at the Markdown parser before entity
or inline decoding. Carry those facts through IR projection and compose rich
wrappers and atoms on one text axis instead of reparsing individual leaves.
Keep unsupported HTML literal across Markdown block boundaries without losing
interior Markdown styles; preserve raw attributes in both linkification modes.

Hydrate partial TDLib rich messages before observer emission and correlate the
leased-DM scenario by run markers and observer identity, not cross-account IDs.
Extend the scenario to eleven rich cases and one same-message edit, with
bounded synthetic diagnostics and explicit full-tree assertions.

Related: #139976

* test(telegram): collect every rich formatting verdict before failing

Continue independent content checks after a formatting mismatch while keeping
all eleven exact native assertions mandatory for success. Keep the first
observed message as the edit target, retain original proof labels and separate
failed received trees from successful cases. Transport failures still abort.

Cover the observed custom-emoji downgrade, a first-message content failure,
and a transport error through the existing scenario harness. No credential
lookup, fixture replacement, production code, or entitlement change.

Related: #139976

* test(telegram): keep literal image fixture free of attachments

Use a valid empty Markdown image destination so the literal-alternative case
reaches the rich renderer instead of fetching a nonexistent example image.
Retain the exact expected native tree and all other cases. Exercise the real
outbound payload planner in the existing flow test to guard the no-media
contract before synthetic delivery.

Related: #139976

* test(telegram): align rich emoji fixture with documented alternative

Pair the existing custom emoji ID with the thumbs-up alternative documented
by the pinned Telegram types. Keep the native custom-node requirement and
update the existing degraded-plaintext negative control to match. This is
a fixture-alignment hypothesis; Test Server acceptance remains required.

Related: #139976

* test(telegram): use an observed Test Server custom emoji

Replace the unavailable documented emoji fixture with a default custom emoji
returned by the Test Server. Keep the strict native custom-node, URL/spoiler,
all-eleven-case and original-message edit assertions unchanged. Update the
synthetic plaintext-degradation controls to the same alternative.

The live lookup proves fixture availability, not sender entitlement or the
full native scenario. Production and net test LOC remain unchanged.

* fix(markdown): preserve raw lexemes during transcript annotation

Keep lexer-owned HTML attributes and comments outside transcript styling
while retaining their full text length for subsequent offset mapping.
This prevents opening-tag metadata loss and formatting inside raw comments.

Validated three failing preimage cases and129 passing owner/sibling checks,
plus fresh P0-P2 review. Production delta: -3 lines.

* test(telegram): model serve group-access precondition
2026-09-07 23:24:50 -07:00
openclaw-mantis[bot]
8e47d6998a
docs: update maturity scorecard (#141910)
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
2026-09-07 23:19:56 -07:00
Peter Steinberger
e32c358e15
refactor: share heartbeat and cron execution with reply runtime (#139624)
* refactor: share heartbeat and cron execution with reply runtime

Consolidate monitoring execution and delivery with ordinary replies, route isolated cron through the shared agent entry, and keep immediate cron retries in the canonical wake queue. Preserve quiet/alert policy, scratch ownership, cancellation, and durable delivery custody while removing duplicate runtime paths.

Carry actual admission outcomes through monitoring so foreground work winning late leaves cron and exec events pending for retry. Preserve CLI/native auth and lifecycle fences for saved context and route recovery through the canonical agent/store owner. Preserve separate physical stores when reconciling exhausted recovery attempts.

No new database schema, configuration, protocol, or migration requirement. Existing heartbeat public contracts remain supported. Production decreases by 56 lines in this slice and 68 across the two slices. Related: #135933, #138451.

* refactor: decouple heartbeat policy from reply run state

* fix: retain source session ownership for bound ACP admission

Keep lifecycle reset and admission in the source agent store while bound ACP execution retains its target. Update fixtures to canonical queue, store-owner, and no-send proof contracts, and register the added QA and regression coverage.
2026-09-07 14:58:09 -07:00
Peter Steinberger
49570ecc64
fix(qa): restore Gateway credential scenario selection (#141555) 2026-09-07 14:10:50 -07:00
Peter Steinberger
0d0e2852b2
fix(release): restore verification inputs and supported upgrade proof (#140981)
* fix(qa): verify leased Telegram tester group access

Check tester membership and effective text permission in the leased user driver before group readiness. Reuse the check in doctor and select trusted skill scripts for frozen release candidate QA. Preserve supported private DM turns and leave credential repair to the pool owner.

* ci: install Chromium for native live browser tests

The native-live-test shard includes the real Gateway widget restart proof,
which launches Playwright after restoring the interrupted session. Chromium
was installed only in separate Repo E2E jobs, leaving the live shard without
its required executable.

Install the candidate UI's pinned Chromium and system dependencies only for
the selected native-live-test row. Preserve profile selection and all test
gates. Verified the unchanged dc420 widget restart test with real OpenAI,
Gateway process replacement, and interactive Chromium dashboard assertions
on Blacksmith Testbox; canonical workflow checks and formatting also pass.

* test(release): prove supported cross-OS 9.2 transition

Select the explicitly supported external package-manager and fresh Doctor
transition only for 2026.9.2 to 2026.9.3. Preserve the old updater's safe
schema-15 refusal as negative evidence and never report self-update passed.
Verify backup, exact package identity, schema migration, retained settings
and session history, candidate serving version, and a fresh persisted turn.
Keep credential-bearing backup archives out of CI evidence and remove them
on both successful and failed runs. Other release pairs retain existing
updater and Windows fallback behavior.

Validation: 13 cross-OS upgrade lane tests, two pre-fix regression failures,
full changed checks, and independent review. Real packaged cross-OS/live
qualification remains required against the aggregate candidate.

* test(installer): prove supported historical package transitions

* test(release): exercise same-schema updater boundaries

Keep updater-specific plugin repair and consent checks on the prepared
candidate, then target explicit synthetic later versions with the same
runtime and storage schema. Retain installed-version and plugin-policy
assertions, and remove the obsolete legacy post-update fallback.

Add versioned future-tarball generation to the existing first-hop fixture
owner, with immutable input and source/target digest receipts. Restore the
reserved runtime-promotion staging ignore in package-derived Git fixtures
so the updater's unchanged-source check stays meaningful.

Validation: 26 focused tests, a pre-fix Git staging regression, full changed
checks, exact candidate tarball fixture generation, and P2 review. Actual
Docker update/consent/channel qualification remains required in aggregate CI.

* test(release): verify channel update staging cleanup

Assert the real channel flow refuses ordinary untracked user files without changing source, then require successful channel updates to remove every reserved runtime staging entry. Preserve recovery data when the assertion detects leftovers. No ignore or production cleanup policy is broadened.

* test(onboard): select configured model through explicit picker

* ci: pin release performance checks to reviewed Kova accounting

Select the reviewed Linux process-lifetime CPU accounting and bounded worker-watchdog implementation from Kova PR #110 for canonical, legacy-list, and trusted live fixtures. Keep existing trust classification, timeout values, and performance thresholds.

* fix(ci): distinguish bundle build modes in job names

* test(release): build matching future runtime fixtures

Derive a synthetic Codex runtime cohort from its verified candidate package by changing only package.version and openclaw.build.openclawVersion. Require matching source metadata and preserve payload, constraints, immutable input, and source/target digest receipts. Reuse the existing private tarball lifecycle and sequence validation. Validated 11 focused tests, pre-change CLI failures, full changed checks, and P2 review.

* docs(release): prepare private QA before local E2E

* fix(release): admit package metadata before validation dispatch

Port the reviewed early source admission into trusted Tooling, preserving the current selected REST transport. Reject invalid version notes and misaligned core packages before any Git push, non-GET API write, or workflow dispatch; retain explicitly allowed substantive draft notes.

* test(release): qualify supported historical upgrade transitions

Prove the published 2026.8.2/2026.9.2 to 2026.9.3 external package-manager and fresh Doctor contract through one shared helper. Preserve the old in-process updater refusal and schema 15, require owner-stopped migration with a verified private backup, and verify schema 16 plus retained Gateway history afterward.

Keep current-runtime update coverage distinct: managed restart uses separately identified future core and matching runtime fixtures, with real updater, canary, service replacement, authentication, and serving checks. Fix fixture-only channel environment pollution and mock serving nonce responses.

Sealed Docker journey, historical first-hop, current-candidate survivor, root-managed VPS, and onboarding pass. Managed-auth historical preservation passes; its future canary exposes an independent existing plugin projection blocker, retained as failing proof for the candidate repair owner. No product schema or refusal guard changes.

* test(release): narrow recorded update command before indexing

* fix(ci): align historical upgrade fixture ownership

Keep all seven external transition assertion cases with the first-hop package fixtures: both own the historical upgrade boundary and now share its temporary-state lifecycle. Preserve helper test selection and all assertions while avoiding an additional standalone tooling group.

Provide the new inference and future-package preparation boundaries in isolated service, mobile, and cron bootstrap probes, and verify each updater argument without changing readiness or migration assertions. Register the path-launched assertion CLI as a Knip entry.

Validation: 67 related tests; 70 exact-merge planner and combined-suite tests; five 80-job stress plans; full Knip unused-file scans; scoped checks and P2 review.

* fix(ci): model upgrade preparation in isolated probes

Complete the companion fixture and Knip repairs for the preceding transition-test consolidation. Model the independent inference and future-package preparation producers, verify exact updater argument boundaries, and retain service-readiness, authored-state, and cron migration assertions. Declare the path-launched transition assertion CLI in the existing Knip entrypoint list.

* test(release): preserve fixture registry on managed restart

Capture the fake service manager registry alongside its endpoint paths, then refresh that owned context after preparing the future cohort. Projected native service callers cannot drop or replace it; transient update state still stays outside the managed child.

The executable service regression fails on the original shim and passes with missing and conflicting caller registries. Four focused tests, the changed gate, and P2 review pass. Reuse the existing d928 package for the final managed-upgrade replay.

* test(release): preserve prepared UI identity in future fixtures

Advance synthetic package versions while retaining the opaque build ID already embedded in unchanged Control UI assets. The canonical asset-health regression fails the prior builder as stale, then verifies readiness and unchanged UI bytes for both future fixture sequences. Package-version update admission and managed-restart proof remain unchanged.

* test(release): retain manager-owned fixture policy

Capture the existing automation and offline-channel policy with registry identity when generating the manager shim. Native service clients project these values away; losing them activated retained synthetic channel credentials and correctly failed readiness. Restore only this explicit manager allowlist and keep transient update/compatibility state excluded.

The executable service regression fails on missing manager flags and passes across projected or conflicting caller environments. Four focused tests, the full changed gate, and fresh P2 review pass. Actual corrected managed Docker proof remains pending.

* docs(ci): keep bundle identity with routing guidance

* fix(release): retry transient GraphQL EOF failures

Classify the observed gh unexpected EOF transport failure within the existing five-attempt retry budget. Preserve authentication, invalid-response, inventory and snapshot handling. Exercise recovery, exhaustion and fail-fast errors through the real verifier CLI with a fake gh transport; also modernize the existing disposable-array sort required by touched-test lint.

* test(release): align survivor fixtures with manager preparation

Supply manager environment JSON to directly launched supervisors, and model the independent manager preparation phase in config-parking and companion fixtures. Preserve service readiness, restart, diagnostics, and strict phase-order assertions without changing production behavior or test limits.

* test(release): align upgrades with deferred schema publication

Restore installed-updater survivor and first-hop coverage after shared migration ownership moved into the durable update ledger. Record applied schema content separately from published schema version, retain the genuine missing-chunk negative control, and label external installation as an alternate with no self-update attempt. Preserve production-owned unsupported migration refusals.

* test(release): restore supported packaged upgrade coverage

Restore real 2026.9.2 to 2026.9.3 self-update after ledger-driven schema migration support. Keep retained conversations, private backup, settings, candidate identity, and serving inference proof. Require migrated content independently of publication grace and remove obsolete external-install diversion from cross-OS and installer update smoke. Shared external helper remains separately owned.

* test(release): refresh legacy manager before updater restart

Bind the service manager to the candidate registry before the published baseline updater starts its managed service. Preserve restart verification for successful and recoverable outcomes alongside explicit future targets. Exercise stale registry capture, replacement failures, pre-update identity, and attribution without changing timeouts or caps.
2026-09-07 11:47:27 -07:00
Peter Steinberger
cb1b29d9a1
fix: preserve account selection and recovery lifecycle ownership (#140973)
* fix: preserve runtime ownership across auth and recovery paths

Retain missing explicit model-account selections through chat and command
admission, preserve durable user context on internal retries, and fail over
exhausted usage windows without delaying temporary throttles. Keep retired
hook listeners alive for their locator grace period and await host-owned
Claude process cleanup before replacement.

Preserve supplied lazy plugin runtime services and Gateway protocol aliases.
Refuse unreadable shared state before Doctor update admission, migrate
credentials before model-route repair, and release staged QA auth-store
leases before child maintenance.

Align isolated Doctor, Claws, Discord, plugin, UI permission, compaction,
yield, and Windows worker-compilation fixtures with their existing contracts.
Keep deadlines, explicit concurrency coverage, authority checks, and
observable assertions intact. Account for the canonical injected provider
resolver owner introduced by #140667 in its exact static boundary inventory.

Release-note context: these repairs prevent silent account substitution,
lost retry context, stale relay connections, premature CLI process reuse,
and Doctor migration ordering failures. This forward-port carries the
reviewed release validation repairs independently to main without changing
release versions or publication state.

* test(doctor): assert manual recovery for unreadable state

Match the current manual recovery contract: require process-stop and verified-backup guidance, and reject a misleading repeat-Doctor loop. Preserve the database bytes, failure cause, selected path, and no-migration assertions. Both old command-copy expectations reproduced failures after the rebase; all six admission cases, scoped checks, and P2 review now pass.
2026-09-07 08:53:08 -07:00
Peter Steinberger
0f087c9625
chore(deps): refresh four seven-day-cooled packages (#140732)
* chore(deps): refresh four seven-day-cooled packages

Update nostr-tools, ignore, tsx and web-tree-sitter across seven manifest
bindings. Preserve the pinned Bash grammar and all existing dependency
policy, overrides, patches and compatibility holds.

The ignore update improves Git-compatible bracket/POSIX matching while
retaining fail-closed matcher composition. Nostr signing/relay entrypoints
remain unchanged; parser ABI, cancellation/reset and explicit deletion
retain their current contracts. Tooling cache policy is unchanged.

Frozen cutoff: 2026-08-31T02:32:01Z. Registry integrity and the complete
lock graph are verified. Focused loopback/owner and independent parser
proof plus managed review are recorded with final validation in the PR.

Closes #140691

* test(ui): keep cache probe independent of filesystem calls

The tsx update replaces synchronous directory indexing with bounded reads.
Require the positive control to observe cache access, rather than a
particular syscall. Keep all no-cache, sentinel, validator and exit-status
assertions unchanged.

The original suite reproduced the Linux/Windows CI assertion failure.
The repaired suite passed all 28 tests on an unchanged retry; the first
repaired run's timeout and cleanup errors remain recorded in PR evidence.
Scoped checks and independent review passed.
2026-09-06 21:45:52 -07:00
Peter Steinberger
512d8f2ee5
chore(deps): refresh seven-day-cooled dependencies (#138199)
* chore(deps): refresh seven-day-cooled dependencies

* fix: resolve dependency refresh CI blockers

Recheck caller cancellation after the OpenAI SSE iterator ends and before
Chat Completions can promote provisional tool calls. OpenAI 7.8 may end
an aborted iterator normally; preserve the shared transport's abort contract.

Wait for the fake WebSocket receive callback before emitting replies in
three Watch journal fixtures, retaining their existing timeout and assertions.
Point the QA release-policy catalog at the repository-owned plugin guide
after main removed its duplicate ClawHub publishing page.

The existing OpenAI regression fails before the owner fix and passes after,
with 243 owner/sibling tests and both transports exercised over real loopback
HTTP. The 58-test QA catalog suite, changed gates, AI package build, and
independent scoped P0 review pass. Fresh exact-head hosted CI, including
iOS lifecycle and production advisory checks, remains required before merge.

* test(mattermost): control loopback timeout deadlines

* test(ai): cover cancellation at normal stream completion

Prove the shared Chat Completions parser rejects an abort immediately before normal iterator return and never finalizes the provisional tool call. The regression fails with only the post-loop guard removed; 312 owner and sibling tests and the changed-file gate pass with the guard intact.

* fix(ui): preserve focused popovers during sidebar updates

Keep community invitation geometry deferred while a DOM-owned popover item has keyboard focus, even when Chromium reports no sidebar focus-within. Exercise background presence updates after real menu focus.

Distinguish independent Swarm child hydration from canonical roster refreshes in the held unread acknowledgement test. Preserve immediate badge clearing and prove an extra canonical refresh still fails the assertion.

Validation: 12 Control UI E2E cases and 431 related UI tests pass; the focused invitation case fails on the previous production condition. Independent Codex P0 review is scoped-clean.

* chore(deps): upgrade CUA and preserve published docs anchors

Upgrade CUA 0.22.0 to 0.22.2 with its coordinated accepted native
artifact records, and slugify 2.2.0 to 2.2.1 under the frozen seven-day
cutoff. Preserve Mint's published heading and component anchors before
counter allocation.

Synchronize the docs publisher's independent slugify manifest and npm
lock atomically with its parser; reject unrelated dependency drift after
rebasing the publish commit. Preserve every existing product security
exception, patch, toolchain document, and public configuration contract.

209 selected owner tests, source/publisher anchor corpus comparisons,
full changed checks and build, the exact dependency age/integrity audit,
and independent managed P0 review passed. Native CUA execution and
fresh exact-head CI remain required before landing PR #138199.

* chore(deps): refresh newly cooled September 5 dependencies

Advance the frozen seven-day selection to 2026-08-29T18:40:39Z. Update AWS, ACP, TUI, Discord types, Matrix WASM and duration formatting; align standalone broker Node types. Preserve main security exceptions and hold incompatible direct Zod upgrades.

* fix: align dependency refresh CI fixtures

Model the atomic publisher manifest/lock handoff in the process-fault fixture and execute its real validator before push. Preserve strict command matching, drain ordering and terminal rejection semantics. Inline the single-use reasoning-effort resolver to keep the shared stream below its existing line limit without changing cancellation or reasoning behavior.

* test(ui): await settled skill-menu geometry

Wait for the existing semantic and animation readiness boundary before comparing list and action widths. Preserve exact width tolerance, viewport bounds and read-only pin assertions; do not fast-forward animation or alter production styling.

* test(ui): bind live browser disclosures to their owners

Adapt the run/tool identity repair from 90c51add82 while preserving the existing 15-second overall observation budget. Remove live page-wide positional disclosure polling; keep history and inert-route assertions, and capture optional synthetic proof.
2026-09-05 22:59:41 -07:00
Peter Steinberger
95f963e342
fix(qa): give hot-reload maintenance proofs their own budgets (#139580)
* test(qa): isolate hot-reload maintenance proofs

Give browser tab cleanup and attachment retention their own scenario budgets while preserving real maintenance cadence, identity, and cleanup. Keep the remaining Gateway hot-reload umbrella at 15 minutes.

Wait for the environment stripe, use a startup-owned deferred-restart control, isolate SSH policy writes from unrelated rate-limit traffic, and declare the Avahi prerequisite.

Validation: real Gateway and Chromium tab cleanup, a full 60-minute retention sweep with fresh-file control, 75 passing umbrella checks, focused catalog and evidence tests, changed checks, and independent review.

* test(discord): bind WAV cancellation to the write boundary

Schedule leave through the existing pre-import workspace mock after real
WAV bytes are written and before the receive owner's write await returns.
Keep the original transcription, caption, and cleanup assertions, and
verify session retirement and file cleanup before unblocking queued work.

The full Discord shard passes on Linux. Removing the existing post-write
ownership check fails the immediate cleanup assertion; production remains
unchanged.
2026-09-05 19:37:45 -07:00
Vincent Koc
8e5a3f3638
chore(test): migrate to stable Vitest 5 (#138264)
* build(test): migrate to stable Vitest 5

Co-authored-by: Peter Steinberger <steipete@gmail.com>

* docs(test): document Vitest 5 contracts

* test(test): prove Vitest cache invalidation

* test(test): stabilize report config load proof

* test(test): restore Vitest 5 compatibility

Co-authored-by: Peter Steinberger <steipete@gmail.com>

* test(test): isolate pnpm cache fixture registry

* fix(test): keep pnpm 12 env lock portable

* fix(test): preserve explicit Vitest project roots

* fix(test): preserve nested Vitest project identity

* test(test): expect captured Vitest name prefix

---------

Co-authored-by: Peter Steinberger <steipete@gmail.com>
2026-09-06 04:17:59 +09:00
weiqinl
9d98269802
fix(compaction): heartbeat sessions ignore the active transcript byte cap (#136533)
* fix(compaction): enforce heartbeat transcript byte cap

Run the host transcript-byte preflight for heartbeat and Codex-owned sessions while retaining the unchanged-projection latch. Restricted Codex sessions compact locally before the native synchronization attempt.

Co-authored-by: zhangguiping-xydt <275915537+zhangguiping-xydt@users.noreply.github.com>

* fix(compaction): enforce heartbeat transcript byte cap

* fix(compaction): preserve byte preflight through wrappers

* fix(compaction): harden transcript byte preflight

* test(compaction): fix exact-head type coverage

* test(compaction): type legacy delegate mock

* test(gateway): cover Codex heartbeat compaction authority

* docs(codex): document transcript byte preflight

* test(gateway): tighten Codex heartbeat compaction proof

* fix(compaction): persist host accounting before native sync

* fix(compaction): preserve typed host commit metadata

* test(qa): tighten heartbeat proof typing

* test(qa): join heartbeat provider errors

* fix(compaction): refresh byte latch after maintenance

* test(compaction): satisfy successor entry contract

* fix(compaction): commit accounting with transcript boundary

* fix(sessions): type compaction writer fences

* test(qa): describe pre-checkpoint restart boundary

* test(compaction): type persistence validator

---------

Co-authored-by: LiuwqGit <7065327+LiuwqGit@users.noreply.github.com>
Co-authored-by: zhangguiping-xydt <275915537+zhangguiping-xydt@users.noreply.github.com>
Co-authored-by: Vincent Koc <vincentkoc@ieee.org>
2026-09-06 00:29:20 +09:00
Peter Steinberger
46aaa582b9
feat(gateway): hot reload diagnostics service settings (#138556)
* feat(gateway): hot reload diagnostics service settings

* fix(gateway): preserve reload policy precedence and scoped diagnostics

* test(config): remove duplicate retained-reader binding after rebase
2026-09-04 19:09:45 -07:00
Peter Steinberger
45ae22e03d
fix: full-access tasks lose tools after Gateway restart (#138701)
* fix: preserve full access when continuing interrupted tasks

* test: exercise full-access continuation across a live restart

* fix: retain replay safety beyond the recent transcript window

* test: make live restart checkpoints suspend explicitly
2026-09-04 17:53:02 -07:00
Peter Steinberger
805c6fb9cf
fix: preserve tasks following text directives (#138545)
* fix: preserve tasks following text directives

Keep native argument ownership strict while preserving complete text directive arguments and mixed tasks. Retain positional exec rejection and cover the full reply pipeline.

Closes #138530

* fix(qa): follow canonical directive argument helper

* fix: validate raw text exec argument boundaries
2026-09-04 17:16:03 -07:00
Peter Steinberger
4d6a68b9e6
test(qa): verify terminal fallback receipts and visible acknowledgments (#138598)
Match terminal delivery to its exact child/run direct-fallback receipt,
including after restart, instead of inferring the owner from text alone.
Count live acknowledgments without counting deleted preview history twice.
Keep history-wide silence and metadata checks, original deadlines, and the
canonical parent acknowledgment from #137342 unchanged.

Both four-case mock-channel flows, empty/restart behavior, 718 owner and
catalog cases, formatting and an independent review pass on the exact
source composition. No production or SQLite change.
2026-09-04 14:18:18 -07:00
Ayaan Zaidi
f804e08cba
fix(channels): keep quiet progress and the tool log under toolProgress (#137342)
* fix(channels): unify quiet progress and tool-log presentation

Use the existing toolProgress setting for quiet progress defaults and the
opt-in tool log. Retain authored text, plan milestones, and independent
approval/failure attention through full plans and activity bursts.

Move Teams onto the shared compositor and keep Matrix and Mattermost event
callbacks active in quiet mode. Give native Slack attention stable event
identities, settle recovery without repeating append-only output, and keep
attention separate from Block Kit activity limits.

Preserve shipped SDK resolver/summary/plain contracts and Slack's explicit
compact fallback. Make Telegram Doctor advice follow the streaming mode.
Retain current-main Slack Markdown and recovery fixes; align docs and hints.

Co-authored-by: Ayaan Zaidi <hi@obviy.us>

* test(slack): separate card and native progress coverage

Move the Block Kit cases into their own suite and share unchanged progress fixtures. Preserve all assertions while keeping both suites within the lint line limit.

* fix(channels): retain reasoning and terminal progress status

Route Teams reasoning deltas and snapshots through the shared compositor,
reset the reasoning burst at its lifecycle boundary, and retain failure
status when a native Slack card finishes with no work rows.

Align quiet-default help and generated metadata, and document Slack's
current compact final delivery after rebasing onto main. Config keys and
counts are unchanged. Preserve explicit tool-log opt-in and SDK contracts.

Co-authored-by: Ayaan Zaidi <hi@obviy.us>

* test(telegram): verify quiet and detailed progress through userbot QA

Exercise omitted and enabled toolProgress with the existing commentary,
exec, and final-response fixture through the Telegram Test Server adapter.
Check every observed message revision, preserve commentary through tool
updates, and correlate the two provider calls with the successful exec.

Accept coalesced drafts and partial commentary. The serve driver exposes
messages and edits, so this scenario makes no deletion-observation claim.

* test(telegram): check credentials before progress proof

* fix(channels): preserve the shipped SDK tool-progress default

Restore the public resolver's omitted default of true and pass the resolved
mode's default explicitly from bundled callers. This keeps existing plugin
callers compatible while bundled progress drafts remain quiet by default.

Add public-barrel regressions for omitted and forwarded undefined arguments,
replace the obsolete internal expectation, and clarify the SDK contract.

Co-authored-by: Ayaan Zaidi <hi@obviy.us>

* test(qa): acknowledge terminal-reply parent fixtures

Return a visible parent acknowledgement after spawning so native requester
settlement can complete. NO_REPLY deliberately retains an implicit
continuation, which prevented this direct-fallback scenario from reaching
its terminal delivery assertions through the real channel driver.

Keep worker silence, completion-agent fallback, and child-session-bound
release ordering unchanged. Five regression cases fail before the fixture
repair; all 284 mock-provider tests pass after it.

Format the Slack trace declaration with the branch's pinned formatter.

Co-authored-by: Ayaan Zaidi <hi@obviy.us>

* perf(ui): defer provider settings copy until its page loads

Reuse the lazy catalog repair from #136969 to clear the inherited Control UI
startup budget failure while landing #137342. Preserve every English key,
value and catalog ordering; keep shared navigation labels eager.

Verify cold Settings search and English fallback, then exercise New Session,
Chat and Models Settings through Chromium with a production bundle.

* test(ui): keep lazy settings coverage within suite limits

* test(ui): capture lazy scripts before browser delivery

* fix(ui): authenticate widgets before sandbox settings load

Keep managed scripted widgets on the authenticated document owner in strict mode, retire prompt authority on policy changes, and preserve inert rendering until scripts are allowed. Cover delayed bootstrap and sandbox lifecycle transitions.

Co-authored-by: Ayaan Zaidi <hi@obviy.us>

---------

Co-authored-by: Peter Steinberger <steipete@gmail.com>
Co-authored-by: Ayaan Zaidi <hi@obviy.us>
2026-09-04 11:19:59 -07:00
Vincent Koc
6f841c697d
fix(qa): preserve monotonic gateway log cursors (#137622)
* fix(qa): preserve monotonic gateway log cursors

* fix(qa): redact gateway logs across cursor boundaries

* fix(qa): keep invalid log cursors live

* fix(qa): discard split gateway log lines

* fix(qa): keep restart marker reads private
2026-09-04 21:15:00 +08:00
Peter Steinberger
085641a41a
fix: keep Code Mode tasks running after tool failures (#138044)
* fix: keep Code Mode tasks running after tool failures

Remove forced read-only reconciliation, recovery_resume, and the one-mutation continuation limit. Failed cells return to the normal agent loop so users can inspect partial effects and complete multi-step work with their configured tools. Preserve cancellation, permissions, explicit terminal outcomes, lifecycle receipts, and automatic whole-run replay protection.

* style: format Code Mode callback with pinned oxfmt

* test: update failed-tool QA references after recovery removal

* test(android): observe connection admission after stop

A startup WebSocket request can reach MockWebServer after Stop retires it,
causing the foreground-reentry assertion to report a reconnect that never
happened. Reuse the existing call-through session shadow to observe actual
connection admissions while preserving the real runtime and offline checks.

The same test bytes have a reproduced before failure, a rejected deliberate
reconnect, and 134 passing service/session sibling tests. Android lint passes;
the updated PR CI will verify this candidate before landing.

* test: repair ClawHub QA documentation reference
2026-09-04 04:08:04 -07:00
Peter Steinberger
fa70549fe4
fix(a2a): avoid duplicate source replies after message delivery (#137923)
* fix(a2a): avoid duplicate source replies after message delivery

Record confirmed final external source delivery in the run terminal receipt and consume it with the terminal reply. Remove history snapshots, fingerprints, and mirror inference from A2A; preserve internal forwarding and delivery facts through Pi, Codex, CLI, and nested tools.

* test(a2a): align delivery harnesses with settled receipts

* fix(agents): retain source delivery for replies and polls

Use the existing action-specific route predicates and settled plugin results for final source-delivery receipts. Add regressions for duplicate announcements, other targets, progress, partial sends, and dry runs. Recheck QA native target matching without scenario retries.

* test(agents): reuse direct harness fixtures
2026-09-04 00:50:04 -07:00
Peter Steinberger
e25c2543d1
fix(markdown): preserve inline code content across renderers (#137812)
* fix(markdown): preserve inline code content across renderers

Retain code-owned table edge whitespace through the shared IR slicer and
remove duplicate table clipping. Update all Markdown-it consumers to
15.0.1 for code-span, unclosed-label, and IPv6 URL normalization fixes.

Add shared-renderer regressions, bundled-browser proof, and a four-message
Telegram Gateway scenario. Application production delta: -26 lines.

Fixes #137728.

* fix(telegram): preserve native inline-code whitespace

Choose code-span padding from CommonMark-normalized space facts so all-space
entities stay exact and CR/LF edges retain their normalized content. Keep
UTF-16 entity offsets and the existing delimiter and fence owners unchanged.

Preserve the existing text/caption regressions and cover normalized line
breaks plus tab and non-breaking-space controls.
2026-09-03 21:39:26 -07:00
Peter Steinberger
3bbeba713e
fix: preserve Telegram accounts and reduce idle cleanup memory (#137860)
* fix: preserve Telegram accounts and reduce idle cleanup memory

Forward-port the remaining 2026.9.2 release repairs from f647f755da. Preserve arbitrary Telegram identifiers during Doctor migration, avoid idle reclamation workers while retaining archive publication, and wait for complete QA heap snapshots. Restore release fixtures to current runtime ownership and ordering contracts.

* test: repair native fixture boundary and browser route cleanup

Use the existing plugin path fixture to resolve the pinned native executable without importing the forbidden discovery helper. Release the browser bootstrap handoff and drain its page routes before context disposal so pending proxy responses remain available. Preserve all scenario assertions and timeout budgets.
2026-09-03 21:26:41 -07:00
Peter Steinberger
5a9168fea3
feat(gateway): hot-reload node, browser, and access settings (#137160)
Apply existing node command/tool/skill policies, browser defaults and launch
settings, detached-terminal retention, and Control UI origin/shared-credential
changes without restarting the Gateway.

Publish policy at its runtime owner, retain paired declaration ceilings across
WebSocket and Watch transports, cancel revoked work, and deliver accepted config
writer results before retiring stale connections. Preserve unaffected clients,
external browsers, and the original detached-terminal deadline anchor.

Remove obsolete health-monitor reload plumbing and duplicated policy, projection,
and startup-gating code. Auth mode, listener settings, terminal enablement, and
browser security globals remain restart-owned.

Verified the exact head in Crabbox with a full build, 36 real-runtime scenario
groups, and 52 authority tests. Includes signed Watch HTTP deny/reconnect/reallow,
actual invoke/poll/result recovery, and positive Gateway restart controls.

Related: #136832
2026-09-03 02:48:00 -07:00
Ayaan Zaidi
f6bda06bde
fix(qa): run Telegram RTT marker in direct messages (#137046)
Use an isolated direct message for the Telegram reply-chain RTT scenario. Record direct replies with the QA bus `dm:` target spelling so the live observer and scenario wait share one logical conversation.

Co-authored-by: Ayaan Zaidi <hi@obviy.us>
2026-09-03 11:47:42 +05:30
Peter Steinberger
2fffd4e8ba
feat(sessions): enable cross-agent session access by default (#136755)
* feat(sessions): enable cross-agent session access by default

`tools.sessions.visibility` now defaults to `all` and
`tools.agentToAgent.enabled` to `true`; both widen access.
Narrow access via `tools.sessions.visibility` (agent|tree|self),
`tools.agentToAgent.allow`, or `enabled: false`.

Document that an omitted/empty allow list permits every agent pair.
Denial copy for narrowed visibility no longer instructs enabling the
already-on policy. Regenerate prompt-snapshot fixtures for the visibility
hedge. Maintainer-directed.

* feat(security): audit default cross-agent session access

Add `security.trust_model.cross_agent_session_access_default`: `info`
for plain multi-agent defaults, `warn` with sandbox/tool-restriction/
multi-user ingress signals. No new config keys.

* test(gateway): drain detached a2a flow between agentId send rows

The announce/ping-pong flow outlives the sessions_send tool request; the second agentId row picked up the first row's follow-up agent call for agent:orion:main, so each row now waits for gateway active work to drain before releasing its test state.

* test(security): mock the cross-agent access collector in the non-deep facade

The readonly-setup-fallback test mocks audit.nondeep.runtime with an explicit factory; it now exports collectCrossAgentSessionAccessFindings so the registered collector resolves under the mock (CI run 33699015755, checks-node-compact-large-21).

* fix(security): scope the cross-agent audit to unsandboxed sessions

The audit finding now names which agents can reach other agents
(unsandboxed sessions, or any session when
agents.defaults.sandbox.sessionToolsVisibility is "all"), emits nothing
when every agent is fully sandboxed under the default clamp, and says
sandboxed transcripts stay readable by unsandboxed callers.

Docs qualify the agent-to-agent reference with the requester-owned
native/ACP child exception and correct the security overview's sandbox
wording. Addresses both ClawSweeper rank-up moves on #136755.

* docs(security): qualify the fully sandboxed audit exemption

State in the CLI reference and high-level security audit summary that
fully sandboxed rosters under the default spawn-tree clamp produce no
cross-agent access finding. Disabling that clamp removes the exemption.

Addresses the mechanical ClawSweeper rank-up on #136755 at 3208e19534e.

* fix(security): report per-agent session tool reach in the cross-agent audit

The finding now lists which agents can reach other agents (unclamped sessions that still have a session tool allowed) with their calling context and allowed tools, lists non-reaching agents with the reason, and emits nothing when nobody reaches; help text no longer claims enabled=false isolates agents because requester-owned native/ACP child sessions stay reachable under tree or all visibility; gateway final-effect proof that a disabled policy or restrictive allow list never dispatches to the target. Addresses the ClawSweeper re-review on #136755.

* test(gateway): drain detached a2a flow in an afterEach hook

The in-row drain shared the row's 10s budget and could time out under load, leaking the next row's mock calls; the hook has its own bounded timeout.

* docs: stop describing disabled agent-to-agent access as isolation

enabled: false blocks ordinary cross-agent access, but requester-owned native subagent and ACP child sessions stay reachable under tree or all visibility; every introduced claim now says so and points strict separation to tools.sessions.visibility or separate gateways. Addresses the ClawSweeper P2 on #136755.

* test(qa-lab): prove default cross-agent send and policy denials end to end

Three mock-openai flow scenarios run a two-agent QA Gateway: default config dispatches sessions_send to agent:orion:main (accepted, target run observed, target main session created); enabled=false and a restrictive allow list return forbidden before any target work. Addresses the ClawSweeper P1 merge risk on #136755.

* fix(config): stop listing tree visibility as strict separation

Strict separation is agent or self; tree still admits requester-owned native subagent and ACP child sessions across agents. Addresses a ClawSweeper rank-up move on #136755.
2026-09-02 22:39:03 -07:00
Peter Steinberger
8f7f96a545
feat(gateway): hot reload pairing policy, HSTS, and terminal shells (#136832)
* feat(gateway): hot reload pairing policy, HSTS, and terminal shells

Revalidate automatic pairing policy inside the approval owner after lock waits, scope SSH probes to their exact policy, and publish HTTP headers and new-terminal shell selection from committed config. Add focused regressions and a real Gateway, browser, PTY, and SSH proof scenario for all newly reloadable Gateway policy groups.

* fix(gateway): complete hot reload behavior and live proof

Honor the external-embed opt-in in the shared assistant-message renderer.
Exercise all 15 hot setting groups through a real Gateway in Testbox,
including held SSH probes, PTYs, browser pages and a restart control.
Align config-patch, catalog startup and upstream fixtures with their
owning contracts, and update the affected CI regression tests.

Validation: complete Testbox scenario passed; 69 focused Gateway tests
and 13 focused UI tests passed. No config keys or schemas added.
2026-09-02 22:12:29 -07:00
Peter Steinberger
0e7564ac11
fix: reuse prepared catalogs and repair Gateway E2E fixtures (#136915)
* test(agents): model confirmed source replies in announce fixtures

The successful announce fixture supplied only aggregate message-tool evidence,
which no longer proves a committed final reply to the source conversation.
Supply the receipt owned by the embedded runner without changing delivery
policy or requester-wake sibling rerouting.

Proof: the format suite failed 21 of 87 tests at both c92d0ae9c0 and
a41af7eb71; all 87 pass with the corrected fixture.

* fix(models): reuse prepared catalogs when browsing inventory

Inventory callers forced full discovery on every browse while trying to
refresh auth-stale content. Give the prepared catalog owner a distinct stale
refresh intent, preserve explicit forced refresh, and leave ordinary turn
reads on their existing facts. Apply the policy to commands, CLI, and RPC.

Correct the Ollama fixture response contract and await full catalog readiness.
Verify discovered-only models survive callbacks and RPC while discovery is
unavailable, and verify explicit refresh still reaches the provider.

Proof: before the fix, the real Gateway made 8 discovery calls after a 3-call
warmup, and the owner policy regressions failed. Prepared catalog and auth
reload owner tests now pass (43 tests).

* test(sandbox): verify writable private workspace isolation

The Docker scenario treated workspaceAccess=none as a read-only private
workspace. The sandbox mount owner intentionally provides private writable
storage while hiding the host and agent workspaces.

Assert a private write/read succeeds and the same-named host sentinel stays
unchanged. Keep read-only write rejection and read-write host persistence.
Use a unique temporary workspace in the owner test so ambient /tmp skills
cannot add unexpected mounts.

Proof: OrbStack Docker reproduced the old assertion on successful exit zero;
the complete none/read-only/read-write scenario now passes.

* test(acpx): use the plugin-owned adapter launcher

The cancellation fixture hardcoded a root node_modules adapter path, but
acpx owns that dependency. This caused deterministic MODULE_NOT_FOUND startup
failures in both release validation runs, not an agent startup flake.

Remove the override so the canonical plugin launcher resolves its adapter;
retain the synthetic Codex app-server process through CODEX_PATH.

Proof: the real installed adapter process passes cancellation authorization,
rejects stale replacements without interrupting them, and preserves queued
successors when the preceding run is cancelled.
2026-09-02 21:24:39 -07:00
Peter Steinberger
f52713cd64
chore(deps): refresh seven-day eligible packages (#135177)
* chore(deps): refresh cooled packages and trusted Codex

Refresh 17 direct targets and owner-constrained transitive families using the
fixed 2026-08-25T02:09:07Z cutoff. Preserve the seven-day policy, trusted Codex
family and exact grammY exceptions. Pair native digests, runtime constants,
current-version documentation, UI boot manifest and Vercel lock fingerprint.
Apply only the approved TypeBox/Codex override bumps and remove the obsolete
Mailparser HTML-converter override now owned directly by Mailparser 3.9.16.

Consumer validation exposed an empty-reply outcome bug: final payload filtering
could report failure without notifying dispatch, hiding the diagnostic from
Gateway clients. Record failed outcomes at both existing payload failure
producers. Preserve deliberate silence, continuations and committed delivery.
Regression coverage checks directive-only output, real Gateway/TUI errors and
successful subsequent turns. Align the reset assertion with its existing
clear-context boundary; no reset behavior or schema changes.

Clarify release-only changelog edits in contributor guidance. No changelog,
OpenClaw release version, new configuration, or protocol version changes.

Proof: full builds, 993 dependency-owner tests, 111 reply/Gateway tests, 28
original-order PTY cases, full static/package checks and 29 repair checks;
exact-tag Codex protocol gate; real native SDK/Codex/CUA probes and inspected
synthetic Control UI before/after screenshots/video. The loaded-host PTY retry,
three preexisting exploratory library defects and partial advisory coverage
remain documented, not presented as a clean upstream security sweep.

* fix(codex): align catalog and fixture runtime versions

* test: expose Windows gateway cleanup failures

Preserve original cron assertion errors and report bounded taskkill/process/pipe diagnostics without changing shutdown policy or deadlines. Correct the Codex model/list cache-side-effect wording. The original Windows failure still requires diagnosis from native CI evidence.

* test: align native validation and late-filter proof

Keep Windows projects serial within each machine while retaining both matrix jobs and all assertions. Use valid Responses events and a late-filtered heartbeat marker to protect the terminal-failure callback after upstream streaming changes. Assert the exact error and successful next turn, not erasure of earlier PTY stream history.

* test(ci): align Windows guard with serial projects

* chore(deps): regenerate boot groups after integration

* test: remove empty package manifest suites

Delete eight plugin-only registrations that declare no assertions. The existing manifest helper registers only dependency-ownership and host-floor checks, so these rows fail Vitest collection while protecting no contract. Preserve all 501 manifest/dependency assertions; the original scoped command now passes. Independent Codex review found no actionable P0 issues.

* fix(talk): retain playback ownership until the player drains

Integrate the focused playback-owner repair from
59c2767beed12c101dc52225dc20d1e5692af3e6 in #136049 to unblock the native
CI failure in dependency refresh #135177. Remove estimated-duration
completion; only the generation-checked PCM player result completes normal
output. Explicit cancellation, clear, replacement and teardown retain their
existing ownership.

Keep the turn, playback marks and microphone echo suppression while queued
audio remains pending. The deterministic regression uses the existing
microphone timestamp seam; the original stale-player failure case is
unchanged. Document actual-drain behavior for Apple clients.

Also retain the canonical boot-generator refresh after the required conflict
rebase: main's gateway-suspend schema and download helper join the captured
shared boot group. No manual budget change or new dependency selection.

Local focused Swift proof passed 35 tests in four suites with synthetic
transport/capture/player boundaries. Full candidate isolated P0 review is
scoped-clean. Hosted toolchain parity and remaining landing gates remain
required; no merge-recovery or publication bypass.
2026-09-02 18:43:03 -07:00
RoboClaw
b7fa0f52c6
fix(telegram): prioritize finals over CLI commentary (#134826)
Worked on by:
- @VACInc

Co-authored-by: roboclaw-bot <309084314+roboclaw-bot@users.noreply.github.com>
Co-authored-by: VACInc <3279061+VACInc@users.noreply.github.com>
2026-09-02 20:26:24 -04:00
Louisfy
13dd7d2e2d
fix(agents): record empty subagent completion (#135846)
## What Problem This Solves

Fixes #135633.

Successful subagent runs can finish with no assistant output. The old path classified that result as a missing visible reply, then entered generic fallback and retry.

## Why This Change Was Made

This change keeps blank successful subagent output distinct from explicit silence and actual missing output. The lifecycle owner records one intentional non-delivery before task finalization, cleanup preserves that fact, and the task remains successful across gateway restart.

## User Impact

Failed empty output still sends the generic requester notice. Message-tool output that is actually missing remains retryable. Yielded requester and requester-wake ownership stay unchanged.

Pre-change pending `empty` records also remain retryable because they do not contain the new explicit intentional non-delivery fact.

## Evidence

- Current-main regression reached `visible_reply_missing` and fallback/retry.
- 445 lifecycle and announce tests passed.
- 47 terminal producer tests passed.
- 101 formatting and lifecycle end-to-end tests passed.
- 311 QA provider and catalog tests passed.
- The mock-provider, ephemeral-gateway, QA-channel scenario passed with one visible parent acknowledgement, zero child terminal sends, `succeeded/not_applicable`, and no replay after gateway restart.
- The scenario asserts the exact parent acknowledgement once, and the sibling terminal-delivery matrix remains green without the retired empty-success expectation.
- Exact-head CI passed on run 33619750562.
- Exact-head P0 Autoreview reported no findings.
- Formatting, Oxlint, repository guards, and remote build passed.
- CI owns TypeScript type-checks under repository host policy.

Production code is +50/-17 against current main. The growth records the producer distinction, lifecycle-owned terminal fact, durable cleanup and task projection, while removing the downstream retry special case.

Co-authored-by: Ayaan Zaidi <hi@obviy.us>
2026-09-02 16:15:30 +05:30
Peter Steinberger
65db7dafd5
fix(media): parse Markdown images through the shared syntax owner (#136228)
Preserve code examples and escaped image syntax as text, and retain complete balanced image destinations. Share Markdown syntax parsing and carry native Telegram content types into the existing semantic proof.

Fixes #136139. Fixes #136140.
2026-09-02 03:18:48 -07:00
Peter Steinberger
51177301f2
fix(codex): preserve inbound audio for automatic voice replies (#135884)
* fix(codex): preserve inbound audio in dynamic tool context

Carry inbound-audio state from the shared attempt context so reconstructed
Codex message tools apply automatic TTS to voice input. Keep accepted audio
steering bound to its originating reply operation and remove the embedded
runner's duplicate forwarding path.

Add ordered real-tool regression coverage and a Gateway scenario that sends
audio followed by text through the native Codex app server.

Refs #135571. Thanks to @hyper-sdn for reporting the missing voice reply.

* test: run native voice QA scenario serially

* test(codex): cover inbound TTS after accepted audio steering
2026-09-01 22:33:04 -07:00
Peter Steinberger
a9469a2119
fix(markdown): preserve task lists across chunk boundaries (#135200)
* fix(markdown): preserve task lists across chunk boundaries

Compose canonical Markdown IR slices without losing list, block, annotation, or link metadata. Reuse one append owner for disjoint chunks and fallback edits, and retain semantic whitespace only when the rendered output is nonempty. Remove the duplicate partial metadata mergers.

* fix(markdown): keep composition proof in its owning test graph

Use checked metadata assignments and an unshadowed source-range name. Retain the plugin formatter and HTTP regressions while removing the duplicate plugin import from the core outbound planner test graph.

* test(markdown): preserve authored link boundaries during chunking

* test(markdown): avoid codifying dropped whitespace

* test(qa): verify native Telegram formatting boundaries
2026-09-01 15:14:11 -07:00
Josh Lehman
397f298cfd
refactor(plugins): centralize release cohort convergence (#134664)
Move plugin payload verification and post-core convergence to their owning boundaries, expose a reusable cohort convergence service, and prevent ambient service-repair policy from leaking into old-package Doctor runs. Behavior and exact-pin semantics remain unchanged.
2026-09-01 12:58:03 -07:00
Peter Steinberger
0ccc900ea3
fix(qa): require inherited child context evidence (#134708)
Reject forked-context passes built from parent prose or spawn planning alone. Correlate the accepted child with its raw inherited-history provider request, then verify source-bound completion and the canonical parent result across the existing mock-openai OpenClaw/Codex runtime pair.

Keep product history acquisition and other subagent QA work separate. Refs #134692.
2026-08-31 21:53:02 -07:00
Colin Johnson
5dd4c6e3b0
improve(ios): align composer controls with WebUI (#132683)
* feat(ios): align composer controls with web UI

* fix(ios): preserve composer capability ownership

* fix(ios): bind composer mutations to session owner

* fix(gateway): fence session settings mutations

* fix(ios): fence composer settings delivery

* refactor(gateway): isolate session patch expectations

* fix(ios): bind queued sends to session settings

* fix(ios): trust authoritative tool denials

* fix(gateway): freeze admitted session settings

* fix(ios): preserve durable composer authority

* fix(gateway): keep permission CAS write scoped

* fix(ios): fence mixed-version outbox delivery

* fix(ios): distinguish settings downgrade recovery

* fix(ios): persist switched-session settings failures

* fix(acp): enforce admitted session restrictions

* fix(ios): park queued sends before settings changes

* chore(review): document internal reply options cast

* test(gateway): keep session scope coverage focused

* chore(ios): refresh composer localization inventory

* refactor(ios): centralize current route metadata access

* fix(ios): reconcile composer camera control API

* test(gateway): prove session settings rollout

* fix(gateway): keep admitted settings private from hooks

* test(gateway): inspect restricted tool names

* fix(gateway): finish composer rebase cleanup

* refactor(ios): split shared chat state owners

* style(ios): format shared chat owner split

* fix(apple): fence queued sends by session settings

* fix(chat): preserve admitted session authority

* refactor(chat): isolate runtime takeover policy

* test(ios): prove inline model selection semantics

* fix(ios): expose inline model selection semantics

* fix(ios): label selected inline models

* test(ios): follow explicit inline model selection

* test(ios): reacquire model menu after selection

* test(ios): wait for inline model selection

* fix(ios): refresh inline model menu selection

* fix(ios): match qualified model selections

* fix(ios): restore native inline model picker

* test(ios): assert native model selection trait

* test(ios): verify inline model state transitions

* fix(chat): gate pre-dispatch session authority

* chore(ios): refresh composer inventory

* refactor(chat): remove dead composer seams

* refactor(gateway): isolate chat settings admission

* refactor(gateway): isolate active-leaf admission

* chore: format reply dispatch imports

* fix(ci): serialize OpenClawKit tests on macOS retries

* test(ios): bound outbox model patch fixtures

* test(gateway): unwrap settings rejection reason

* test(macos): connect before capability send
2026-08-31 20:21:10 -04:00
Peter Steinberger
1f71c763ea
chore(deps): refresh eligible seven-day npm dependencies (#133772)
* chore(deps): refresh eligible seven-day npm dependencies

* docs(plugins): align embedded TypeBox dependency pins

* test(deps): align evidence and Escape ownership

* fix(ci): repair native PID imports and cancellation assertions

* test(ui): make effort Escape ownership explicit

* fix(agents): keep error presentation on prepared policy

* fix(agents): preserve loaded provider policy in error presentation

* fix(agents): carry prepared provider owners into lifecycle errors

Preserve endpoint-owned recovery guidance for custom provider routes in terminal events and callbacks. Reuse the prepared model handle and full-signal classifier, with a real Agent/AgentSession boundary regression.

* fix(agents): reconcile explicit diagnostic ownership and structured errors

Keep presentation on explicit prepared owners, preserve full assistant error facts ahead of generic request wrappers, and retain raw-schema diagnostics. Carry prepared owners into terminal observations and prove source/compiled scope boundaries. Complete the shared attempt fixture with the real model-handle getter.

* fix(agents): carry full classified facts into safe failure copy

Share explicit-owner assistant classification between direct formatting and the user-facing wrapper. Preserve structured codes, types and body evidence in safe provider/model/status copy, including message-less failures, while retaining raw-schema diagnostics and ownerless policy boundaries.

* fix(ui): keep Home work context lazy and current

Let the existing deferred assistant panel prepare page work context once,
using the shell's validated route facts. Keep explicit agent ownership
through global/main aliases and refresh the quoted reference when session,
agent or Gateway snapshots change.

Reuse the frozen refinement from PR #134059:
6e2a8f9550e6e6da957b0352fa674024f198118e.
Add source-bound roster-refresh/send proof, extend the existing owner
fixture for snapshot updates and cleanup, and regenerate the boot manifest
for the pinned dependency graph. Startup gzip is 347243 B under the
unchanged 347353 B gate. The UI repair removes two production lines net.

* test(agents): align generation fixtures with prepared metadata

Use the captured main generation-scope contract in lifecycle and
source/compiled provider-owner fixtures. Remove its retired config input
while preserving provider selection and empty-generation fencing.

Integrate captured main 10564e2 with the Home context refinement from
PR #134059 and the assistant dock cleanup from PR #134435. Preserve the
existing contributor credit and canonical catalog owner already on main.

The integrated candidate passes the normal full build, scoped checks,
679 original-order model cases, 400 backend owner cases, both catalog
E2Es, 290 UI cases and 18 browser cases. Final grouped startup gzip is
347299 B under the unchanged 347353 B enforcement limit.

* refactor(ui): keep submission projection in lazy chat owner

Keep the app store responsible for bounded retained bytes and client lifetime.
Move receipt adaptation and display retirement into the existing history
projection owner, shared by both lazy chat consumers. Preserve missing-store
behavior and the retained-prompt, attachment and reconnect contracts.

Continue the retained-submission owner from PR #134059
(0f3e17e56b).

Validation: 808 owner tests, 21 Chromium cases, grouped-bundle Home and
retired-prompt proof, changed checks and fresh full-candidate autoreview.
Startup gzip: 347588 -> 347320 bytes; unchanged limit 347353.
2026-08-31 16:48:58 -07:00
Peter Steinberger
4716cbfe4b
refactor(agents): avoid unnecessary failure classification (#134008) 2026-08-31 05:05:14 -07:00
Peter Steinberger
0ec41639b1
fix(reply): report failures after a turn is accepted (#133979)
* fix(reply): report failures after a turn is accepted

* test(codex): prepare runtime policy in credential wait fixture
2026-08-31 02:57:26 -07:00
wanyongstar
a36fedbbf8
fix(outbound): skip code-span placeholders consumed by tag stripping (#125184)
* fix(outbound): skip code-span placeholders consumed by tag stripping

When a markdown code span sits inside an HTML tag's attribute region,
stripRemainingHtmlTags removes the tag together with the restore
placeholder. The restore loop then sliced with indexOf(...)= -1,
dropping the final character and rewinding the cursor, which garbled
and duplicated the remaining outbound text. Skip the missing
placeholder instead.

* fix(outbound): restore surviving code spans by identity

Keep protected regions associated with their own markers when HTML stripping
consumes an earlier attribute span. Restore markers and original NUL bytes
in one pass so literal marker-shaped text remains inert.

Extend owner regressions, public delivery integration, and the existing
mock Gateway scenario for hidden and interleaved code spans.

Co-authored-by: 万拥 0668000723 <wan.yong@xydigit.com>

* fix(outbound): document intentional sanitizer delimiters

Keep the deliberate NUL marker regex exempt from the control-character lint rule.

Co-authored-by: 万拥 0668000723 <wan.yong@xydigit.com>

* test(outbound): register intentional sanitizer delimiter suppression

Keep the exact suppression inventory aligned with the documented NUL marker regex.

Co-authored-by: 万拥 0668000723 <wan.yong@xydigit.com>

---------

Co-authored-by: Peter Steinberger <steipete@gmail.com>
Co-authored-by: 万拥 0668000723 <wan.yong@xydigit.com>
2026-08-30 22:25:42 -07:00
Jason (Json)
e1a15aaced
fix(agent): always reply after settled tool work (#133520)
* fix(agent): always finalize settled tool turns

* fix(agent): preserve silent helper failures

* fix(agent): preserve cancellation during fallback persistence

* test(qa): update settled fallback contracts

* chore: leave changelog to release generation

* fix(agent): preserve fallback writer ownership

* fix(agent): fence settled final reply delivery

* fix(agent): fence settled replies at transport

* fix(outbound): fence recovered final replies

* fix(outbound): fence final adapter handoff

* refactor(agent): keep run loop within lint budget

* fix(channels): fence provider sends after custody refresh

* test(infra): expect delivery handoff fence
2026-08-30 23:18:49 -06:00
Peter Steinberger
be3d86b8d3
refactor(context-engine): simplify compaction and turn handoffs (#133678)
* refactor(context-engine): simplify compaction and turn handoffs

Reject contradictory session identities before invoking compaction, then
forward the direct compactor's resolved target instead of rebuilding it
from obsolete marker outputs and session-row scans. Preserve thread routing
separately from storage identity and keep external successor fencing intact.

Trim private turn candidates to the eight fields their acceptance owner
consumes and build legacy lifecycle inputs only for legacy callers.

Cover durable summary/endpoint compaction, cancellation, accepted-turn
siblings, and the next real Gateway turn using a deterministic HTTP provider.

* test(context-engine): cover projected compaction identity

Exercise the real registry host-parameter projection with valid public calls. Preserve every omission, fallback, and rejection case without adding type suppressions or growing the inventory.
2026-08-30 22:08:27 -07:00