* fix(agents): keep spawned workers out of the current conversation
Agent-spawned workers (sessions_spawn thread=true, native and ACP) may only bind a new child thread. Channels whose spawn placement is the current conversation (Telegram, Feishu, LINE, generic current-conversation bindings) now reject thread=true instead of handing the user's chat to the worker. Legacy spawn-created bindings on those channels are ignored by the binding service, so the conversation routes to its normal agent again. Discord/Matrix child-thread sessions and defaultSpawnContext are unchanged.
* test(telegram): expect parent routing after worker-takeover upgrade
* test(agents): refresh prompt snapshots for child-only thread spawns
* test(telegram): keep parent-phase checkpoint codes in public upgrade evidence
Use Podman native init for PID 1 and verify HostConfig.Init before starting the isolated test container. Native diagnosis found the terminated watchdog retained as a zombie with the same process-start identity; the unchanged disappearance assertion passes with init. Preserve all isolation, cancellation, joining, snapshot-retention and assertion budgets; document the existing init-helper prerequisite.
Co-authored-by: vincentkoc <25068+vincentkoc@users.noreply.github.com>
Scope the missing-load-path fixture to published drivers that admit invalid configuration before staging. Retain ordinary update and migration checks for older drivers, preserve exact-row reruns, and record fixture applicability in the published receipt.
* refactor(telegram): move bundled thread bindings to workers
Await binding hydration and persistence through the existing plugin-state worker. Keep public synchronous SDK compatibility on the same owner while bundled callers use awaited operations. Preserve FIFO admission, current authority, committed acknowledgements, and shutdown drainage.
Validation on the frozen source: Focused binding, bot lifecycle and ingress tests, extension production and test typechecks, scoped lint, fresh review, a normal full build, and the built Telegram import profile.
This private checkpoint retains proof from base 78d5bedb34. It does not resolve the inherited TS1619/current-main qualification limit or the unchanged baseline TSGO_CORE_TEST_MAX_ROOTS unused-export finding. ACP startup metadata reads remain separately owned.
* test(telegram): dispatch sticker checks through the admitted handler
* test(telegram): verify published-driver binding upgrades
* test(telegram): keep upgrade runner private
* fix(telegram): await binding activity on bundled routes
* test(telegram): bind stop fixture to the active routing facade
Missing prerequisites for two test lanes, a copy-paste slip in an SDK
sample, three undeclared identifiers in a quick start, a colon promising
keys named two paragraphs later, duplicated Related lists, two unlinked
pages that exist, an unbracketed placeholder, and a bare doctor invocation.
The thirteenth fix, an ffmpeg prerequisite on docs/tools/tts/quickstart.md,
is held back: every PR touching that page is refused by the secret scanner
before review.
* ci: narrow PR tests by import and dependency impact
Resolve affected tests through the existing runtime import graph, workspace
and SDK aliases, configured Vitest inputs, and explicit policy owners.
Compare exact-base dependency closures instead of treating every lockfile
or package metadata change as global. Keep documentation and localization
data out of PR Node test rows.
Preserve canonical execution metadata and complete automatic main coverage.
Retain diagnostic fallbacks for global inputs and unverifiable ownership.
The 40-run replay reduces explicit fallbacks from 30 to 6 and median selected
jobs from 132 to 122, with no unexplained historical failure omissions.
The remaining large import closures do not meet the requested cost target.
* fix(ci): run test selector parsing under Node
Bun workers cannot provide the Node TypeScript parser used by the selector. Resolve the existing Node executable at the scanner boundary and retain that built-in-only dependency in materialized wrappers and standalone fixtures. Preserve real shard inventory exports in the command planner fixture.
* test(ci): isolate changed-gate compiler fixtures
Source-aware selection reaches inline compiler checks that bypassed the fixtures’ subprocess stubs, repeatedly compiling the real repository during gate-order and root-lint tests. Bind that owner to the shared synthetic recorder while preserving the real CLI, lint execution, serial-order assertions, and failure sentinels.
* fix(ci): retain compiler gates for TypeScript catalogs
Catalog-only PRs suppress Node test rows, but their TypeScript modules still need compiler validation. Admit the existing check owner independently of Node tests while preserving documentation and JSON-only skips. The catalog fixtures retain main and manual coverage and distinguish compiler admission from test execution.
* refactor: retire pre-June config and upgrade test support
Retire obsolete Doctor keys and their runtime fallbacks. Older configurations use the documented 2026.9.5 Doctor bridge; current upgrade verification starts at June 2026 while historical receipts remain readable.
* fix: preserve retired config when the upgrade bridge is skipped
* test: align upgrade fixtures with retained migration contracts
The project runner replaced timed-out configurations with another attempt
and could turn their failed hooks into a green job. CI run 35928401414,
job 107409413294, masked seven settlement-hook failures this way.
Remove replacement dispatch and its expanded retry deadline. Keep the first
watchdog failure even when a child exits zero, retain incomplete reports,
and preserve process/cache cleanup.
Proof: the temporary fail-first test failed the real CI shard entry point
with exit 143 and one attempt. The runner owner passed 20 standalone runs
and three CI-config replays on Testbox; shared-helper importers passed.
* ci: shorten real-Gateway UI validation
Use runtime-only preparation while build-artifacts retains SDK declaration checks. Preserve four exhaustive Gateway tours in full manual and release validation with their faster owner-boundary siblings in ordinary CI.
* test(ci): align real-Gateway preparation guards
Keep the paired runtime and UI build contract, assert the existing runtime-only mode, and document SDK validation ownership in the artifact job.
* ci: split real-Gateway UI validation into two jobs
* fix(ci): declare the shared real-Gateway parallel inventory
* test(ci): require manifest strings before decoding
* ci: balance standalone UI proofs across Gateway shards
* ci: defer desktop transport tour to release validation
* perf(test): reuse prebuilt Control UI assets
* test(ui): align Gateway fixtures with catalog and build contracts
* refactor: retire pre-June import and verification compatibility
Remove pre-June task, flow, and plugin-state sidecar imports, obsolete
runtime chunks, package/installer validation exceptions, the old MCP
attachment fallback, and the April self-upgrade lane with its orphan helpers.
Leave retired data files untouched and document migration through 2026.6.1.
Preserve June-and-later contracts and September delivery recovery receipts.
Refs #156190
* docs: route legacy upgrades through 2026.9.5
* test: await Telegram fixture lifecycle events
Replace the setup stopwatch with the actual stop event or terminal run outcome. Keep cancellation assertions and outer execution bounds, and prove early terminal outcomes fail promptly.
Use process workers on Windows so concurrent Vitest workers do not share temporary inheritable subprocess pipe handles. Preserve worker limits, file parallelism, and strict process/output cleanup. Route shared worker-policy changes through Windows CI.
* fix: keep web chat steering working after completed turns
* fix: observe completed progress refresh runs
* test: wait for durable startup recovery admission
Observe committed session rows before stopping startup recovery. Dispatch invocation precedes durable admission, so stopping after a call-count assertion can cancel the write and leave abortedLastRun set. Preserve restored-store capacity sequencing and wait for both stores without assuming recovery order.
Controlled original synchronization reproduces the CI assertion failure; corrected synchronization passes. Ten startup and cancellation cases pass in 26.64 seconds, with owner types, lint, formatting, line-cap guards and independent review through P2 clean.
Use the shared cleanup runner for provider test shards and track xAI fetch mocks so they cannot leak into later transport tests. The same 73-file shard improved from 88.658s to 41.593s across three runs per setting. All four native provider partitions passed: 290 files, 5,234 tests, and one existing opt-in live-test skip.
A test failure without a related change is a defect: re-running, re-pushing,
or refreshing a PR to get green is prohibited, and a failure outside the diff
is not "unrelated" until its cause and owning fix are identified. New or
changed tests state their measured cost and stay within the budgets in the
testing guide, which gains a cost-budget section and a flake-triage recipe.
* perf(test): reuse verified runtime worker artifacts locally
Retain joined compiler generations in exclusive checkout-local slots and
reuse their content-verified outputs on unchanged invocations. Preserve
source, output, resolution, toolchain, and borrower-lifetime checks while
avoiding repeated native compilation. CI routing and fresh-build policy,
Bun, custom loaders, test assertions, and timeouts remain unchanged.
Batch prepared fixture copies, relocate the complete fs-safe runtime
closure with its native packages, and inject one compiler fault child
per lifetime so receipt writers cannot race.
Related: #132712, #139428.
* fix(test): repair worker cache harness CI fixtures
Use measured CI-only six/eight-worker memory tiers for roomy serial self-hosted jobs. Retain local behavior, actual-host fallback limits, Gateway exclusivity, unproven group caps, and historical timing floors. The twelve predefined Blacksmith probe samples passed; native PR CI remains a separate validation step.
* fix: avoid proxy-blocked local UI E2E readiness waits
Fail early on restricted loopback transport, provide a secretless rootless isolated runner, and preserve meaningful HTTP readiness failures without changing Gateway semantics or proxy policy.
Co-authored-by: steipete <58493+steipete@users.noreply.github.com>
* fix: avoid proxy-blocked local UI E2E readiness waits
Worked on by:
- @steipete
Co-authored-by: steipete <58493+steipete@users.noreply.github.com>
OpenClaw-Publication: fc4d675b-1e2a-45e0-9494-70f01e9fdd35
* fix(test): report local Gateway transport failures early
Worked on by:
- @steipete
Co-authored-by: steipete <58493+steipete@users.noreply.github.com>
OpenClaw-Publication: 69d67c9d-c404-4ccf-ab4e-34ca9f134cdf
---------
Co-authored-by: steipete <58493+steipete@users.noreply.github.com>
* fix(test): run local Gateway tests behind protected exec egress
Worked on by:
- @steipete
Co-authored-by: steipete <58493+steipete@users.noreply.github.com>
OpenClaw-Publication: b694aecd-a905-4c37-b65d-e05463cc1a7a
* fix(test): register isolated Vitest process entry for audits
Model the actual path-launched container entry in the existing development-root registry so full-tree dead-code checks follow its imports.
Co-authored-by: steipete <58493+steipete@users.noreply.github.com>
* fix(test): retain isolated inputs until Podman processes join
Preserve nested unjoined host-command failures even when the container is absent. Retain the snapshot and original error rather than releasing inputs or suppressing unresolved cleanup as a signal exit.
Co-authored-by: steipete <58493+steipete@users.noreply.github.com>
---------
Co-authored-by: steipete <58493+steipete@users.noreply.github.com>
* fix(ci): ratchet line caps without blocking main
Independently green PRs can combine to exceed max-lines and block main.
Use explicit hosted warnings and an oxlint-backed changed-file growth check.
Keep canonical caps, suppression inventory, and strict local/landing lint.
* fix(ci): preserve warning counts and isolated lint runners
Count native colored diagnostics and load cumulative config only when
warning mode is requested. Restore strict-runner fixture isolation and
accurate hosted warning totals.
* fix(ci): align line-cap ratchet with global warnings
Main now warns for all max-lines scopes, making hosted-only demotion redundant. Keep the PR growth ratchet and document the shared warning policy and shrink-only maintenance.
* chore(lint): make max-lines a warning
Peter decided on 2026-09-16 that max-lines should warn in CI rather than
block main. The chat-pane-render.ts failure exposed unnecessary lint
burden after #149697.
Change all six max-lines scopes to warn while preserving every threshold
and all other rules. Document warning behavior and the retained suppression
ratchet. The runners and CI already preserve warning output and exit codes.
Proof: check-changed, formatting, and git diff --check pass. The current
chat-pane-render.ts passes single-file core lint; the exact 702-line
revision from failed main run 35060497703 prints one max-lines warning and
exits 0 through run-oxlint with CI-style output. The sample UI tsconfig
command hit the existing nested-checkout declaration boundary, so proof
uses CI's source-only core tsconfig.
* test(lint): expect max-lines warnings
Align the existing config-policy test with Peter’s 2026-09-16 decision. Preserve all six budgets and exclusions. The prior PR run correctly exposed the stale error expectations.
Proof: all 13 oxlint config tests, check-changed, and git diff --check pass. Fresh independent review found no actionable P0/P1 findings.
Publish bounded, host-redacted receipts after successful published-baseline
upgrade runs. Reuse the diagnostic owner and retain the initial post-core
result separately from repair and recovery output. Preserve Docker outcomes
and existing failure reports.
* test(docker): check the configurable ClawHub sweep spec
* fix(e2e): preserve live ClawHub package default
* fix(plugins): preserve live ClawHub package metadata
* docs(testing): document live ClawHub plugin opt-in
* chore: refresh ClawHub PR CI
Refresh the PR projection and checks against main with the wizard recovery test routing fix.
---------
Co-authored-by: Vincent Koc <vincentkoc@ieee.org>
Co-authored-by: Peter Steinberger <steipete@gmail.com>
* perf(test): admit isolated Gateway readers to the existing worker pool
* test: align Gateway scheduling and sandbox cache expectations
* test(ui): retain safe Gateway failure state in CI logs
The hosted job-cap regression rebuilt the growing repository's full compact
plan twelve times, resetting the module graph for every inventory variation.
Its discovery cache avoided repeat enumeration but left import, ownership,
and packing work unbounded; the case timed out at 120 seconds in hosted CI.
Use fixed timing inputs and a bounded synthetic inventory through the real
planner. Fifty-six full-budget anchors leave twenty-four tooling jobs at the
80-job cap, including a stronger-runner tail. Retain all twelve variants and
every original coverage, ownership, runner, and execution-budget assertion.
Document the fixture contract and assert exact baseline capacity/membership.
Measured case runtime fell from 8.858 seconds to 0.575 seconds locally. The
complete owner file passes all 61 tests; the unmocked repository planner still
produces 78 jobs within the cap. No runtime or workflow policy changes.
* test: cover transcript cold storage in release Docker lanes
* test: remove stale cloud picker test seams
Clear baseline CI failures introduced by #145183: reuse the existing picker fixture, test cloud configuration through its production renderer, and remove the unused Connect menu renderer.
* test: preserve historical transcript fixture modification times
Doctor imports filesystem mutation time rather than timestamps inside messages. Date the legacy files before import, and handle CLI failure diagnostics without adding a lint suppression or retaining credential-bearing command arguments.
* fix: keep cold sessions available to Activity title probes
Handle cold transcripts at the optional title-reader boundary without restoring payloads or caching missing previews. Verify real archival, mixed hot/cold listings and cache recovery; extend packaged release proof across restart and portable backup recovery. Preserve JSON CLI errors from stdout in the test harness.
* fix(ui): use the action cursor for session details
Clear the existing cursor-policy failure from #145183 without weakening its regression test.
* test: preserve frozen release targets in cold-storage lanes
Share capability resolution with source preflight and omit only unsupported cold subcases under the existing explicit frozen-target policy. Keep current coverage required, isolate ordinary retention in the live fixture, and strengthen fixture typing without suppressions.
* test: complete cold-storage release entrypoint contracts
Register the shell entrypoints for dependency analysis, type the frozen-source fixture cases explicitly, and align existing cloud test callback ordering with the same repair now on main.
* test: isolate archived recall from external tool policy
* test: retain the cloud machine locator for proof capture
live-workflows.md ended with "prefer narrowing live tests via the
allowlist env vars described below" — but the page ends on the next
line, and those env vars now live on docs/help/testing/docker.md after
the testing guide was split.
This is the failure mode no gate in this repository detects. The
sentence is valid prose and contains no link, so the link audit, the
formatter, markdownlint and the MDX check all pass. It reached main
through a full review.
Found with a new .audit/check-orphan-refs.py, which flags directional
phrasing for a human to confirm. Scoped to the six split trees it
reports 8 candidates, of which this was the one real defect; the rest
resolve on their own page.
* docs(reference): split the testing reference by reader job
docs/reference/test.md was 92,671 characters and mixed agent proof policy,
how-to steps, reference tables, and runner-internals explanation on one page.
It is now a short index over six children, one per reader job:
- reference/test/local Routine local order, core commands, PR gate
- reference/test/lanes Control UI, TUI, extension, Gateway, E2E lanes
- reference/test/docker Docker scheduler knobs and the notable lanes
- reference/test/performance Profiling, shard timings, benchmark scripts
- reference/test/runner-internals Build locks, test state, JSON report merging
- reference/test/remote-proof Crabbox/Testbox policy, wrapper, lease, trust
Anchor strategy: per-anchor redirect routes are impossible here, because
redirectSource() in scripts/lib/docs-redirects.mjs rejects any source
containing [?#]. Every pre-split anchor instead stays alive on the parent as
an authored <a id="..." /> stub in the "Where each section moved" list, the
same mechanism docs/ci.md uses. All 30 ids were enumerated with
parseDocsDocument, never a hand-rolled slug, so the four punctuated headings
keep both the encoded and the cleaned id (for example
full-docker-suite-(pnpm-test%3Adocker%3Aall) and
full-docker-suite-pnpm-testdockerall). 29 ids are stubbed; `related` is not,
because the index still publishes that heading itself, and stubbing it would
raise a duplicate authored/canonical ID collision.
Verified independently of docs-link-audit, which cannot see the regression:
the split rewrote the repo's own links, so the audit reads clean even when
external deep links break. Resolving all 30 pre-split ids against the parsed
post-split index gives 30 resolved, 0 dead, 0 collisions, and all 24 onward
deep links land on a real fragment of a real child.
Losslessness (bodies, frontmatter excluded):
code fences 20 -> 20 (+0)
table rows 47 -> 47 (+0)
inline code spans 530 -> 530 (+0, byte-identical multiset)
fenced blocks 10 -> 10 (byte-identical, so commands are unchanged)
chars 92,523 -> 92,253 on children + 5,093 on the index
words 10,008 -> 9,981 on children + 381 on the index
links 13 -> 9 on children + 35 on the index
The chars/words/links deltas reconcile exactly: the two intro bullets (170
chars) and the Related list (139 chars, 3 links) stay on the index, and one
declared edit adds 40 chars and one link.
The single declared prose edit: "Local test commands below are the normal
trusted development path" pointed at a section the split moves to another
page, so "below" became a link to /reference/test/local. No other prose was
rewritten; the remaining prose findings stay open for a follow-up.
Test pins repointed. test/scripts/docs-sync-publish.test.ts pins the exact
Release & CI navigation page list and breaks on the docs.json nav addition;
its route list now includes the six children. The two QA Lab producers point
their docsRefs at reference/test/local.md, the page that now owns the
commands they run, instead of the index. changed-lanes.test.ts needs no
change: it uses the path only as a docs-path example and asserts nothing
about the content.
Closes audit findings: r3-0695, r3-0696
* docs(reference): link the relocated local test commands
ClawSweeper found a second orphaned cross-reference the split missed.
remote-proof.md said "Local test commands below", but after the split
those commands live at /reference/test/local while this page continues
with remote-proof instructions, so "below" pointed at nothing.
docs-link-audit --anchors: 0 broken links.
docs/help/testing.md was 79,946 characters and mixed how-to, reference,
contributor conventions, and roadmap content in one page. It is now a
short index over six pages, one per reader job:
- help/testing/suites: quick start, the suite reference, which suite to
run, the live-test pointer, docs sanity, and offline regressions.
- help/testing/live-workflows: live provider debugging through the
Docker and Parallels lanes.
- help/testing/docker: the Docker "works in Linux" runners, their
weighted scheduler, the lane catalog, and their env vars.
- help/testing/qa-runners: the qa-lab command surface, the shared Convex
credential contract, and adding a channel to QA.
- help/testing/contracts: plugin and channel contract tests.
- help/testing/writing-tests: temp-directory rules, the agent
reliability eval gaps, and how to add a regression.
Anchor strategy: per-anchor routes are impossible because redirectSource()
rejects any source containing [?#]. Instead every one of the 49 ids the
old page published stays alive on the index as an authored <a id="..." />
stub in "Where each section moved". Ids were computed with
parseDocsDocument, not a slug approximation, so punctuated headings keep
both emitted forms (for example
docker-runners-(optional-%22works-in-linux%22-checks) and
docker-runners-optional-works-in-linux-checks). The index still publishes
`related` itself, so that id is deliberately not stubbed and no
duplicate authored/canonical ID is raised.
Losslessness, asserted mechanically rather than by eye: all 14 original
section bodies are character-identical after the move (0 lost, 0
changed), and the page lede is byte-identical. Word count 9,632 -> 9,632,
code fences 18 -> 18, links 13 -> 13, table rows 0 -> 0. All 751 inline
code spans and fenced blocks compare as an identical set, so every
command in the guide is unchanged. The only body delta is six trailing
newlines removed by scripts/format-docs.mts.
Verified independently of docs-link-audit, which a split makes
uninformative because it rewrites the repo's own links: the 49 pre-split
ids were enumerated from HEAD, the post-split index and every child were
re-parsed, and each id was asserted to resolve. 0 unresolved, 0 stub
targets that miss their child, 0 collisions.
Prose findings are deliberately left alone and deferred: r3-0398,
r3-0399, r3-0400, r3-0401, r3-1686, r3-1687, r3-1688, r3-1689, r3-1690.
Closes audit findings: r3-0397