Commit graph

925 commits

Author SHA1 Message Date
Peter Steinberger
ec66133ba0
chore: cover worker state in published upgrade tests (#149487)
* test: add worker state upgrade survivor cells

Add opt-in Projects Doctor and terminal task/flow restoration cells using the unchanged published updater, verified package bytes, and the canonical survivor lifecycle. Preserve failed synthetic state until the outer Docker owner has joined.

Validation: 161 focused tests, selected changed checks, and P2 review. Actual package upgrade cells remain separately qualified against frozen candidate artifacts.

* test: activate taskflow survivor fixture on Gateway startup

* test: clean worker runtime with container ownership

* fix(test): preserve Docker status through exit traps
2026-09-15 21:51:41 -07:00
Vincent Koc
f3579e762e
fix(e2e): use a maintained ClawHub install fixture (#148902)
* fix(e2e): use a maintained ClawHub install fixture

* test(e2e): model Docker package mount in wrapper probe
2026-09-16 11:57:19 +08:00
Peter Steinberger
1ac3cf63b8
fix(ci): publish receipts for successful upgrade checks (#148245)
Publish bounded, host-redacted receipts after successful published-baseline
upgrade runs. Reuse the diagnostic owner and retain the initial post-core
result separately from repair and recovery output. Preserve Docker outcomes
and existing failure reports.
2026-09-14 13:01:06 -07:00
Saurabh Yadav
98ab09cae2
fix(plugins): preserve live ClawHub package metadata (#148118)
* test(docker): check the configurable ClawHub sweep spec

* fix(e2e): preserve live ClawHub package default

* fix(plugins): preserve live ClawHub package metadata

* docs(testing): document live ClawHub plugin opt-in

* chore: refresh ClawHub PR CI

Refresh the PR projection and checks against main with the wizard recovery test routing fix.

---------

Co-authored-by: Vincent Koc <vincentkoc@ieee.org>
Co-authored-by: Peter Steinberger <steipete@gmail.com>
2026-09-14 10:48:18 -07:00
Peter Steinberger
a5cc2078b8
fix: preserve outbound provider policy when the current target is missing (#147939)
* fix: enforce outbound provider policy without a current target

Preserve correlated recovery source providers independently of complete delivery authority. Keep cross-provider defaults and explicit overrides unchanged across restart and requester-settle continuations.

* test: preserve configured cross-provider opt-ins during recovery

* test: cover provider policy on explicit source reply routes

The explicit WebChat route fixture depended on the missing-target policy gap. Verify both the denial without opt-in and the unchanged normal outbound route when explicitly allowed.

* fix: preserve unbound CLI messaging context

* test: distinguish unbound and WebChat agent requests
2026-09-14 05:28:08 -07:00
Peter Steinberger
7043f8fb4f
feat(subagents): explain waits and separate execution from result delivery (#147571)
* feat(subagents): make task waits, controls, and delivery observable

* test(subagents): make live status and evidence probes explicit

* test(subagents): refresh model-facing tool snapshots

* test(tasks): reuse the task-owned Gateway fixture

* test(subagents): reuse the Gateway test-client owner

* refactor(tasks): consolidate restored indexes and settle lifecycle proof

* refactor(ui): reuse task status copy in subagent tooltips

* build(workboard): refresh assets after subagent protocol integration

* perf(ui): load background task copy with its deferred views
2026-09-13 23:29:48 -07:00
Peter Steinberger
97faa3e5d0
test(agents): cover live subagent yield and operator resume (#146874)
* test(agents): cover live subagent yield and operator resume

* test(agents): match live fanout admission to workload
2026-09-13 02:33:27 -07:00
Peter Steinberger
3c16ced267
perf(test): admit isolated Gateway readers to the existing worker pool (#146550)
* perf(test): admit isolated Gateway readers to the existing worker pool

* test: align Gateway scheduling and sandbox cache expectations

* test(ui): retain safe Gateway failure state in CI logs
2026-09-12 23:03:10 -07:00
Peter Steinberger
bd002c80a1
fix(update): preserve env shorthand after service reinstall (#146660)
A systemd Gateway whose unit lists a managed environment key and whose config uses `${VAR}` shorthand can start again after the 2026.9.4 service reinstall. Startup, service planning, and secrets audit share the existing config-resolution provenance to recognize authored environment references after substitution; escaped literals remain literal, stale keys are still removed, and configuration files remain unchanged.

Fixes #146612

Reported by @foxsky (#146612).

Validation: focused startup, installer, Doctor, config, audit, and runtime suites passed; built Linux before/after startup proof retained the key and reached readiness with unchanged config bytes. Exact-head CI was reverified green. A complete published-updater service-regeneration run remains a documented proof gap.
2026-09-12 20:05:59 -07:00
Peter Steinberger
055250e760
chore: cover custom-plugin sibling imports in full release validation (#146636)
* test: cover custom-plugin sibling imports in release validation

* fix: complete sibling scenario release harness integration
2026-09-12 19:11:00 -07:00
Peter Steinberger
54a986072f
docs(plugins): remove obsolete Gateway restart guidance (#146516)
* docs(plugins): remove obsolete Gateway restart guidance

* docs(plugins): simplify apply hints and update Session Share guidance
2026-09-12 16:48:13 -07:00
Peter Steinberger
3818b8e91d
fix(test): fail expired Vitest runs after clean child exits (#145999) 2026-09-12 07:14:24 -07:00
Peter Steinberger
760f8b7715
test(ci): bound hosted tooling growth inventory (#145759)
The hosted job-cap regression rebuilt the growing repository's full compact
plan twelve times, resetting the module graph for every inventory variation.
Its discovery cache avoided repeat enumeration but left import, ownership,
and packing work unbounded; the case timed out at 120 seconds in hosted CI.

Use fixed timing inputs and a bounded synthetic inventory through the real
planner. Fifty-six full-budget anchors leave twenty-four tooling jobs at the
80-job cap, including a stronger-runner tail. Retain all twelve variants and
every original coverage, ownership, runner, and execution-budget assertion.
Document the fixture contract and assert exact baseline capacity/membership.

Measured case runtime fell from 8.858 seconds to 0.575 seconds locally. The
complete owner file passes all 61 tests; the unmocked repository planner still
produces 78 jobs within the cap. No runtime or workflow policy changes.
2026-09-12 00:42:16 -07:00
Peter Steinberger
f9ccd07ded
fix: skip official setup approvals and default to Astra (#145646)
* fix: skip official setup approvals and default to Astra

* test: align default-model expectations across consumers

* test: align attachment catalog with Astra default

* fix: preserve configured Codex catalog model selections
2026-09-12 00:34:25 -07:00
Peter Steinberger
31848a1520
fix: keep cold sessions visible and verify release recovery (#145462)
* test: cover transcript cold storage in release Docker lanes

* test: remove stale cloud picker test seams

Clear baseline CI failures introduced by #145183: reuse the existing picker fixture, test cloud configuration through its production renderer, and remove the unused Connect menu renderer.

* test: preserve historical transcript fixture modification times

Doctor imports filesystem mutation time rather than timestamps inside messages. Date the legacy files before import, and handle CLI failure diagnostics without adding a lint suppression or retaining credential-bearing command arguments.

* fix: keep cold sessions available to Activity title probes

Handle cold transcripts at the optional title-reader boundary without restoring payloads or caching missing previews. Verify real archival, mixed hot/cold listings and cache recovery; extend packaged release proof across restart and portable backup recovery. Preserve JSON CLI errors from stdout in the test harness.

* fix(ui): use the action cursor for session details

Clear the existing cursor-policy failure from #145183 without weakening its regression test.

* test: preserve frozen release targets in cold-storage lanes

Share capability resolution with source preflight and omit only unsupported cold subcases under the existing explicit frozen-target policy. Keep current coverage required, isolate ordinary retention in the live fixture, and strengthen fixture typing without suppressions.

* test: complete cold-storage release entrypoint contracts

Register the shell entrypoints for dependency analysis, type the frozen-source fixture cases explicitly, and align existing cloud test callback ordering with the same repair now on main.

* test: isolate archived recall from external tool policy

* test: retain the cloud machine locator for proof capture
2026-09-11 20:25:40 -07:00
Ayaan Zaidi
ab0ea18607
fix(models): require explicit GitHub Copilot activation (#144749)
Require explicit Copilot configuration, a saved profile, or COPILOT_GITHUB_TOKEN within the existing agent scope. Generic GitHub credentials continue to serve other tools and remain covered by secret audit and cleanup.

Retire the discovery opt-out through the shared Doctor migration and add one-time upgrade guidance. Update provider documentation and generated configuration inventories.

Validated with focused plugin and secret tests, real Gateway starts across all 16 sanitized configs, and independent CLI checks for activation, Doctor notices, secret detection, and matching-value cleanup.

Related: #144726

Co-authored-by: Ayaan Zaidi <hi@obviy.us>
2026-09-11 13:51:44 +05:30
Peter Steinberger
a58c71a8f9
improve: use compiled Doctor contracts in built checkouts (#144676)
* perf(plugins): reuse compiled Doctor contracts consistently

* refactor(plugins): separate runtime bindings from artifact selection

Keep Doctor filesystem selection independent of runtime record types, migrate every binding caller, and verify repaired persistence through the built CLI instead of a source-mode observer.
2026-09-11 00:49:15 -07:00
Vyctor H. Brzezowski
26c0a27ae9
docs: remove repeated Docker plugin test summary (#144406) 2026-09-10 18:32:26 -03:00
Vincent Koc
efefa622e2
docs: fix 20 link defects confirmed live in the verified backlog (#144133)
Mis-pointed targets where a correct target exists, first mentions of
documented surfaces that were never linked, and pages with no Related
section at all. Adds 57 internal links, none broken.

These sat in the ledger's verified bucket, which an earlier round did not
have in scope when it reported this category closed.

Co-authored-by: Vincent Koc <vincent@openclaw.org>
2026-09-11 03:48:10 +08:00
Vincent Koc
f486f76460
docs: close the remaining ia and ste findings (#144089)
Two nav repairs (an orphaned reference page with a real inbound link, and
two start/ pages sitting under Help > Community), Related lists and index
cards that omitted whole top-level areas, and a plain-English pass over the
pages carrying rate findings.

Hard STE violations across the 46 measured pages: 649 -> 334 at the 25-word
cap, 767 -> 388 at 20. Every page named by a rate row is now under 1.5 at
both caps except reference/templates/AGENTS.md, refuted separately.

Co-authored-by: Vincent Koc <vincent@openclaw.org>
2026-09-10 22:28:26 +08:00
Vincent Koc
c0b128344c
docs(gateway,concepts,install,help): fix information-architecture findings (#143977)
Structure-only pass over 50 open `ia` audit rows in docs/gateway/,
docs/concepts/, docs/install/ and docs/help/. No page splits, no page
moves, no URL changes.

On-page structure:
- concepts/multi-user: six H3s inside the 1,288-word per-person accounts section
- concepts/model-failover: one notices section; H2 blocks reordered to
  storage -> rotation -> cooldowns -> fallback -> notices
- concepts/compaction: 'Provider checkpoints' and 'Successor transcripts'
  regrouped under a new 'Provider and engine behavior' H2
- concepts/memory-builtin: 'When to use' moved after 'What it provides'
- concepts/queue: 'Scope and guarantees' split into 'Input durability' and
  'Lanes and scope'
- concepts/session, concepts/session-tool: 'Further reading' merged into 'Related'
- help/debugging: sections reordered (watch mode first); 'Safety notes'
  demoted to H3 under raw stream logging; Node/tsx errors beside VSCode
- help/environment: OPENCLAW_HOME moved under 'Paths and instances'
- help/testing-updates-plugins: 'On this page' section index
- gateway/logging, gateway/multi-tenant-hosting, install/backups: duplicate
  body H1 removed (precedent 204971f2a9), old id kept as an <a id> stub
- install/upstash: 'Next steps'; vps.md: Upstash Box and Render cards

Navigation (docs.json), no URL changes:
- install/nix -> Runtimes; install/ansible -> Hosting > Self-hosted and local
- gateway/clients and gateway/external-apps ahead of gateway/protocol
- gateway/cli-backends, local-models, local-model-services -> new
  'Models and local providers' group
- gateway/heartbeat -> Capabilities > Automation
- gateway/portals -> Web interfaces
- gateway/security/dependency-locking -> Release & CI > Release process
- concepts/typing-indicators -> Messages and delivery
- concepts/usage-tracking, concepts/timezone -> Technical reference

Anchors: 15 changed pages enumerated with parseDocsDocument before and
after. Zero ids lost, zero collisions; 11 added.
2026-09-10 19:20:49 +08:00
Vincent Koc
47ff7bfb11
docs: close remaining cross-link gaps across concepts, gateway, and security (#143923)
Adds the missing reciprocal links and one mis-targeted link fix for the
last open `link` audit findings.

- Related back-links: system prompt (context engine, timezone), diagnostics
  flags (gateway diagnostics, gateway troubleshooting), cloud workers
  (operator scopes), auth credential semantics (secrets, auth storage),
  agent runtime architecture (agent runtimes), agent runtime workflow
  (testing), network (remote access, architecture), threat model (gateway
  security index, network proxy), backups/updating/doctor (database schemas).
- docs/security/network-proxy.md gains a Related section.
- docs/security/incident-response.md links back to the three sibling pages
  that already link to it.
- docs/concepts/typing-indicators.md links the heartbeat and groups pages
  that its Defaults section describes.
- docs/diagnostics/flags.md links the environment-variable reference from
  the timeline section that names three OPENCLAW_DIAGNOSTICS_* variables.
- docs/concepts/main-session.md names `session.maintenance.maxDiskBytes` and
  links the maintenance reference instead of stating a bare 10 GB default.
- docs/network.md pointed its "Gateway config reference" entry at
  /gateway/configuration; retargeted to /gateway/configuration-reference.
- docs/openclaw-agent-runtime.md merges its References list into Related and
  keeps the old `#references` anchor as a stub.
- Six zh-CN glossary sources added beside their related existing terms.
2026-09-10 18:34:34 +08:00
Vincent Koc
fa3bc9a244
docs: fix and reciprocate cross-page links across install, help, reference, start, and web (#143773)
Closes link-kind audit findings for docs/install/, docs/help/, docs/reference/,
docs/start/, docs/web/ and docs/nodes/.

- start/hubs: point the Model providers hub entry at the provider directory
  (/providers) instead of the /providers/models quickstart duplicate.
- Add the missing reciprocal links the audit found: Cloudflare Containers,
  Kubernetes and Ansible from Docker and the Linux server page; macOS VMs from
  iMessage and the Linux server page; Podman from the sandbox Podman backend;
  Updating from Migrating; Docker from the Configuration page; Backups,
  Bootstrapping and Default AGENTS.md from Agent workspace; Bootstrapping from
  the BOOTSTRAP template; Tests from the two help testing pages; Session
  management deep dive from Context engine; Transcript hygiene from Session and
  Session pruning; SecretRef credential surface from Auth credential semantics;
  Device model database from the Nodes macOS section; RPC adapters from Signal
  and iMessage; Personal assistant setup from Getting started; Onboarding from
  the macOS platform page; The Lobster from Control UI settings; Release
  performance sweep from Dependency locking; Release policy from Release
  channels; Full release validation and Update and plugin tests from RELEASING.
- help/index: list the Scripts page under Testing.
- help/faq-first-run: add the Models FAQ to Related (was one-directional).
- reference/credits: replace the two off-topic Related links with the lore and
  pull-request-review-flow pages.
- reference/rich-output-protocol: replace the unrelated RPC adapters link with
  the Control UI hosted-embeds section that actually renders [embed ...].
- web/lobster: add a Related section.
- glossary: 17 append-only zh-CN sources for the new list-item link labels, each
  inserted beside a related existing term rather than at the end of the array.
2026-09-10 15:35:25 +09:00
Dallin Romney
0b92bfeda7
fix(mac-gateway): stop discarding launchd stderr (#143519)
macOS launchd discarded Gateway stderr during both installation and restart, leaving startup failures invisible before logging initialized. Route stderr to the existing supervisor stdout log, matching the diagnostic reader and retaining the existing log-retention owner. Update status hints and troubleshooting documentation to identify the shared log.

Ordinary restart rewrites and reloads the generated plist; preserveDefinition retains its existing behavior. Hosted CI passed on a2616fa8dd86c1b76fc0f4f83e61b6ef640b9c60, with install/restart plist and diagnostic coverage: https://github.com/openclaw/openclaw/actions/runs/34438914271. Actual after-fix launchd capture remains unverified.

Fixes #90711.
2026-09-09 22:52:38 -07:00
Shakker
61f7614641
fix: distinguish interrupted dev runners from completed shutdowns (#143676)
Preserve Unix signal termination through the native runner and watcher, and distinguish acknowledged UI shutdown from interrupted or unverified cleanup.
2026-09-10 06:04:23 +01:00
Vincent Koc
faef46e142
docs: split help/faq into 13 topic pages (#143150)
* docs: split help/faq into 13 topic pages

docs/help/faq.md was 89,648 characters and 1,613 lines - the largest
hand-written page left in the tree. It is now a 31,704-character index:
a table of the thirteen topic pages, the triage ladder it opens with,
the two pointer sections to the first-run and models FAQs, and an
anchor-compatibility list.

Content-preserving. The thirteen children are verbatim slices cut on
H2 boundaries, so no AccordionGroup is split and no prose is rewritten,
reordered, or reformatted.

All 143 pre-split anchor ids still resolve on /help/faq: 14 are still
published by the index itself, and the other 129 are authored <a id>
stubs pointing at the page that now holds the answer. Zero id
collisions on the parent and on every child.

Matches the docs/help/testing/ and docs/help/testing-live/ split shape
already on main.

* docs: link the relocated data-locality answer from the Foundation answer

Addresses the ClawSweeper P3 finding on
docs/help/faq/what-is-openclaw.md:68. The Foundation answer said to see
"Is all data used with OpenClaw saved locally?" below; after the split
that answer is on docs/help/faq/where-things-live-on-disk.md, so the
directional reference is replaced with an explicit link to it.

This is the second and last declared prose rewrite in this PR. With both
reversed, the children still reassemble byte-identically to the original
docs/help/faq.md (sha256 fe2dd3d6...).
2026-09-09 23:40:39 +09:00
Vincent Koc
82050740b5
docs(help): split the first-run FAQ into two child pages (#143148)
`docs/help/faq-first-run.md` was 40,388 characters in one `## ` section
holding 52 accordions across two `<AccordionGroup>` blocks.

An `<AccordionGroup>` cannot be split across files without inventing a
second wrapper, so the only content-preserving cut is the boundary
between the two existing groups:

- `help/faq-first-run/quick-start` - install, onboarding, first-run
  failures, builds, and subscription basics (group 1).
- `help/faq-first-run/providers-and-hosting` - provider auth and limits,
  model choice, hardware, and where to run the Gateway (group 2).

The parent stays as an index. Every one of the 56 pre-split ids still
resolves on it: `related` is still published there, and the other 55 are
authored `<a id="...">` stubs pointing at the child that now holds the
content, matching `docs/help/testing.md`.

Content-preserving: the two child bodies, with frontmatter and ledes
removed, concatenate byte-identically to the original body
(sha256 040670d96d7a8819c40d5f6bb924cccf1e9cd48596f5c264517a388cfa773911).
No prose was rewritten, reordered, or added inside the moved content.
2026-09-09 23:12:13 +09:00
Vincent Koc
aabfc3e753
docs(help): split the live testing suites page by reader job (#143042)
docs/help/testing-live.md was 46,841 characters of 17 independent live
lanes. Split it into docs/help/testing-live/ children, one per reader job,
and keep the parent as an index that still carries the safety and
credential sections plus an anchor stub for every pre-split heading id.

Content-preserving: the children are verbatim slices of the original body.

Co-authored-by: Vincent Koc <vincent@openclaw.org>
2026-09-09 19:36:37 +09:00
Vincent Koc
8dc256d5c8
docs: fix accuracy findings in concepts, start, install, and help (#143029)
* docs: fix accuracy findings in concepts, start, install, and help

Corrects statements in the onboarding docs that disagree with the code, and
scopes version-locked claims to the release that changed them.

Each change is backed by a source reference; findings that turned out to be
wrong about the code are listed in the PR body rather than "fixed".

* docs(retry): separate Discord Gateway reconnects from request retry defaults

The "Applies to" column listed gateway reconnects alongside Discord sends.
The WebSocket reconnect loop does not use that envelope: gateway.ts allows
50 attempts and backs off from 2000 ms with no jitter, while retry.ts governs
per-request retries only.

Addresses the ClawSweeper P2 finding.
2026-09-09 19:34:49 +09:00
Vincent Koc
26b84d3f38
docs(tools): split the browser page by reader job (#142984)
* docs(tools): split the browser page by reader job

docs/tools/browser.md was 55,693 characters across 21 H2 sections mixing
explanation, how-to, reference, and troubleshooting. Move each section into
docs/tools/browser/ and keep /tools/browser as the index.

The split is content-preserving: the nine child pages plus the three blocks
the index still publishes (lede, What you get, Related) reconstruct the
original body byte-for-byte (sha256 a28d5ea4...). Only the four H2 lines that
became child page titles are not reproduced verbatim; each keeps its anchor
on the index.

Every one of the 43 pre-split anchor IDs -- headings, their percent-encoded
and cleaned variants, and the Accordion and Tab titles -- is authored as an
<a id> stub on the index, so existing /tools/browser#... links still resolve.
Per-anchor redirects are not possible: redirectSource() rejects sources
containing [?#].

* docs(tools): retarget cross-references orphaned by the browser split

Eleven link targets, no prose changes. Three were same-page or same-route
anchors inside the moved content: /tools/browser/setup's [Configuration] link
was genuinely broken (its target moved to the configuration page), and the
[Profiles] and [Custom Chrome MCP launch] links resolved only through the
index's anchor stub. The remaining eight are inbound deep links from other
pages, retargeted at the page that now holds the section, matching the
code-mode split precedent.

* docs(i18n): add zh-CN glossary entries for the browser child page titles

check-docs-i18n-glossary requires a source term for every changed doc label.
2026-09-09 18:48:49 +09:00
Vincent Koc
e36233ea68
docs(help): link the allowlist env vars orphaned by the testing split (#142560)
live-workflows.md ended with "prefer narrowing live tests via the
allowlist env vars described below" — but the page ends on the next
line, and those env vars now live on docs/help/testing/docker.md after
the testing guide was split.

This is the failure mode no gate in this repository detects. The
sentence is valid prose and contains no link, so the link audit, the
formatter, markdownlint and the MDX check all pass. It reached main
through a full review.

Found with a new .audit/check-orphan-refs.py, which flags directional
phrasing for a human to confirm. Scoped to the six split trees it
reports 8 candidates, of which this was the one real defect; the rest
resolve on their own page.
2026-09-09 05:20:39 +08:00
Sally O'Malley
8da4fc5577
feat(config): support externally managed read-only configuration (#140719)
Add OPENCLAW_CONFIG_READONLY=1 to the shared immutable-config policy so externally managed deployments do not need unrelated Nix behavior.

Keep the selector host-owned across config reload and service regeneration. Gate normal writes and recovery rewrites before altering source bytes or creating recovery artifacts, while preserving mutable defaults and existing Nix guidance. Include compatibility and failing-before recovery regressions.

Closes #140706

Worked on by:
- @sallyom

Co-authored-by: sallyom <11166065+sallyom@users.noreply.github.com>
2026-09-08 12:49:31 -07:00
Vincent Koc
caee06871a
docs(gateway): split the security overview by reader job (#141162)
docs/gateway/security/index.md was 89,057 characters, 9,905 words and 25 H2
sections mixing explanation, how-to, reference and vulnerability-report triage
policy for two audiences. The page already lived in a directory with five
siblings, so the split extends that directory rather than creating a parallel
one, and index.md becomes a real index.

New children in docs/gateway/security/ (alongside audit-checks, exposure-runbook,
rate-limiting, secure-file-operations and dependency-locking):

- trust-model.md (Security trust model) - scope, trust boundary matrix, findings
  closed as no-action, gateway/node trust, threat model, reporting.
- running-the-audit.md - the `openclaw security audit` command, what it checks,
  triage order.
- hardened-baseline.md - both copy/paste baselines and the requester-scoped
  controls note.
- access-control.md - DM policy, allowlists, DM session isolation, context
  visibility, command authorization.
- prompt-injection.md - prompt injection, external-content wrapping, bypass flags.
- tool-permissions.md - control-plane tools, node execution, dynamic skills,
  plugins, sandboxing, per-agent access profiles.
- browser-control.md - browser control risks and the SSRF policy.
- network-exposure.md - bind/firewall, Docker/UFW, mDNS, Gateway auth, Tailscale,
  reverse proxy, HSTS, Control UI over HTTP, dangerous flags.
- secrets-and-storage.md - host trust, secrets on disk, credential map,
  permissions, workspace .env, logs and transcripts, secret scanning.
- operator-incident-response.md - contain, rotate, audit, collect.

Anchor strategy

Per-anchor redirects are not possible: redirectSource() in
scripts/lib/docs-redirects.mjs rejects any source containing [?#]. Every anchor
the old page published is therefore kept alive on the index itself as an
authored <a id="..." /> stub inside a "Where each section moved" list, each
pointing at its new home. Ids were computed with parseDocsDocument, not a slug
approximation, so percent-encoded punctuation and compatibility aliases are
preserved exactly. The index publishes no original heading itself, so no stub
can collide with a canonical id.

- pre-split ids on /gateway/security: 80 (57 headings, 20 compatibility
  aliases, 3 accordion targets)
- still published by /gateway/security after the split: 80
- index collisions reported by parseDocsDocument: 0
- stub destinations checked against the destination pages: 60 (all resolve)

Losslessness

Reassembling the child section bodies in original order reproduces the original
body byte for byte. Counts, original body vs children:

- words 9,665 -> 9,664
- characters 87,422 -> 87,491
- code fences 32 -> 32
- markdown links 44 -> 45
- table rows 38 -> 38

The single word/link delta is the one declared prose edit: a cross-reference the
split orphaned, `(see "Per-agent access profiles" above)` in the secure-baseline
section, is now a link to /gateway/security/tool-permissions. The index intro
line "The rest of this page is the deep end" became "The pages below are the
deep end" for the same reason. No security control, threat-model statement,
default or guarantee was reworded, softened or reordered; sections keep their
original relative order within each child.

Also updated: the "Security and sandboxing" nav group in docs/docs.json, eleven
in-repo deep links repointed at the new pages (release notes left untouched, and
their anchors still resolve through the stubs), and
src/docs/environment-docs.test.ts, which asserts on the workspace-dotenv text
that now lives in secrets-and-storage.md.

Closes audit findings: r3-0302, r3-0308
2026-09-08 19:35:27 +00:00
Vincent Koc
d72936b37e
docs(reference): split the testing reference by reader job (#141221)
* docs(reference): split the testing reference by reader job

docs/reference/test.md was 92,671 characters and mixed agent proof policy,
how-to steps, reference tables, and runner-internals explanation on one page.
It is now a short index over six children, one per reader job:

- reference/test/local            Routine local order, core commands, PR gate
- reference/test/lanes            Control UI, TUI, extension, Gateway, E2E lanes
- reference/test/docker           Docker scheduler knobs and the notable lanes
- reference/test/performance      Profiling, shard timings, benchmark scripts
- reference/test/runner-internals Build locks, test state, JSON report merging
- reference/test/remote-proof     Crabbox/Testbox policy, wrapper, lease, trust

Anchor strategy: per-anchor redirect routes are impossible here, because
redirectSource() in scripts/lib/docs-redirects.mjs rejects any source
containing [?#]. Every pre-split anchor instead stays alive on the parent as
an authored <a id="..." /> stub in the "Where each section moved" list, the
same mechanism docs/ci.md uses. All 30 ids were enumerated with
parseDocsDocument, never a hand-rolled slug, so the four punctuated headings
keep both the encoded and the cleaned id (for example
full-docker-suite-(pnpm-test%3Adocker%3Aall) and
full-docker-suite-pnpm-testdockerall). 29 ids are stubbed; `related` is not,
because the index still publishes that heading itself, and stubbing it would
raise a duplicate authored/canonical ID collision.

Verified independently of docs-link-audit, which cannot see the regression:
the split rewrote the repo's own links, so the audit reads clean even when
external deep links break. Resolving all 30 pre-split ids against the parsed
post-split index gives 30 resolved, 0 dead, 0 collisions, and all 24 onward
deep links land on a real fragment of a real child.

Losslessness (bodies, frontmatter excluded):
  code fences        20 -> 20   (+0)
  table rows         47 -> 47   (+0)
  inline code spans 530 -> 530  (+0, byte-identical multiset)
  fenced blocks      10 -> 10   (byte-identical, so commands are unchanged)
  chars           92,523 -> 92,253 on children + 5,093 on the index
  words           10,008 ->  9,981 on children +   381 on the index
  links               13 ->      9 on children +    35 on the index
The chars/words/links deltas reconcile exactly: the two intro bullets (170
chars) and the Related list (139 chars, 3 links) stay on the index, and one
declared edit adds 40 chars and one link.

The single declared prose edit: "Local test commands below are the normal
trusted development path" pointed at a section the split moves to another
page, so "below" became a link to /reference/test/local. No other prose was
rewritten; the remaining prose findings stay open for a follow-up.

Test pins repointed. test/scripts/docs-sync-publish.test.ts pins the exact
Release & CI navigation page list and breaks on the docs.json nav addition;
its route list now includes the six children. The two QA Lab producers point
their docsRefs at reference/test/local.md, the page that now owns the
commands they run, instead of the index. changed-lanes.test.ts needs no
change: it uses the path only as a docs-path example and asserts nothing
about the content.

Closes audit findings: r3-0695, r3-0696

* docs(reference): link the relocated local test commands

ClawSweeper found a second orphaned cross-reference the split missed.
remote-proof.md said "Local test commands below", but after the split
those commands live at /reference/test/local while this page continues
with remote-proof instructions, so "below" pointed at nothing.

docs-link-audit --anchors: 0 broken links.
2026-09-09 02:15:05 +08:00
Peter Steinberger
c21fa54066
docs(install): add Node.js and Bun compatibility reference pages (#142156)
The Node requirement has changed nine times in 2026 and the reasons (the
SQLite WAL-reset corruption floor and the separate node:sqlite TEXT NUL
decoder bug) were buried in install prose; the Bun page's Caveats section
had become the de facto Bun contract while sitting under Containers.

- Add docs/install/node-compatibility.md: supported lines, why the floors
  exist, platform consequences, what each installer provisions, the guard
  diagnostic, and a sourced history of the requirement across releases.
- Add docs/install/bun-compatibility.md: Bun requirements, per-platform
  SQLite builds, macOS library selection and OPENCLAW_SQLITE_LIBRARY with
  the preload migration, the memory scan fallback, known limitations, and
  release history.
- Keep docs/install/node.md and docs/install/bun.md as install how-tos;
  move the contract paragraphs to the new pages and link them.
- Add a Runtimes nav group (Node, Node compatibility, Bun, Bun
  compatibility) and move Bun out of Containers; no URLs change.
- Point the environment reference, memory config, and install overview at
  the new pages; add zh-CN glossary entries; add a docs guide bullet to
  refresh the tables when the runtime floors in code change.
2026-09-08 05:53:53 -07:00
Vincent Koc
cf4abffc00
docs(help): split the testing guide by reader job (#141968)
docs/help/testing.md was 79,946 characters and mixed how-to, reference,
contributor conventions, and roadmap content in one page. It is now a
short index over six pages, one per reader job:

- help/testing/suites: quick start, the suite reference, which suite to
  run, the live-test pointer, docs sanity, and offline regressions.
- help/testing/live-workflows: live provider debugging through the
  Docker and Parallels lanes.
- help/testing/docker: the Docker "works in Linux" runners, their
  weighted scheduler, the lane catalog, and their env vars.
- help/testing/qa-runners: the qa-lab command surface, the shared Convex
  credential contract, and adding a channel to QA.
- help/testing/contracts: plugin and channel contract tests.
- help/testing/writing-tests: temp-directory rules, the agent
  reliability eval gaps, and how to add a regression.

Anchor strategy: per-anchor routes are impossible because redirectSource()
rejects any source containing [?#]. Instead every one of the 49 ids the
old page published stays alive on the index as an authored <a id="..." />
stub in "Where each section moved". Ids were computed with
parseDocsDocument, not a slug approximation, so punctuated headings keep
both emitted forms (for example
docker-runners-(optional-%22works-in-linux%22-checks) and
docker-runners-optional-works-in-linux-checks). The index still publishes
`related` itself, so that id is deliberately not stubbed and no
duplicate authored/canonical ID is raised.

Losslessness, asserted mechanically rather than by eye: all 14 original
section bodies are character-identical after the move (0 lost, 0
changed), and the page lede is byte-identical. Word count 9,632 -> 9,632,
code fences 18 -> 18, links 13 -> 13, table rows 0 -> 0. All 751 inline
code spans and fenced blocks compare as an identical set, so every
command in the guide is unchanged. The only body delta is six trailing
newlines removed by scripts/format-docs.mts.

Verified independently of docs-link-audit, which a split makes
uninformative because it rewrites the repo's own links: the 49 pre-split
ids were enumerated from HEAD, the post-split index and every child were
re-parsed, and each id was asserted to resolve. 0 unresolved, 0 stub
targets that miss their child, 0 collisions.

Prose findings are deliberately left alone and deferred: r3-0398,
r3-0399, r3-0400, r3-0401, r3-1686, r3-1687, r3-1688, r3-1689, r3-1690.

Closes audit findings: r3-0397
2026-09-08 15:37:59 +08:00
Vincent Koc
c5af6a25ba
docs: correct 11 verified factual errors (#141214)
Each fix replaces a statement that contradicts the code, the CLI reference,
or the config schema. Authorities are quoted with file:line.

1. docs/start/setup.md:64 - bare `openclaw setup` routing
Said: "Bare `openclaw setup`, without `--baseline`, is an alias for
`openclaw onboard` and runs the full interactive wizard."
Authority: src/cli/program/register.setup.ts:23-40 `resolveSetupCommandRoute`
returns "system-agent" when `input.configured && (input.interactive ||
input.json)`, and falls to "onboarding" otherwise. docs/cli/setup.md:13-16
states the same routing. Corrected to the chat-first routing with a link to
the CLI reference.

2. docs/install/installer.md:238 - Alpine container tag
Said: "use an official `node:24-alpine` container".
Authority: scripts/install.sh:19 `NODE_DEFAULT_MAJOR=26` and
scripts/install.sh:2263 `echo "Use an official node:${NODE_DEFAULT_MAJOR}-alpine
container ..."`. docs/install/installer.md:104 already says `node:26-alpine`
for the same vulnerable-SQLite fallback. Changed 24 to 26.

3. docs/specs/codex-supervision.md:178 - `codex sessions` synopsis
The synopsis omitted `[--agent <id>]` while line 191 of the same page says all
three commands accept it.
Authority: extensions/codex/src/session-cli.ts:318-320 registers
`.command("sessions")` with `.option("--agent <id>", "Agent id that owns the
Codex sessions")`. docs/plugins/codex-supervision.md:220 shows the flag.
Added `[--agent <id>]`.

4. docs/providers/moonshot.md:38-40 - Kimi K3 thinking levels
Said: "OpenClaw exposes those exact levels and maps `/think xhigh` to `max`"
after listing the vendor's `low`, `high`, and `max` reasoning efforts.
Authority: extensions/moonshot/provider-policy-api.ts:8-12 registers
ALWAYS_THINKING_PROFILES with `[KIMI_K3_MODEL_ID]: { id: "max", label: "max" }`,
and lines 32-41 return `levels: [profile]` with `defaultLevel: profile.id`, so
only `max` is exposed. src/llm/providers/stream-wrappers/moonshot-thinking.ts:13-17
maps `"kimi-k3": "max"` and lines 62-70 overwrite the outgoing payload with
`payload.reasoning_effort = effort` after deleting `thinking` and
`reasoningEffort`. The accordion at line 320 already stated the max-only
behavior correctly and is left unchanged; the contradicting sentence at the top
of the page was corrected instead. The catalog thinkingLevelMap in
extensions/moonshot/openclaw.plugin.json is model metadata and does not make the
lower efforts usable on the direct Moonshot route.

5. docs/providers/perplexity-provider.md:82 - `count` transport
The filter table marked `count` "Native only" while line 90 says native-only
filters return an error on the chat-completions path.
Authority: extensions/perplexity/src/perplexity-web-search-provider.runtime.ts:305-316
`unsupportedOptions` covers country, language, date_after/date_before,
domain_filter, and max_tokens/max_tokens_per_page - `count` is absent, so it
does not error. Line 375 of the same file passes count only on the structured
path. docs/tools/perplexity-search.md:144-146 says count is accepted and
compatibility-only on the Sonar path. Changed the row to "Both" and noted that
the chat-completions path ignores it.

6. docs/providers/qianfan.md:76 - onboarding default model
Said the config example "explicitly selects the current DeepSeek flagship
instead of the onboarding compatibility default", but it selects
`qianfan/deepseek-v4-pro`, which is the onboarding default.
Authority: extensions/qianfan/provider-catalog.ts:7
`export const QIANFAN_DEFAULT_MODEL_ID = "deepseek-v4-pro";`, consumed at
extensions/qianfan/onboard.ts:49 as `defaultModelId`. The same page states this
at lines 17 and 44.

7. docs/help/faq.md:821 - Gateway host platforms
Said: "The Gateway runs on macOS/Linux (Windows via WSL2)".
Authority: docs/platforms/windows.md:118 "## Native Windows CLI and Gateway"
with the PowerShell install and `openclaw gateway status` runbook;
docs/help/faq.md:1319 documents "Native Windows CLI/Gateway: runs directly in
Windows". Corrected to macOS, Linux, and Windows (native or WSL2).

8. docs/help/faq.md:900 - node local tools by platform
Said: "Local node tools are currently macOS-only."
Authority: docs/nodes/index.md:772-779 gives default node command allowlists for
iOS, watchOS, Android, macOS, Windows, and Linux, including `camera.list` and
`location.get` on iOS/Android/Windows and `computer.act` on Windows and Linux.
docs/help/faq.md:719 says Macs/iOS/Android nodes expose local tools. Replaced
with a platform-dependent statement plus a link to /nodes.

9. docs/concepts/features.md:35 - channels shipped in core
Said only Telegram and WebChat ship with the core install.
Authority: package.json `files` excludes non-bundled plugins as
`!dist/extensions/<id>/**` (for example `!dist/extensions/discord/**`,
`!dist/extensions/slack/**`); `a2a`, `reef`, and `telegram` carry no such
exclusion, so they ship in the package. docs/channels/index.md:12 defines
"bundled plugin"/"included in core" as shipping with the core install, and
lines 38, 54, 59, and 62 mark A2A, Reef, Telegram, and WebChat that way.

10. docs/gateway/logging.md:142 - settable console styles
Said: "Console styles: `pretty` | `compact` | `json`."
Authority: src/config/zod-schema.root-shape.ts:117
`consoleStyle: z.union([z.literal("pretty"), z.literal("json")]).optional()`
and src/config/types.base.ts:287 `consoleStyle?: "pretty" | "json";`.
docs/gateway/logging.md:82 already says `compact` is no longer settable and is
applied automatically off a TTY.

11. docs/gateway/multiple-gateways.md:34 - rescue-profile port spacing
Said: "Use a base port at least 20 higher than the main bot", while line 20 of
the same page requires 120.
Authority: extensions/browser/src/config/port-defaults.ts:27-29
`deriveDefaultBrowserControlPort` returns `gatewayPort + 2`, and lines 32-34
`deriveDefaultBrowserCdpPortRange` returns `start = controlPort + 9`,
`end = start + DEFAULT_BROWSER_CDP_PORT_RANGE_SPAN` where the span is
18899 - 18800 = 99. Each Gateway therefore reaches base + 110, so 20-port
spacing overlaps. Changed 20 to 120.

Closes audit findings: r3-0772, r3-0428, r3-0754, r3-0671, r5-0045, r3-0682, r3-0393, r3-0394, r3-0278, r3-0386, r5-0004
2026-09-08 15:35:38 +08:00
Ayaan Zaidi
f310ecc641
docs(models): explain route readiness and catalog refresh results (#141927)
## What Problem This Solves

Operators can see a stored credential with `static` health while `models status --check` exits 1 because route readiness is unknown. The reference currently omits that result. The FAQ also recommends unscoped status immediately after agent-specific login guidance, which can inspect a different model selection.

## Why This Change Was Made

The model CLI reference now distinguishes inventory, credential sources, profile health, model-route issues, and runtime availability. It documents check exit codes and command-resolution failures, and links chat-session inspection to the existing `/model status` guidance. The FAQ uses an explicit agent and links to this canonical explanation.

Hosted refresh now has its own section, preserving the distinction between downloading metadata and activating it after a Gateway restart. The FAQ's obsolete numeric picker was already corrected by #139232 and remains unchanged.

## User Impact

Operators can inspect the intended agent, understand an indeterminate result without treating it as rejected credentials, and distinguish a catalog download from authentication or live activation. No product behavior, model-selection policy, auth policy, or configuration changes.

## Evidence

- Source audit at `8e47d6998a` covers CLI registration, status output and exit-code selection, credential-source/health owners, and hosted refresh.
- Reused the actual isolated `models status --agent work --json --check` observation at `4e402dda9e0ee8c8a5b01e39d7ccd61ccf6d40ad`: stored exec reference retained, profile `static`, route `indeterminate`, exit 1, and no execution of the stored reference. The relevant status/health/overview/config-loading source is unchanged at the documentation baseline.
- Reused the completed hosted refresh proof from #141885: a separate CLI download produces a Gateway restart notice without activating rows or prices. No new runtime, provider request, or credential access was needed for this docs-only change.
- `pnpm docs:list`, changed-file `oxfmt --check`, `node scripts/check-docs-mdx.mjs docs/cli/models.md docs/help/faq-models.md`, the glossary check against the pinned baseline, and `git diff --check` pass.
- Full link/anchor audits check 9,427 candidate links and 9,424 baseline links. Both report the same 39 pre-existing missing ClawHub mirror links; all three new links pass. The local external documentation mirror is unavailable. No product build or typecheck ran locally.

Co-authored-by: Ayaan Zaidi <hi@obviy.us>
2026-09-08 12:06:19 +05:30
Peter Steinberger
757e482ab6
feat(sqlite): select an extension-capable SQLite library for Bun on macOS (#141854)
* feat(sqlite): select an extension-capable SQLite library for Bun on macOS

Bun on macOS dlopens Apple's SQLite, which omits extension loading, so
sqlite-vec never loaded and memory search fell back to the batched scan.
Bun's only hook is bun:sqlite Database.setCustomSQLite: one-shot, before
the first open, and fatal to every later open on a bad path.

- Add src/infra/bun-sqlite-library.ts: validate each candidate through
  bun:ffi (WAL-reset-safe version, extension loading present) before
  committing; resolve explicit path > OPENCLAW_SQLITE_LIBRARY >
  Homebrew/MacPorts discovery; memoize process-wide. No-op on Node and
  on Linux/Windows Bun, whose static SQLite already loads extensions.
- Select before the first open in the runtime guard probe and in
  requireNodeSqlite; an unusable override becomes a clean runtime-guard
  diagnostic with exit 1 instead of a stack trace.
- Forward the selected library to the memory KNN child through its typed
  stdin input and select there before opening; the child env stays
  stripped. The #141104 scan fallback remains the no-library path.
- Report the selection at Gateway startup and in doctor.
- Retire the undocumented legacy OPENCLAW_CLAUDE_CLI_LOG_OUTPUT alias so
  the OPENCLAW_* name budget stays at 493.
- Docs: Bun install caveats, environment reference, memory config.

* test(sqlite): retain the KNN stdin spy for payload assertions

* docs(bun): explain migration from custom SQLite preloads
2026-09-07 22:53:04 -07:00
Peter Steinberger
a6de2a2192
fix(update): restore missing official plugins at the selected release (#141478)
* fix(update): restore untracked official plugins from the core release cohort

Keep update admission and Doctor repair on the selected stable core version when formerly bundled official plugins have no install record. Preserve existing selectors and capability policy. Add DuckDuckGo published-upgrade and standalone Doctor coverage for bundled and already-missing states.

Related: #135002, #116740

* test(update): recognize missing optional plugin warnings

* test(update): isolate plugin repair proof and satisfy helper lint

* test(update): preserve beta tag selection in doctor fixture

* chore(ci): register plugin doctor survivor entrypoint
2026-09-07 13:08:35 -07:00
Peter Steinberger
0d0e2852b2
fix(release): restore verification inputs and supported upgrade proof (#140981)
* fix(qa): verify leased Telegram tester group access

Check tester membership and effective text permission in the leased user driver before group readiness. Reuse the check in doctor and select trusted skill scripts for frozen release candidate QA. Preserve supported private DM turns and leave credential repair to the pool owner.

* ci: install Chromium for native live browser tests

The native-live-test shard includes the real Gateway widget restart proof,
which launches Playwright after restoring the interrupted session. Chromium
was installed only in separate Repo E2E jobs, leaving the live shard without
its required executable.

Install the candidate UI's pinned Chromium and system dependencies only for
the selected native-live-test row. Preserve profile selection and all test
gates. Verified the unchanged dc420 widget restart test with real OpenAI,
Gateway process replacement, and interactive Chromium dashboard assertions
on Blacksmith Testbox; canonical workflow checks and formatting also pass.

* test(release): prove supported cross-OS 9.2 transition

Select the explicitly supported external package-manager and fresh Doctor
transition only for 2026.9.2 to 2026.9.3. Preserve the old updater's safe
schema-15 refusal as negative evidence and never report self-update passed.
Verify backup, exact package identity, schema migration, retained settings
and session history, candidate serving version, and a fresh persisted turn.
Keep credential-bearing backup archives out of CI evidence and remove them
on both successful and failed runs. Other release pairs retain existing
updater and Windows fallback behavior.

Validation: 13 cross-OS upgrade lane tests, two pre-fix regression failures,
full changed checks, and independent review. Real packaged cross-OS/live
qualification remains required against the aggregate candidate.

* test(installer): prove supported historical package transitions

* test(release): exercise same-schema updater boundaries

Keep updater-specific plugin repair and consent checks on the prepared
candidate, then target explicit synthetic later versions with the same
runtime and storage schema. Retain installed-version and plugin-policy
assertions, and remove the obsolete legacy post-update fallback.

Add versioned future-tarball generation to the existing first-hop fixture
owner, with immutable input and source/target digest receipts. Restore the
reserved runtime-promotion staging ignore in package-derived Git fixtures
so the updater's unchanged-source check stays meaningful.

Validation: 26 focused tests, a pre-fix Git staging regression, full changed
checks, exact candidate tarball fixture generation, and P2 review. Actual
Docker update/consent/channel qualification remains required in aggregate CI.

* test(release): verify channel update staging cleanup

Assert the real channel flow refuses ordinary untracked user files without changing source, then require successful channel updates to remove every reserved runtime staging entry. Preserve recovery data when the assertion detects leftovers. No ignore or production cleanup policy is broadened.

* test(onboard): select configured model through explicit picker

* ci: pin release performance checks to reviewed Kova accounting

Select the reviewed Linux process-lifetime CPU accounting and bounded worker-watchdog implementation from Kova PR #110 for canonical, legacy-list, and trusted live fixtures. Keep existing trust classification, timeout values, and performance thresholds.

* fix(ci): distinguish bundle build modes in job names

* test(release): build matching future runtime fixtures

Derive a synthetic Codex runtime cohort from its verified candidate package by changing only package.version and openclaw.build.openclawVersion. Require matching source metadata and preserve payload, constraints, immutable input, and source/target digest receipts. Reuse the existing private tarball lifecycle and sequence validation. Validated 11 focused tests, pre-change CLI failures, full changed checks, and P2 review.

* docs(release): prepare private QA before local E2E

* fix(release): admit package metadata before validation dispatch

Port the reviewed early source admission into trusted Tooling, preserving the current selected REST transport. Reject invalid version notes and misaligned core packages before any Git push, non-GET API write, or workflow dispatch; retain explicitly allowed substantive draft notes.

* test(release): qualify supported historical upgrade transitions

Prove the published 2026.8.2/2026.9.2 to 2026.9.3 external package-manager and fresh Doctor contract through one shared helper. Preserve the old in-process updater refusal and schema 15, require owner-stopped migration with a verified private backup, and verify schema 16 plus retained Gateway history afterward.

Keep current-runtime update coverage distinct: managed restart uses separately identified future core and matching runtime fixtures, with real updater, canary, service replacement, authentication, and serving checks. Fix fixture-only channel environment pollution and mock serving nonce responses.

Sealed Docker journey, historical first-hop, current-candidate survivor, root-managed VPS, and onboarding pass. Managed-auth historical preservation passes; its future canary exposes an independent existing plugin projection blocker, retained as failing proof for the candidate repair owner. No product schema or refusal guard changes.

* test(release): narrow recorded update command before indexing

* fix(ci): align historical upgrade fixture ownership

Keep all seven external transition assertion cases with the first-hop package fixtures: both own the historical upgrade boundary and now share its temporary-state lifecycle. Preserve helper test selection and all assertions while avoiding an additional standalone tooling group.

Provide the new inference and future-package preparation boundaries in isolated service, mobile, and cron bootstrap probes, and verify each updater argument without changing readiness or migration assertions. Register the path-launched assertion CLI as a Knip entry.

Validation: 67 related tests; 70 exact-merge planner and combined-suite tests; five 80-job stress plans; full Knip unused-file scans; scoped checks and P2 review.

* fix(ci): model upgrade preparation in isolated probes

Complete the companion fixture and Knip repairs for the preceding transition-test consolidation. Model the independent inference and future-package preparation producers, verify exact updater argument boundaries, and retain service-readiness, authored-state, and cron migration assertions. Declare the path-launched transition assertion CLI in the existing Knip entrypoint list.

* test(release): preserve fixture registry on managed restart

Capture the fake service manager registry alongside its endpoint paths, then refresh that owned context after preparing the future cohort. Projected native service callers cannot drop or replace it; transient update state still stays outside the managed child.

The executable service regression fails on the original shim and passes with missing and conflicting caller registries. Four focused tests, the changed gate, and P2 review pass. Reuse the existing d928 package for the final managed-upgrade replay.

* test(release): preserve prepared UI identity in future fixtures

Advance synthetic package versions while retaining the opaque build ID already embedded in unchanged Control UI assets. The canonical asset-health regression fails the prior builder as stale, then verifies readiness and unchanged UI bytes for both future fixture sequences. Package-version update admission and managed-restart proof remain unchanged.

* test(release): retain manager-owned fixture policy

Capture the existing automation and offline-channel policy with registry identity when generating the manager shim. Native service clients project these values away; losing them activated retained synthetic channel credentials and correctly failed readiness. Restore only this explicit manager allowlist and keep transient update/compatibility state excluded.

The executable service regression fails on missing manager flags and passes across projected or conflicting caller environments. Four focused tests, the full changed gate, and fresh P2 review pass. Actual corrected managed Docker proof remains pending.

* docs(ci): keep bundle identity with routing guidance

* fix(release): retry transient GraphQL EOF failures

Classify the observed gh unexpected EOF transport failure within the existing five-attempt retry budget. Preserve authentication, invalid-response, inventory and snapshot handling. Exercise recovery, exhaustion and fail-fast errors through the real verifier CLI with a fake gh transport; also modernize the existing disposable-array sort required by touched-test lint.

* test(release): align survivor fixtures with manager preparation

Supply manager environment JSON to directly launched supervisors, and model the independent manager preparation phase in config-parking and companion fixtures. Preserve service readiness, restart, diagnostics, and strict phase-order assertions without changing production behavior or test limits.

* test(release): align upgrades with deferred schema publication

Restore installed-updater survivor and first-hop coverage after shared migration ownership moved into the durable update ledger. Record applied schema content separately from published schema version, retain the genuine missing-chunk negative control, and label external installation as an alternate with no self-update attempt. Preserve production-owned unsupported migration refusals.

* test(release): restore supported packaged upgrade coverage

Restore real 2026.9.2 to 2026.9.3 self-update after ledger-driven schema migration support. Keep retained conversations, private backup, settings, candidate identity, and serving inference proof. Require migrated content independently of publication grace and remove obsolete external-install diversion from cross-OS and installer update smoke. Shared external helper remains separately owned.

* test(release): refresh legacy manager before updater restart

Bind the service manager to the candidate registry before the published baseline updater starts its managed service. Preserve restart verification for successful and recoverable outcomes alongside explicit future targets. Exercise stale registry capture, replacement failures, pre-update identity, and attribution without changing timeouts or caps.
2026-09-07 11:47:27 -07:00
Peter Steinberger
ce0e84d073
fix(runtime): require Node builds with lossless SQLite reads (#140672)
* fix(runtime): require Node builds with lossless SQLite reads

* fix(runtime): preserve upgrades and guard sealed workers

Validate downloaded Node before switching the active runtime alias, reject unsupported sealed-worker runtimes, and keep the Gateway error fixture on a supported Node release. Document the approved ARMv7 and older macOS compatibility losses and decoder fix boundaries.

* test(runtime): use typed process exports in worker fixture

* test(runtime): align installer fixtures without growing test shards

* test(runtime): align release and guest runtime fixtures

* fix(test): canonicalize Windows temp roots for Node 24

Expand Windows short paths before creating test directories and owned child environments. Node 24 filesystem watchers otherwise abort when native long event paths differ from inherited short temporary paths. Preserve explicit custom-root spelling and existing cleanup ownership.

* test(ci): run Windows temp-root regressions in the native lane
2026-09-07 10:31:31 -07:00
Peter Steinberger
c24b3f2e1a
fix: catch published upgrade regressions before merging (#141146)
* ci: gate published upgrades with legacy operator state

Add a latest-release PR/main tripwire and runtime-resolved weekly and release baselines. Exercise baseline-authored agents, approvals, cron owners, a plugin, and mock agent turns; preserve the narrow unfenced-updater refusal contract.

Related: #140778, #140784, #140886, #140825.

* test: verify native published upgrade lifecycle

Persist JSON-era CLI approval IDs, retain a moving plugin selector, and exercise the baseline updater-owned restart. Query cron owners before repair, reject malformed canonical approvals, and require clean Doctor evidence.

* test: keep survivor installs in disposable runtime state

Retain verification logs and hashes while letting the container own the legacy scenario npm prefix. Other survivor scenarios keep their existing artifact layout.

* test: preserve immutable sorting in CI docs guard

* fix: preserve historical release upgrade qualification

* test: handle absent release scenario output in assertions

* ci: preserve survivor coverage within existing limits

Keep all native operator and schema cases in their existing owner suites. Register the shell observer entry and current inert scenario catalog, and update fixture, request, and routing contracts without changing runner caps or assertions.

* ci: pair expanded upgrade baselines with native operator state

* test: classify upgrade baseline provenance as internal

* ci: preserve docs-only main push exclusions

* test: require successful upgrades from every supported baseline

Remove the temporary 2026.9.2 schema-refusal pass after #141109. Observe applied shared-state content while publication is deferred, retaining candidate, state-preservation, restart, and validated plugin-consent recovery checks.

* test: require legacy agent migration before upgrade probes

Snapshot the seeded roster and legacy session specimens independently of existing SQLite files. Require migrated agent stores, imported history, and preserved recovery artifacts before candidate probes can initialize state.

* test: scope upgrade specimens to session migration storage

Ignore native model catalogs and unrelated per-agent artifacts while preserving strict classification inside the legacy sessions directory and the existing SQLite runtime inventory.
2026-09-07 09:44:31 -07:00
Vincent Koc
29ba80f738
docs(gateway): split the agents configuration reference by domain (#141149)
The agents configuration reference was 1,586 lines / 94,062 characters with
47 headings, and one H2 (`Agent defaults`) held 7,858 of its 11,143 words.
That is roughly five times the 20k split threshold. It is now a short index
plus eight domain pages under `docs/gateway/config-agents/`, matching the
`docs/web/control-ui` split and reusing the wording of the
`configuration-reference` split.

Children (all under the 20k threshold):

- `workspace-and-bootstrap` — workspace, cwd, repoRoot, skills, bootstrap
  injection, the context budget map, image handling, timezone (11,163 chars)
- `models` — `agents.defaults.model`, fallbacks, media model slots,
  `modelSelectionScope` (18,517 chars)
- `runtime-and-cli-backends` — runtime policy, CLI backend selection, GPT-5
  personality (5,172 chars)
- `heartbeat-compaction-and-streaming` — heartbeat, system agent, compaction,
  context pruning, block streaming, typing indicators (15,129 chars)
- `sandbox` — the `agents.defaults.sandbox` block (10,718 chars)
- `entries-and-multi-agent` — `agents.entries`, `multiAgent` bindings, match
  fields, access profiles (10,826 chars)
- `sessions` — `session.*` (9,869 chars)
- `messages-and-talk` — `messages.*` and `talk.*` (11,964 chars)

Anchors. `redirectSource()` in `scripts/lib/docs-redirects.mjs` rejects any
source containing `[?#]`, so a fragment can never reach a path redirect and no
redirect is added; the parent keeps its route at `/gateway/config-agents`.
Every anchor the single page published stays resolvable on the index instead.
All 82 ids were enumerated with `parseDocsDocument`, not a hand-rolled slug:
81 are retained as authored `<a id>` stubs in a "Where each section moved"
list, and `related` stays a published heading, so no id is both stubbed and
published and the parser reports zero collisions. Each stub links to the page
that now owns the heading, and each of those 63 in-set links was asserted to
resolve against the child's own parsed id set. 25 anchored in-repo references
across 18 files are repointed at the owning child; the 18 links to
`#agent-defaults` keep pointing at the index, which is now the map of the five
pages that H2 became.

Losslessness. Each child body is byte-identical to its original line range,
with two exceptions: 34 headings are raised one level so each child has a
top-level structure, and one cross-reference the split orphaned
(`agents.entries.cwd` -> "Working directory") is repointed from
`/gateway/config-agents#agents.defaults.cwd` to the workspace page. No prose
was rewritten. Fenced blocks 80 before / 80 after. Table rows 29 before /
29 after, unchanged per table: context budget ownership map 6, model selection
scope 6, runtime policy 10, response prefix 7.

`src/config/talk-defaults.test.ts` hard-coded `docs/gateway/config-agents.md`
as the Talk defaults parity source and now reads
`docs/gateway/config-agents/messages-and-talk.md`. The eight children are
registered in `docs/docs.json` as a nested group, mirroring Control UI.
zh-CN glossary entries were added for the eight new titles; those translations
are unreviewed.

Closes audit findings: r3-0297, r3-0299
2026-09-07 11:03:47 +00:00
Peter Steinberger
233dabe750
docs: state who runs OpenClaw, what it sends home, and how it is funded (#141053) 2026-09-07 02:25:57 -07:00
Peter Steinberger
34975ff95b
fix: live test suites pass green when provider credentials are missing (#141024)
The github-copilot connection-bound-ids live suite returned green when
OPENCLAW_LIVE_TEST=1 was set but no Copilot token could be resolved from env or
the auth profile, so a Testbox run could report a passing live test that never
contacted the provider. The music and video generation live sweeps had the same
shape: an unfiltered run with zero attempts warned and passed.

Use the Vitest test-context skip(reason) in the Copilot test, and thread the
context skip through the music/video sweep summary helpers so unfiltered
zero-attempt runs show as skipped in the reporter with the skipped entries.
Filtered zero-attempt runs still throw; a present-but-invalid token still fails.
Unit coverage asserts the skip callback contract for both sweep helpers, and
docs/help/testing-live.md records the rule that a live suite without
credentials must skip visibly or fail, never pass green.
2026-09-07 01:42:27 -07:00
Vincent Koc
cc358246f6
docs(gateway): split the configuration reference by domain (#140440)
The configuration reference was 2,077 lines / 146,192 characters with 55
headings, roughly seven times the 20k split threshold and well past the
40-heading split signal. It is now a short index parent plus nine
domain pages. No prose was rewritten: every moved H2 section body is
byte-identical to its source.

Children (all moved verbatim from configuration-reference.md):

- gateway/config-runtime (1,117 words): worktreeRoot, Models, Discovery,
  Update, ACP, Wizard, Bridge (legacy, removed)
- gateway/config-extensions (2,221 words): MCP, Skills, Plugins,
  Canvas widget presenter
- gateway/config-browser-ui-desktop (2,012 words): Browser, UI, Desktop
- gateway/config-gateway (3,600 words): Gateway, incl. OpenAI-compatible
  endpoints, multi-instance isolation, gateway.tls, gateway.reload
- gateway/config-cloud-workers (1,495 words): Cloud worker environments
- gateway/config-hooks (3,273 words): Hooks, incl. HTTP contract, agent
  payload, session policy, mapping, retries and fan-out, Gmail
- gateway/config-secrets-env (1,042 words): Environment, Secrets,
  Auth storage, Config includes ($include)
- gateway/config-observability (1,190 words): Audit, Logging,
  Diagnostics, Telemetry
- gateway/config-automation (958 words): Automations (cron), Media model
  template variables

Anchor preservation. docs.json redirects match on pathname only - 0 of
the 281 existing redirects carry a fragment in `source`, and fragments
never reach a redirect matcher - so no path redirect is added and the
parent keeps its route. Instead every anchor stays resolvable on
/gateway/configuration-reference: all 32 original H2 headings remain as
one-line pointer sections, and all 23 original H3 anchors are retained
as authored <a id> stubs (the pattern already used in
docs/help/faq-first-run.md). All 46 anchored in-repo references across
27 files are rewritten to the child page that now owns the heading;
docs-link-audit --anchors reports 8,717 links checked, 0 broken.

Losslessness (parent + nine children vs. the old single file):

- code fences: 40 -> 40
- MDX components: 4 -> 4
- distinct config keys documented: 490 -> 511 (0 lost, 21 gained from
  the new index text)
- words: 16,756 -> 17,542 (+786, all new index and lede text)
- headings: 55 -> 83 (55 original + 27 pointer H2 + 1 index H2)
- parent page: 2,077 -> 226 lines, 146,192 -> 9,701 chars, 55 -> 33
  headings

src/docs/cloud-workers-config.test.ts pinned the cloudWorkers examples
to configuration-reference.md by path; it now names
config-cloud-workers.md. Twelve zh-CN glossary entries were added for
the new and existing "Configuration - <domain>" page titles.

Closes audit findings: r3-0286, r3-1510
2026-09-07 05:18:18 +08:00
Vincent Koc
11f454ee82
docs(start): repair the first-run path from landing page to first channel (#140390)
Walks the new reader path end to end and fixes each step where the docs
sent the reader somewhere the previous step had not prepared.

Landing page (docs/index.md): the Quick start installed with npm and then
ran `openclaw onboard --install-daemon`, which selects the classic wizard
(onboarding-overview.md), so the landing reader never saw the Quick start
and Custom setup choice that Getting Started narrates. It now uses the
installer script that install/index.md calls "Recommended", names which
wizard that opens, and adds the missing Gateway service step. The mobile
hub card "Get started" pointed at `/`, and the "Channels" card pointed at
one channel page while promising the catalog; both now point at the hub
they describe.

Install page (docs/install/index.md): "Verify the install" ended the page
with no route onward. Adds a next-step card group to Getting started and
to the channel hub. docs/install/node.md pointed "installer script" at the
alternative-methods anchor instead of the installer script section.

Getting Started (docs/start/getting-started.md): Step 2 left the Gateway
in the foreground and buried "Ctrl+C, then `openclaw gateway install`" in
prose, while Steps 3 and 4 assumed a running background Gateway. That is
now its own numbered step between onboarding and verification.

Channel hub (docs/channels/index.md): the "fastest setup is usually
Telegram" guidance sat at the bottom, below the catalog and a long group
introductions explanation, and the page carried no `openclaw channels add`
example. Both now sit above the 31-entry catalog.

Telegram (docs/channels/telegram.md): quick setup showed a JSON5 block
with no file path and no `openclaw channels add`, told the reader to run
`openclaw pairing list` without first sending the bot a message, and
started a second foreground Gateway that conflicts with the service the
Getting Started path installs. Step 2 now names `~/.openclaw/openclaw.json`
and leads with `openclaw channels add --channel telegram --token <token>`.
Restart and pairing are separate steps, the restart uses
`openclaw gateway restart` with `openclaw gateway` named as the
no-service case, and pairing starts by messaging the bot. The opening
line now says what the page is for.

macOS onboarding (docs/start/onboarding.md): the first three steps were
images with empty alt text. Each now says what dialog appears and which
button to click, and read_when addresses first-run readers rather than
the people implementing the flow.

First-run FAQ (docs/help/faq-first-run.md): the install answer ran the
installer and then onboarding again, the "what does onboarding do" answer
described the classic 8-step wizard as the default, install and
onboarding answers sat below heartbeat and exec-approval answers, the
provider-add command disagreed with the wizard pages, and the "I am
stuck" link landed on a heading rather than the accordion.

Also trims duplicated Quick start prose from onboarding-overview.md,
names the WhatsApp plugin prerequisite in openclaw.md, repairs the
localhost dashboard link and the install entries in hubs.md, cuts the
19-item "Start here" list in docs-directory.md, and splits the
single-instruction sentences named by the STE findings in telegram.md,
why-openclaw.md, teams.md, setup.md, and wizard-cli-automation.md.

Closes audit findings: r3-0936, r3-0937, r3-0938, r3-0939, r3-0940,
r3-0941, r3-0942, r3-0943, r3-0944, r3-0945, r3-0946, r3-0947, r3-0948,
r3-0949, r3-0950, r3-0951, r3-0952, r3-0953, r3-0954, r3-0955, r3-0956,
r3-0958, r3-0959, r3-0962, r3-0963, r3-0964, r3-0965
2026-09-07 04:01:34 +08:00
Vincent Koc
e8ffd3cf09
docs: replace private paths and document the ClawHub docs source (#140227)
Docs governance and publish-hygiene pass over docs/AGENTS.md and the
pages its rules cover.

- Replace every `~/Projects` operator path in docs/ with a neutral
  placeholder. 38 occurrences across 13 pages, including the private
  repo path `~/Projects/manager/skills`. `docs/AGENTS.md` forbids local
  paths, and its own Internal Docs bullet named one.
- Record the placeholder convention in the Published Link Rules bullet
  that bans local paths.
- Document the ClawHub docs source in Source Ownership: this repo holds
  no `/clawhub/**` page sources even though `docs/docs.json` lists them,
  so a local preview and `pnpm docs:check-links` report those routes as
  missing until `OPENCLAW_DOCS_SYNC_CLAWHUB_REPO` points at a ClawHub
  checkout.
- Drop the Strict-STE hard violation rate on `docs/AGENTS.md` from 18
  hard (3.4 per 100 words) to 0 by splitting semicolon sentences and
  sentences over 20 words, and by naming the actor in passive
  sentences. No rule changes meaning.
- Convert the Maturity Scorecard paragraph to a bulleted list, matching
  every other section.
- `docs/prose.md`: name v2026.8.1 as the release that removed OpenProse,
  and explain the `--agent codex` flag and the third-party skills CLI.
- Link `/prose` from `tools/skills` and `tools/slash-commands`.
- `docs/docs_map.md`: correct the summary to describe the stub, and drop
  the H1 that repeated the frontmatter title.

Closes audit findings: r3-0734, r3-0736, r3-0907, r3-0910, r3-2277,
r3-2278, r3-2285, r3-2286, r3-2287, r3-2288, r4-clawhub-0001

Partially addresses r3-0906 (private path removed; publish-tree
exclusion left as a follow-up). Not addressed: r3-0909 (generator
change).
2026-09-06 23:48:33 +08:00
Ayaan Zaidi
341f459500
fix(skills): simplify Workshop reviews and skill previews (#139760)
## What Problem This Solves

Fixes an issue where Skill Workshop revisions could be sent to an unrelated current chat, skill previews omitted valid Markdown or important procedure steps, and collection reviews lost inherited model-runtime settings when an agent had an empty model map.

The Workshop also carried a separate historical-review completion path used only by standalone tests, even though production already had one durable completion owner.

## Why This Change Was Made

- Remove the current-chat switch and its preference/state plumbing. Revision requests use an eligible original conversation or the existing new-conversation path, with ownership, stale-revision, and recovery checks preserved.
- Use the shared sanitized document renderer in Board and Today. Today exposes the complete skill through a native disclosure instead of guessing English section names and truncating steps. Shared styles contain images, tables, and code; external images stay click-to-open.
- Remove dead new/seen state, an unused revision callback override, duplicated prose styles, and unreachable empty-state code. Applied-history links use the canonical grouped projection and open the correct filter.
- Require prepared historical-review inputs and explicit completion. The existing batch owner handles durable completion and late failures; authoring rules come from the Workshop tool once.
- Preserve global model metadata during cron preparation. The canonical request-parameter resolver now composes global defaults, shared model parameters, per-agent model parameters, then agent-wide parameters. Thinking and fast-mode controls use the same per-agent model scope, including the final cron execution and fallback path. No cron-only parameter merge is retained.

No SQLite schema, new configuration, permissions, transcript-selection limits, or automatic-apply policy changes. Historical proposal records are retained. Separating current skills from proposal history is a follow-up, not part of this PR.

## User Impact

Users get complete, safely rendered skill documents and one predictable revision route. Background collection reviews keep the configured model runtime, including when the agent's model map is empty. Failed or interrupted historical reviews retain the existing recovery behavior.

Production delta: +215 / -704, net **-489 lines**. Tests: +482 / -207, net +275 lines. Docs/config: +6 / -6. Overall: -214 lines across 51 files.

## Evidence

- **177 focused tests passed on commit 9da2656b6e02d9d353f233dcfac2a31e51c84015**, covering Workshop service and Gateway handlers, revision ownership/recovery, historical scan completion, cron settings, shared Markdown documents, page actions, and Chromium layout.
- Captured failing regressions before fixes: wrong-chat routing from the saved preference, omitted unordered lists, incorrect applied-history navigation, remote-image loading, wide-image overflow, omitted Today sections/steps, and lost inherited cron runtime/parameters.
- A real collection automation initially failed with the wrong effective runtime. With the cron fix built and deployed, the same automation completed on the configured OpenAI model using the embedded runtime. It reviewed five skills and edited three; filesystem comparison confirmed no additions or removals.
- That live runtime was built from 6b3d88538e plus the identical cron patch. It proves the cron repair, not deployment of all other refactors. The final committed owner tests cover the combined candidate.
- Additional **360 parameter/sibling tests passed on 79f3292**, followed by **59 focused tests on the source committed as ef534a5**. The final regression failed without the executor repair: a per-agent fallback expected thinking off but received the shared high setting.
- **Served-browser proof with a real Gateway and provider passed.** Baseline with the saved current-chat preference sent a skill revision into an unrelated conversation. The candidate created a Workshop conversation, reused its eligible origin for a second revision, persisted proposal versions 2 and 3, and left the unrelated conversation byte-identical. Full mixed-language headings, all steps, bullets, code and table content remained readable. Apply followed by View history opened the exact applied version.
- That browser proof used candidate UI source 9da2656 with the retained Gateway described above. Later UI edits only reorganize empty-state returns and do not change these nonempty flows. A dev-build update banner remained visible; the browser lacked a Japanese font, although stored and accessibility text remained exact. These limits are retained with the complete private captures, not hidden.
- **Exact-head built cron proof on ef534a5** exercises the public cron command through a real Gateway and a local HTTP provider fixture. It checks the actual outgoing sampling parameters, cross-alias token-limit precedence, per-agent thinking, and fast-mode metadata. This is protocol-fixture proof, not a claim about a real provider model's capabilities.
- Direct lint, formatting, styles, English catalog baseline, and whitespace checks passed. The first ClawSweeper findings prompted the populated-parameter repair and complete served-browser proof. Fresh exact-head ClawSweeper and CI remain the landing gates.

## Review Follow-through

- The requested current-chat preference retirement is intentional and owner-authorized. Both saved values now use the same safe routing; no stored history is deleted.
- Both rank-up moves are addressed: populated model parameters have failing/passing regression and outgoing-request proof, and complete sanitized browser captures are available to the local reviewer. Raw provider transcripts and private fixture metadata are not uploaded publicly.
- Existing Today → Applied History → Today selection/loading failure is recorded for the separately authorized collection-navigation follow-up, which removes that dual-view state.
- Existing session-list thinking metadata can display the shared default even when a cron request correctly uses a per-agent model setting. The wire proof distinguishes that pre-existing display gap from actual execution; no accurate-display claim is made here.

AI-assisted.

Co-authored-by: Ayaan Zaidi <hi@obviy.us>
2026-09-06 11:33:56 +05:30