* test: add worker state upgrade survivor cells
Add opt-in Projects Doctor and terminal task/flow restoration cells using the unchanged published updater, verified package bytes, and the canonical survivor lifecycle. Preserve failed synthetic state until the outer Docker owner has joined.
Validation: 161 focused tests, selected changed checks, and P2 review. Actual package upgrade cells remain separately qualified against frozen candidate artifacts.
* test: activate taskflow survivor fixture on Gateway startup
* test: clean worker runtime with container ownership
* fix(test): preserve Docker status through exit traps
Publish bounded, host-redacted receipts after successful published-baseline
upgrade runs. Reuse the diagnostic owner and retain the initial post-core
result separately from repair and recovery output. Preserve Docker outcomes
and existing failure reports.
* test(docker): check the configurable ClawHub sweep spec
* fix(e2e): preserve live ClawHub package default
* fix(plugins): preserve live ClawHub package metadata
* docs(testing): document live ClawHub plugin opt-in
* chore: refresh ClawHub PR CI
Refresh the PR projection and checks against main with the wizard recovery test routing fix.
---------
Co-authored-by: Vincent Koc <vincentkoc@ieee.org>
Co-authored-by: Peter Steinberger <steipete@gmail.com>
* fix: enforce outbound provider policy without a current target
Preserve correlated recovery source providers independently of complete delivery authority. Keep cross-provider defaults and explicit overrides unchanged across restart and requester-settle continuations.
* test: preserve configured cross-provider opt-ins during recovery
* test: cover provider policy on explicit source reply routes
The explicit WebChat route fixture depended on the missing-target policy gap. Verify both the denial without opt-in and the unchanged normal outbound route when explicitly allowed.
* fix: preserve unbound CLI messaging context
* test: distinguish unbound and WebChat agent requests
* perf(test): admit isolated Gateway readers to the existing worker pool
* test: align Gateway scheduling and sandbox cache expectations
* test(ui): retain safe Gateway failure state in CI logs
A systemd Gateway whose unit lists a managed environment key and whose config uses `${VAR}` shorthand can start again after the 2026.9.4 service reinstall. Startup, service planning, and secrets audit share the existing config-resolution provenance to recognize authored environment references after substitution; escaped literals remain literal, stale keys are still removed, and configuration files remain unchanged.
Fixes#146612
Reported by @foxsky (#146612).
Validation: focused startup, installer, Doctor, config, audit, and runtime suites passed; built Linux before/after startup proof retained the key and reached readiness with unchanged config bytes. Exact-head CI was reverified green. A complete published-updater service-regeneration run remains a documented proof gap.
The hosted job-cap regression rebuilt the growing repository's full compact
plan twelve times, resetting the module graph for every inventory variation.
Its discovery cache avoided repeat enumeration but left import, ownership,
and packing work unbounded; the case timed out at 120 seconds in hosted CI.
Use fixed timing inputs and a bounded synthetic inventory through the real
planner. Fifty-six full-budget anchors leave twenty-four tooling jobs at the
80-job cap, including a stronger-runner tail. Retain all twelve variants and
every original coverage, ownership, runner, and execution-budget assertion.
Document the fixture contract and assert exact baseline capacity/membership.
Measured case runtime fell from 8.858 seconds to 0.575 seconds locally. The
complete owner file passes all 61 tests; the unmocked repository planner still
produces 78 jobs within the cap. No runtime or workflow policy changes.
* test: cover transcript cold storage in release Docker lanes
* test: remove stale cloud picker test seams
Clear baseline CI failures introduced by #145183: reuse the existing picker fixture, test cloud configuration through its production renderer, and remove the unused Connect menu renderer.
* test: preserve historical transcript fixture modification times
Doctor imports filesystem mutation time rather than timestamps inside messages. Date the legacy files before import, and handle CLI failure diagnostics without adding a lint suppression or retaining credential-bearing command arguments.
* fix: keep cold sessions available to Activity title probes
Handle cold transcripts at the optional title-reader boundary without restoring payloads or caching missing previews. Verify real archival, mixed hot/cold listings and cache recovery; extend packaged release proof across restart and portable backup recovery. Preserve JSON CLI errors from stdout in the test harness.
* fix(ui): use the action cursor for session details
Clear the existing cursor-policy failure from #145183 without weakening its regression test.
* test: preserve frozen release targets in cold-storage lanes
Share capability resolution with source preflight and omit only unsupported cold subcases under the existing explicit frozen-target policy. Keep current coverage required, isolate ordinary retention in the live fixture, and strengthen fixture typing without suppressions.
* test: complete cold-storage release entrypoint contracts
Register the shell entrypoints for dependency analysis, type the frozen-source fixture cases explicitly, and align existing cloud test callback ordering with the same repair now on main.
* test: isolate archived recall from external tool policy
* test: retain the cloud machine locator for proof capture
Require explicit Copilot configuration, a saved profile, or COPILOT_GITHUB_TOKEN within the existing agent scope. Generic GitHub credentials continue to serve other tools and remain covered by secret audit and cleanup.
Retire the discovery opt-out through the shared Doctor migration and add one-time upgrade guidance. Update provider documentation and generated configuration inventories.
Validated with focused plugin and secret tests, real Gateway starts across all 16 sanitized configs, and independent CLI checks for activation, Doctor notices, secret detection, and matching-value cleanup.
Related: #144726
Co-authored-by: Ayaan Zaidi <hi@obviy.us>
* perf(plugins): reuse compiled Doctor contracts consistently
* refactor(plugins): separate runtime bindings from artifact selection
Keep Doctor filesystem selection independent of runtime record types, migrate every binding caller, and verify repaired persistence through the built CLI instead of a source-mode observer.
Mis-pointed targets where a correct target exists, first mentions of
documented surfaces that were never linked, and pages with no Related
section at all. Adds 57 internal links, none broken.
These sat in the ledger's verified bucket, which an earlier round did not
have in scope when it reported this category closed.
Co-authored-by: Vincent Koc <vincent@openclaw.org>
Two nav repairs (an orphaned reference page with a real inbound link, and
two start/ pages sitting under Help > Community), Related lists and index
cards that omitted whole top-level areas, and a plain-English pass over the
pages carrying rate findings.
Hard STE violations across the 46 measured pages: 649 -> 334 at the 25-word
cap, 767 -> 388 at 20. Every page named by a rate row is now under 1.5 at
both caps except reference/templates/AGENTS.md, refuted separately.
Co-authored-by: Vincent Koc <vincent@openclaw.org>
Structure-only pass over 50 open `ia` audit rows in docs/gateway/,
docs/concepts/, docs/install/ and docs/help/. No page splits, no page
moves, no URL changes.
On-page structure:
- concepts/multi-user: six H3s inside the 1,288-word per-person accounts section
- concepts/model-failover: one notices section; H2 blocks reordered to
storage -> rotation -> cooldowns -> fallback -> notices
- concepts/compaction: 'Provider checkpoints' and 'Successor transcripts'
regrouped under a new 'Provider and engine behavior' H2
- concepts/memory-builtin: 'When to use' moved after 'What it provides'
- concepts/queue: 'Scope and guarantees' split into 'Input durability' and
'Lanes and scope'
- concepts/session, concepts/session-tool: 'Further reading' merged into 'Related'
- help/debugging: sections reordered (watch mode first); 'Safety notes'
demoted to H3 under raw stream logging; Node/tsx errors beside VSCode
- help/environment: OPENCLAW_HOME moved under 'Paths and instances'
- help/testing-updates-plugins: 'On this page' section index
- gateway/logging, gateway/multi-tenant-hosting, install/backups: duplicate
body H1 removed (precedent 204971f2a9), old id kept as an <a id> stub
- install/upstash: 'Next steps'; vps.md: Upstash Box and Render cards
Navigation (docs.json), no URL changes:
- install/nix -> Runtimes; install/ansible -> Hosting > Self-hosted and local
- gateway/clients and gateway/external-apps ahead of gateway/protocol
- gateway/cli-backends, local-models, local-model-services -> new
'Models and local providers' group
- gateway/heartbeat -> Capabilities > Automation
- gateway/portals -> Web interfaces
- gateway/security/dependency-locking -> Release & CI > Release process
- concepts/typing-indicators -> Messages and delivery
- concepts/usage-tracking, concepts/timezone -> Technical reference
Anchors: 15 changed pages enumerated with parseDocsDocument before and
after. Zero ids lost, zero collisions; 11 added.
Adds the missing reciprocal links and one mis-targeted link fix for the
last open `link` audit findings.
- Related back-links: system prompt (context engine, timezone), diagnostics
flags (gateway diagnostics, gateway troubleshooting), cloud workers
(operator scopes), auth credential semantics (secrets, auth storage),
agent runtime architecture (agent runtimes), agent runtime workflow
(testing), network (remote access, architecture), threat model (gateway
security index, network proxy), backups/updating/doctor (database schemas).
- docs/security/network-proxy.md gains a Related section.
- docs/security/incident-response.md links back to the three sibling pages
that already link to it.
- docs/concepts/typing-indicators.md links the heartbeat and groups pages
that its Defaults section describes.
- docs/diagnostics/flags.md links the environment-variable reference from
the timeline section that names three OPENCLAW_DIAGNOSTICS_* variables.
- docs/concepts/main-session.md names `session.maintenance.maxDiskBytes` and
links the maintenance reference instead of stating a bare 10 GB default.
- docs/network.md pointed its "Gateway config reference" entry at
/gateway/configuration; retargeted to /gateway/configuration-reference.
- docs/openclaw-agent-runtime.md merges its References list into Related and
keeps the old `#references` anchor as a stub.
- Six zh-CN glossary sources added beside their related existing terms.
Closes link-kind audit findings for docs/install/, docs/help/, docs/reference/,
docs/start/, docs/web/ and docs/nodes/.
- start/hubs: point the Model providers hub entry at the provider directory
(/providers) instead of the /providers/models quickstart duplicate.
- Add the missing reciprocal links the audit found: Cloudflare Containers,
Kubernetes and Ansible from Docker and the Linux server page; macOS VMs from
iMessage and the Linux server page; Podman from the sandbox Podman backend;
Updating from Migrating; Docker from the Configuration page; Backups,
Bootstrapping and Default AGENTS.md from Agent workspace; Bootstrapping from
the BOOTSTRAP template; Tests from the two help testing pages; Session
management deep dive from Context engine; Transcript hygiene from Session and
Session pruning; SecretRef credential surface from Auth credential semantics;
Device model database from the Nodes macOS section; RPC adapters from Signal
and iMessage; Personal assistant setup from Getting started; Onboarding from
the macOS platform page; The Lobster from Control UI settings; Release
performance sweep from Dependency locking; Release policy from Release
channels; Full release validation and Update and plugin tests from RELEASING.
- help/index: list the Scripts page under Testing.
- help/faq-first-run: add the Models FAQ to Related (was one-directional).
- reference/credits: replace the two off-topic Related links with the lore and
pull-request-review-flow pages.
- reference/rich-output-protocol: replace the unrelated RPC adapters link with
the Control UI hosted-embeds section that actually renders [embed ...].
- web/lobster: add a Related section.
- glossary: 17 append-only zh-CN sources for the new list-item link labels, each
inserted beside a related existing term rather than at the end of the array.
macOS launchd discarded Gateway stderr during both installation and restart, leaving startup failures invisible before logging initialized. Route stderr to the existing supervisor stdout log, matching the diagnostic reader and retaining the existing log-retention owner. Update status hints and troubleshooting documentation to identify the shared log.
Ordinary restart rewrites and reloads the generated plist; preserveDefinition retains its existing behavior. Hosted CI passed on a2616fa8dd86c1b76fc0f4f83e61b6ef640b9c60, with install/restart plist and diagnostic coverage: https://github.com/openclaw/openclaw/actions/runs/34438914271. Actual after-fix launchd capture remains unverified.
Fixes#90711.
Preserve Unix signal termination through the native runner and watcher, and distinguish acknowledged UI shutdown from interrupted or unverified cleanup.
* docs: split help/faq into 13 topic pages
docs/help/faq.md was 89,648 characters and 1,613 lines - the largest
hand-written page left in the tree. It is now a 31,704-character index:
a table of the thirteen topic pages, the triage ladder it opens with,
the two pointer sections to the first-run and models FAQs, and an
anchor-compatibility list.
Content-preserving. The thirteen children are verbatim slices cut on
H2 boundaries, so no AccordionGroup is split and no prose is rewritten,
reordered, or reformatted.
All 143 pre-split anchor ids still resolve on /help/faq: 14 are still
published by the index itself, and the other 129 are authored <a id>
stubs pointing at the page that now holds the answer. Zero id
collisions on the parent and on every child.
Matches the docs/help/testing/ and docs/help/testing-live/ split shape
already on main.
* docs: link the relocated data-locality answer from the Foundation answer
Addresses the ClawSweeper P3 finding on
docs/help/faq/what-is-openclaw.md:68. The Foundation answer said to see
"Is all data used with OpenClaw saved locally?" below; after the split
that answer is on docs/help/faq/where-things-live-on-disk.md, so the
directional reference is replaced with an explicit link to it.
This is the second and last declared prose rewrite in this PR. With both
reversed, the children still reassemble byte-identically to the original
docs/help/faq.md (sha256 fe2dd3d6...).
`docs/help/faq-first-run.md` was 40,388 characters in one `## ` section
holding 52 accordions across two `<AccordionGroup>` blocks.
An `<AccordionGroup>` cannot be split across files without inventing a
second wrapper, so the only content-preserving cut is the boundary
between the two existing groups:
- `help/faq-first-run/quick-start` - install, onboarding, first-run
failures, builds, and subscription basics (group 1).
- `help/faq-first-run/providers-and-hosting` - provider auth and limits,
model choice, hardware, and where to run the Gateway (group 2).
The parent stays as an index. Every one of the 56 pre-split ids still
resolves on it: `related` is still published there, and the other 55 are
authored `<a id="...">` stubs pointing at the child that now holds the
content, matching `docs/help/testing.md`.
Content-preserving: the two child bodies, with frontmatter and ledes
removed, concatenate byte-identically to the original body
(sha256 040670d96d7a8819c40d5f6bb924cccf1e9cd48596f5c264517a388cfa773911).
No prose was rewritten, reordered, or added inside the moved content.
docs/help/testing-live.md was 46,841 characters of 17 independent live
lanes. Split it into docs/help/testing-live/ children, one per reader job,
and keep the parent as an index that still carries the safety and
credential sections plus an anchor stub for every pre-split heading id.
Content-preserving: the children are verbatim slices of the original body.
Co-authored-by: Vincent Koc <vincent@openclaw.org>
* docs: fix accuracy findings in concepts, start, install, and help
Corrects statements in the onboarding docs that disagree with the code, and
scopes version-locked claims to the release that changed them.
Each change is backed by a source reference; findings that turned out to be
wrong about the code are listed in the PR body rather than "fixed".
* docs(retry): separate Discord Gateway reconnects from request retry defaults
The "Applies to" column listed gateway reconnects alongside Discord sends.
The WebSocket reconnect loop does not use that envelope: gateway.ts allows
50 attempts and backs off from 2000 ms with no jitter, while retry.ts governs
per-request retries only.
Addresses the ClawSweeper P2 finding.
* docs(tools): split the browser page by reader job
docs/tools/browser.md was 55,693 characters across 21 H2 sections mixing
explanation, how-to, reference, and troubleshooting. Move each section into
docs/tools/browser/ and keep /tools/browser as the index.
The split is content-preserving: the nine child pages plus the three blocks
the index still publishes (lede, What you get, Related) reconstruct the
original body byte-for-byte (sha256 a28d5ea4...). Only the four H2 lines that
became child page titles are not reproduced verbatim; each keeps its anchor
on the index.
Every one of the 43 pre-split anchor IDs -- headings, their percent-encoded
and cleaned variants, and the Accordion and Tab titles -- is authored as an
<a id> stub on the index, so existing /tools/browser#... links still resolve.
Per-anchor redirects are not possible: redirectSource() rejects sources
containing [?#].
* docs(tools): retarget cross-references orphaned by the browser split
Eleven link targets, no prose changes. Three were same-page or same-route
anchors inside the moved content: /tools/browser/setup's [Configuration] link
was genuinely broken (its target moved to the configuration page), and the
[Profiles] and [Custom Chrome MCP launch] links resolved only through the
index's anchor stub. The remaining eight are inbound deep links from other
pages, retargeted at the page that now holds the section, matching the
code-mode split precedent.
* docs(i18n): add zh-CN glossary entries for the browser child page titles
check-docs-i18n-glossary requires a source term for every changed doc label.
live-workflows.md ended with "prefer narrowing live tests via the
allowlist env vars described below" — but the page ends on the next
line, and those env vars now live on docs/help/testing/docker.md after
the testing guide was split.
This is the failure mode no gate in this repository detects. The
sentence is valid prose and contains no link, so the link audit, the
formatter, markdownlint and the MDX check all pass. It reached main
through a full review.
Found with a new .audit/check-orphan-refs.py, which flags directional
phrasing for a human to confirm. Scoped to the six split trees it
reports 8 candidates, of which this was the one real defect; the rest
resolve on their own page.
Add OPENCLAW_CONFIG_READONLY=1 to the shared immutable-config policy so externally managed deployments do not need unrelated Nix behavior.
Keep the selector host-owned across config reload and service regeneration. Gate normal writes and recovery rewrites before altering source bytes or creating recovery artifacts, while preserving mutable defaults and existing Nix guidance. Include compatibility and failing-before recovery regressions.
Closes#140706
Worked on by:
- @sallyom
Co-authored-by: sallyom <11166065+sallyom@users.noreply.github.com>
docs/gateway/security/index.md was 89,057 characters, 9,905 words and 25 H2
sections mixing explanation, how-to, reference and vulnerability-report triage
policy for two audiences. The page already lived in a directory with five
siblings, so the split extends that directory rather than creating a parallel
one, and index.md becomes a real index.
New children in docs/gateway/security/ (alongside audit-checks, exposure-runbook,
rate-limiting, secure-file-operations and dependency-locking):
- trust-model.md (Security trust model) - scope, trust boundary matrix, findings
closed as no-action, gateway/node trust, threat model, reporting.
- running-the-audit.md - the `openclaw security audit` command, what it checks,
triage order.
- hardened-baseline.md - both copy/paste baselines and the requester-scoped
controls note.
- access-control.md - DM policy, allowlists, DM session isolation, context
visibility, command authorization.
- prompt-injection.md - prompt injection, external-content wrapping, bypass flags.
- tool-permissions.md - control-plane tools, node execution, dynamic skills,
plugins, sandboxing, per-agent access profiles.
- browser-control.md - browser control risks and the SSRF policy.
- network-exposure.md - bind/firewall, Docker/UFW, mDNS, Gateway auth, Tailscale,
reverse proxy, HSTS, Control UI over HTTP, dangerous flags.
- secrets-and-storage.md - host trust, secrets on disk, credential map,
permissions, workspace .env, logs and transcripts, secret scanning.
- operator-incident-response.md - contain, rotate, audit, collect.
Anchor strategy
Per-anchor redirects are not possible: redirectSource() in
scripts/lib/docs-redirects.mjs rejects any source containing [?#]. Every anchor
the old page published is therefore kept alive on the index itself as an
authored <a id="..." /> stub inside a "Where each section moved" list, each
pointing at its new home. Ids were computed with parseDocsDocument, not a slug
approximation, so percent-encoded punctuation and compatibility aliases are
preserved exactly. The index publishes no original heading itself, so no stub
can collide with a canonical id.
- pre-split ids on /gateway/security: 80 (57 headings, 20 compatibility
aliases, 3 accordion targets)
- still published by /gateway/security after the split: 80
- index collisions reported by parseDocsDocument: 0
- stub destinations checked against the destination pages: 60 (all resolve)
Losslessness
Reassembling the child section bodies in original order reproduces the original
body byte for byte. Counts, original body vs children:
- words 9,665 -> 9,664
- characters 87,422 -> 87,491
- code fences 32 -> 32
- markdown links 44 -> 45
- table rows 38 -> 38
The single word/link delta is the one declared prose edit: a cross-reference the
split orphaned, `(see "Per-agent access profiles" above)` in the secure-baseline
section, is now a link to /gateway/security/tool-permissions. The index intro
line "The rest of this page is the deep end" became "The pages below are the
deep end" for the same reason. No security control, threat-model statement,
default or guarantee was reworded, softened or reordered; sections keep their
original relative order within each child.
Also updated: the "Security and sandboxing" nav group in docs/docs.json, eleven
in-repo deep links repointed at the new pages (release notes left untouched, and
their anchors still resolve through the stubs), and
src/docs/environment-docs.test.ts, which asserts on the workspace-dotenv text
that now lives in secrets-and-storage.md.
Closes audit findings: r3-0302, r3-0308
* docs(reference): split the testing reference by reader job
docs/reference/test.md was 92,671 characters and mixed agent proof policy,
how-to steps, reference tables, and runner-internals explanation on one page.
It is now a short index over six children, one per reader job:
- reference/test/local Routine local order, core commands, PR gate
- reference/test/lanes Control UI, TUI, extension, Gateway, E2E lanes
- reference/test/docker Docker scheduler knobs and the notable lanes
- reference/test/performance Profiling, shard timings, benchmark scripts
- reference/test/runner-internals Build locks, test state, JSON report merging
- reference/test/remote-proof Crabbox/Testbox policy, wrapper, lease, trust
Anchor strategy: per-anchor redirect routes are impossible here, because
redirectSource() in scripts/lib/docs-redirects.mjs rejects any source
containing [?#]. Every pre-split anchor instead stays alive on the parent as
an authored <a id="..." /> stub in the "Where each section moved" list, the
same mechanism docs/ci.md uses. All 30 ids were enumerated with
parseDocsDocument, never a hand-rolled slug, so the four punctuated headings
keep both the encoded and the cleaned id (for example
full-docker-suite-(pnpm-test%3Adocker%3Aall) and
full-docker-suite-pnpm-testdockerall). 29 ids are stubbed; `related` is not,
because the index still publishes that heading itself, and stubbing it would
raise a duplicate authored/canonical ID collision.
Verified independently of docs-link-audit, which cannot see the regression:
the split rewrote the repo's own links, so the audit reads clean even when
external deep links break. Resolving all 30 pre-split ids against the parsed
post-split index gives 30 resolved, 0 dead, 0 collisions, and all 24 onward
deep links land on a real fragment of a real child.
Losslessness (bodies, frontmatter excluded):
code fences 20 -> 20 (+0)
table rows 47 -> 47 (+0)
inline code spans 530 -> 530 (+0, byte-identical multiset)
fenced blocks 10 -> 10 (byte-identical, so commands are unchanged)
chars 92,523 -> 92,253 on children + 5,093 on the index
words 10,008 -> 9,981 on children + 381 on the index
links 13 -> 9 on children + 35 on the index
The chars/words/links deltas reconcile exactly: the two intro bullets (170
chars) and the Related list (139 chars, 3 links) stay on the index, and one
declared edit adds 40 chars and one link.
The single declared prose edit: "Local test commands below are the normal
trusted development path" pointed at a section the split moves to another
page, so "below" became a link to /reference/test/local. No other prose was
rewritten; the remaining prose findings stay open for a follow-up.
Test pins repointed. test/scripts/docs-sync-publish.test.ts pins the exact
Release & CI navigation page list and breaks on the docs.json nav addition;
its route list now includes the six children. The two QA Lab producers point
their docsRefs at reference/test/local.md, the page that now owns the
commands they run, instead of the index. changed-lanes.test.ts needs no
change: it uses the path only as a docs-path example and asserts nothing
about the content.
Closes audit findings: r3-0695, r3-0696
* docs(reference): link the relocated local test commands
ClawSweeper found a second orphaned cross-reference the split missed.
remote-proof.md said "Local test commands below", but after the split
those commands live at /reference/test/local while this page continues
with remote-proof instructions, so "below" pointed at nothing.
docs-link-audit --anchors: 0 broken links.
The Node requirement has changed nine times in 2026 and the reasons (the
SQLite WAL-reset corruption floor and the separate node:sqlite TEXT NUL
decoder bug) were buried in install prose; the Bun page's Caveats section
had become the de facto Bun contract while sitting under Containers.
- Add docs/install/node-compatibility.md: supported lines, why the floors
exist, platform consequences, what each installer provisions, the guard
diagnostic, and a sourced history of the requirement across releases.
- Add docs/install/bun-compatibility.md: Bun requirements, per-platform
SQLite builds, macOS library selection and OPENCLAW_SQLITE_LIBRARY with
the preload migration, the memory scan fallback, known limitations, and
release history.
- Keep docs/install/node.md and docs/install/bun.md as install how-tos;
move the contract paragraphs to the new pages and link them.
- Add a Runtimes nav group (Node, Node compatibility, Bun, Bun
compatibility) and move Bun out of Containers; no URLs change.
- Point the environment reference, memory config, and install overview at
the new pages; add zh-CN glossary entries; add a docs guide bullet to
refresh the tables when the runtime floors in code change.
docs/help/testing.md was 79,946 characters and mixed how-to, reference,
contributor conventions, and roadmap content in one page. It is now a
short index over six pages, one per reader job:
- help/testing/suites: quick start, the suite reference, which suite to
run, the live-test pointer, docs sanity, and offline regressions.
- help/testing/live-workflows: live provider debugging through the
Docker and Parallels lanes.
- help/testing/docker: the Docker "works in Linux" runners, their
weighted scheduler, the lane catalog, and their env vars.
- help/testing/qa-runners: the qa-lab command surface, the shared Convex
credential contract, and adding a channel to QA.
- help/testing/contracts: plugin and channel contract tests.
- help/testing/writing-tests: temp-directory rules, the agent
reliability eval gaps, and how to add a regression.
Anchor strategy: per-anchor routes are impossible because redirectSource()
rejects any source containing [?#]. Instead every one of the 49 ids the
old page published stays alive on the index as an authored <a id="..." />
stub in "Where each section moved". Ids were computed with
parseDocsDocument, not a slug approximation, so punctuated headings keep
both emitted forms (for example
docker-runners-(optional-%22works-in-linux%22-checks) and
docker-runners-optional-works-in-linux-checks). The index still publishes
`related` itself, so that id is deliberately not stubbed and no
duplicate authored/canonical ID is raised.
Losslessness, asserted mechanically rather than by eye: all 14 original
section bodies are character-identical after the move (0 lost, 0
changed), and the page lede is byte-identical. Word count 9,632 -> 9,632,
code fences 18 -> 18, links 13 -> 13, table rows 0 -> 0. All 751 inline
code spans and fenced blocks compare as an identical set, so every
command in the guide is unchanged. The only body delta is six trailing
newlines removed by scripts/format-docs.mts.
Verified independently of docs-link-audit, which a split makes
uninformative because it rewrites the repo's own links: the 49 pre-split
ids were enumerated from HEAD, the post-split index and every child were
re-parsed, and each id was asserted to resolve. 0 unresolved, 0 stub
targets that miss their child, 0 collisions.
Prose findings are deliberately left alone and deferred: r3-0398,
r3-0399, r3-0400, r3-0401, r3-1686, r3-1687, r3-1688, r3-1689, r3-1690.
Closes audit findings: r3-0397
Each fix replaces a statement that contradicts the code, the CLI reference,
or the config schema. Authorities are quoted with file:line.
1. docs/start/setup.md:64 - bare `openclaw setup` routing
Said: "Bare `openclaw setup`, without `--baseline`, is an alias for
`openclaw onboard` and runs the full interactive wizard."
Authority: src/cli/program/register.setup.ts:23-40 `resolveSetupCommandRoute`
returns "system-agent" when `input.configured && (input.interactive ||
input.json)`, and falls to "onboarding" otherwise. docs/cli/setup.md:13-16
states the same routing. Corrected to the chat-first routing with a link to
the CLI reference.
2. docs/install/installer.md:238 - Alpine container tag
Said: "use an official `node:24-alpine` container".
Authority: scripts/install.sh:19 `NODE_DEFAULT_MAJOR=26` and
scripts/install.sh:2263 `echo "Use an official node:${NODE_DEFAULT_MAJOR}-alpine
container ..."`. docs/install/installer.md:104 already says `node:26-alpine`
for the same vulnerable-SQLite fallback. Changed 24 to 26.
3. docs/specs/codex-supervision.md:178 - `codex sessions` synopsis
The synopsis omitted `[--agent <id>]` while line 191 of the same page says all
three commands accept it.
Authority: extensions/codex/src/session-cli.ts:318-320 registers
`.command("sessions")` with `.option("--agent <id>", "Agent id that owns the
Codex sessions")`. docs/plugins/codex-supervision.md:220 shows the flag.
Added `[--agent <id>]`.
4. docs/providers/moonshot.md:38-40 - Kimi K3 thinking levels
Said: "OpenClaw exposes those exact levels and maps `/think xhigh` to `max`"
after listing the vendor's `low`, `high`, and `max` reasoning efforts.
Authority: extensions/moonshot/provider-policy-api.ts:8-12 registers
ALWAYS_THINKING_PROFILES with `[KIMI_K3_MODEL_ID]: { id: "max", label: "max" }`,
and lines 32-41 return `levels: [profile]` with `defaultLevel: profile.id`, so
only `max` is exposed. src/llm/providers/stream-wrappers/moonshot-thinking.ts:13-17
maps `"kimi-k3": "max"` and lines 62-70 overwrite the outgoing payload with
`payload.reasoning_effort = effort` after deleting `thinking` and
`reasoningEffort`. The accordion at line 320 already stated the max-only
behavior correctly and is left unchanged; the contradicting sentence at the top
of the page was corrected instead. The catalog thinkingLevelMap in
extensions/moonshot/openclaw.plugin.json is model metadata and does not make the
lower efforts usable on the direct Moonshot route.
5. docs/providers/perplexity-provider.md:82 - `count` transport
The filter table marked `count` "Native only" while line 90 says native-only
filters return an error on the chat-completions path.
Authority: extensions/perplexity/src/perplexity-web-search-provider.runtime.ts:305-316
`unsupportedOptions` covers country, language, date_after/date_before,
domain_filter, and max_tokens/max_tokens_per_page - `count` is absent, so it
does not error. Line 375 of the same file passes count only on the structured
path. docs/tools/perplexity-search.md:144-146 says count is accepted and
compatibility-only on the Sonar path. Changed the row to "Both" and noted that
the chat-completions path ignores it.
6. docs/providers/qianfan.md:76 - onboarding default model
Said the config example "explicitly selects the current DeepSeek flagship
instead of the onboarding compatibility default", but it selects
`qianfan/deepseek-v4-pro`, which is the onboarding default.
Authority: extensions/qianfan/provider-catalog.ts:7
`export const QIANFAN_DEFAULT_MODEL_ID = "deepseek-v4-pro";`, consumed at
extensions/qianfan/onboard.ts:49 as `defaultModelId`. The same page states this
at lines 17 and 44.
7. docs/help/faq.md:821 - Gateway host platforms
Said: "The Gateway runs on macOS/Linux (Windows via WSL2)".
Authority: docs/platforms/windows.md:118 "## Native Windows CLI and Gateway"
with the PowerShell install and `openclaw gateway status` runbook;
docs/help/faq.md:1319 documents "Native Windows CLI/Gateway: runs directly in
Windows". Corrected to macOS, Linux, and Windows (native or WSL2).
8. docs/help/faq.md:900 - node local tools by platform
Said: "Local node tools are currently macOS-only."
Authority: docs/nodes/index.md:772-779 gives default node command allowlists for
iOS, watchOS, Android, macOS, Windows, and Linux, including `camera.list` and
`location.get` on iOS/Android/Windows and `computer.act` on Windows and Linux.
docs/help/faq.md:719 says Macs/iOS/Android nodes expose local tools. Replaced
with a platform-dependent statement plus a link to /nodes.
9. docs/concepts/features.md:35 - channels shipped in core
Said only Telegram and WebChat ship with the core install.
Authority: package.json `files` excludes non-bundled plugins as
`!dist/extensions/<id>/**` (for example `!dist/extensions/discord/**`,
`!dist/extensions/slack/**`); `a2a`, `reef`, and `telegram` carry no such
exclusion, so they ship in the package. docs/channels/index.md:12 defines
"bundled plugin"/"included in core" as shipping with the core install, and
lines 38, 54, 59, and 62 mark A2A, Reef, Telegram, and WebChat that way.
10. docs/gateway/logging.md:142 - settable console styles
Said: "Console styles: `pretty` | `compact` | `json`."
Authority: src/config/zod-schema.root-shape.ts:117
`consoleStyle: z.union([z.literal("pretty"), z.literal("json")]).optional()`
and src/config/types.base.ts:287 `consoleStyle?: "pretty" | "json";`.
docs/gateway/logging.md:82 already says `compact` is no longer settable and is
applied automatically off a TTY.
11. docs/gateway/multiple-gateways.md:34 - rescue-profile port spacing
Said: "Use a base port at least 20 higher than the main bot", while line 20 of
the same page requires 120.
Authority: extensions/browser/src/config/port-defaults.ts:27-29
`deriveDefaultBrowserControlPort` returns `gatewayPort + 2`, and lines 32-34
`deriveDefaultBrowserCdpPortRange` returns `start = controlPort + 9`,
`end = start + DEFAULT_BROWSER_CDP_PORT_RANGE_SPAN` where the span is
18899 - 18800 = 99. Each Gateway therefore reaches base + 110, so 20-port
spacing overlaps. Changed 20 to 120.
Closes audit findings: r3-0772, r3-0428, r3-0754, r3-0671, r5-0045, r3-0682, r3-0393, r3-0394, r3-0278, r3-0386, r5-0004
## What Problem This Solves
Operators can see a stored credential with `static` health while `models status --check` exits 1 because route readiness is unknown. The reference currently omits that result. The FAQ also recommends unscoped status immediately after agent-specific login guidance, which can inspect a different model selection.
## Why This Change Was Made
The model CLI reference now distinguishes inventory, credential sources, profile health, model-route issues, and runtime availability. It documents check exit codes and command-resolution failures, and links chat-session inspection to the existing `/model status` guidance. The FAQ uses an explicit agent and links to this canonical explanation.
Hosted refresh now has its own section, preserving the distinction between downloading metadata and activating it after a Gateway restart. The FAQ's obsolete numeric picker was already corrected by #139232 and remains unchanged.
## User Impact
Operators can inspect the intended agent, understand an indeterminate result without treating it as rejected credentials, and distinguish a catalog download from authentication or live activation. No product behavior, model-selection policy, auth policy, or configuration changes.
## Evidence
- Source audit at `8e47d6998a` covers CLI registration, status output and exit-code selection, credential-source/health owners, and hosted refresh.
- Reused the actual isolated `models status --agent work --json --check` observation at `4e402dda9e0ee8c8a5b01e39d7ccd61ccf6d40ad`: stored exec reference retained, profile `static`, route `indeterminate`, exit 1, and no execution of the stored reference. The relevant status/health/overview/config-loading source is unchanged at the documentation baseline.
- Reused the completed hosted refresh proof from #141885: a separate CLI download produces a Gateway restart notice without activating rows or prices. No new runtime, provider request, or credential access was needed for this docs-only change.
- `pnpm docs:list`, changed-file `oxfmt --check`, `node scripts/check-docs-mdx.mjs docs/cli/models.md docs/help/faq-models.md`, the glossary check against the pinned baseline, and `git diff --check` pass.
- Full link/anchor audits check 9,427 candidate links and 9,424 baseline links. Both report the same 39 pre-existing missing ClawHub mirror links; all three new links pass. The local external documentation mirror is unavailable. No product build or typecheck ran locally.
Co-authored-by: Ayaan Zaidi <hi@obviy.us>
* feat(sqlite): select an extension-capable SQLite library for Bun on macOS
Bun on macOS dlopens Apple's SQLite, which omits extension loading, so
sqlite-vec never loaded and memory search fell back to the batched scan.
Bun's only hook is bun:sqlite Database.setCustomSQLite: one-shot, before
the first open, and fatal to every later open on a bad path.
- Add src/infra/bun-sqlite-library.ts: validate each candidate through
bun:ffi (WAL-reset-safe version, extension loading present) before
committing; resolve explicit path > OPENCLAW_SQLITE_LIBRARY >
Homebrew/MacPorts discovery; memoize process-wide. No-op on Node and
on Linux/Windows Bun, whose static SQLite already loads extensions.
- Select before the first open in the runtime guard probe and in
requireNodeSqlite; an unusable override becomes a clean runtime-guard
diagnostic with exit 1 instead of a stack trace.
- Forward the selected library to the memory KNN child through its typed
stdin input and select there before opening; the child env stays
stripped. The #141104 scan fallback remains the no-library path.
- Report the selection at Gateway startup and in doctor.
- Retire the undocumented legacy OPENCLAW_CLAUDE_CLI_LOG_OUTPUT alias so
the OPENCLAW_* name budget stays at 493.
- Docs: Bun install caveats, environment reference, memory config.
* test(sqlite): retain the KNN stdin spy for payload assertions
* docs(bun): explain migration from custom SQLite preloads
* fix(update): restore untracked official plugins from the core release cohort
Keep update admission and Doctor repair on the selected stable core version when formerly bundled official plugins have no install record. Preserve existing selectors and capability policy. Add DuckDuckGo published-upgrade and standalone Doctor coverage for bundled and already-missing states.
Related: #135002, #116740
* test(update): recognize missing optional plugin warnings
* test(update): isolate plugin repair proof and satisfy helper lint
* test(update): preserve beta tag selection in doctor fixture
* chore(ci): register plugin doctor survivor entrypoint
* fix(qa): verify leased Telegram tester group access
Check tester membership and effective text permission in the leased user driver before group readiness. Reuse the check in doctor and select trusted skill scripts for frozen release candidate QA. Preserve supported private DM turns and leave credential repair to the pool owner.
* ci: install Chromium for native live browser tests
The native-live-test shard includes the real Gateway widget restart proof,
which launches Playwright after restoring the interrupted session. Chromium
was installed only in separate Repo E2E jobs, leaving the live shard without
its required executable.
Install the candidate UI's pinned Chromium and system dependencies only for
the selected native-live-test row. Preserve profile selection and all test
gates. Verified the unchanged dc420 widget restart test with real OpenAI,
Gateway process replacement, and interactive Chromium dashboard assertions
on Blacksmith Testbox; canonical workflow checks and formatting also pass.
* test(release): prove supported cross-OS 9.2 transition
Select the explicitly supported external package-manager and fresh Doctor
transition only for 2026.9.2 to 2026.9.3. Preserve the old updater's safe
schema-15 refusal as negative evidence and never report self-update passed.
Verify backup, exact package identity, schema migration, retained settings
and session history, candidate serving version, and a fresh persisted turn.
Keep credential-bearing backup archives out of CI evidence and remove them
on both successful and failed runs. Other release pairs retain existing
updater and Windows fallback behavior.
Validation: 13 cross-OS upgrade lane tests, two pre-fix regression failures,
full changed checks, and independent review. Real packaged cross-OS/live
qualification remains required against the aggregate candidate.
* test(installer): prove supported historical package transitions
* test(release): exercise same-schema updater boundaries
Keep updater-specific plugin repair and consent checks on the prepared
candidate, then target explicit synthetic later versions with the same
runtime and storage schema. Retain installed-version and plugin-policy
assertions, and remove the obsolete legacy post-update fallback.
Add versioned future-tarball generation to the existing first-hop fixture
owner, with immutable input and source/target digest receipts. Restore the
reserved runtime-promotion staging ignore in package-derived Git fixtures
so the updater's unchanged-source check stays meaningful.
Validation: 26 focused tests, a pre-fix Git staging regression, full changed
checks, exact candidate tarball fixture generation, and P2 review. Actual
Docker update/consent/channel qualification remains required in aggregate CI.
* test(release): verify channel update staging cleanup
Assert the real channel flow refuses ordinary untracked user files without changing source, then require successful channel updates to remove every reserved runtime staging entry. Preserve recovery data when the assertion detects leftovers. No ignore or production cleanup policy is broadened.
* test(onboard): select configured model through explicit picker
* ci: pin release performance checks to reviewed Kova accounting
Select the reviewed Linux process-lifetime CPU accounting and bounded worker-watchdog implementation from Kova PR #110 for canonical, legacy-list, and trusted live fixtures. Keep existing trust classification, timeout values, and performance thresholds.
* fix(ci): distinguish bundle build modes in job names
* test(release): build matching future runtime fixtures
Derive a synthetic Codex runtime cohort from its verified candidate package by changing only package.version and openclaw.build.openclawVersion. Require matching source metadata and preserve payload, constraints, immutable input, and source/target digest receipts. Reuse the existing private tarball lifecycle and sequence validation. Validated 11 focused tests, pre-change CLI failures, full changed checks, and P2 review.
* docs(release): prepare private QA before local E2E
* fix(release): admit package metadata before validation dispatch
Port the reviewed early source admission into trusted Tooling, preserving the current selected REST transport. Reject invalid version notes and misaligned core packages before any Git push, non-GET API write, or workflow dispatch; retain explicitly allowed substantive draft notes.
* test(release): qualify supported historical upgrade transitions
Prove the published 2026.8.2/2026.9.2 to 2026.9.3 external package-manager and fresh Doctor contract through one shared helper. Preserve the old in-process updater refusal and schema 15, require owner-stopped migration with a verified private backup, and verify schema 16 plus retained Gateway history afterward.
Keep current-runtime update coverage distinct: managed restart uses separately identified future core and matching runtime fixtures, with real updater, canary, service replacement, authentication, and serving checks. Fix fixture-only channel environment pollution and mock serving nonce responses.
Sealed Docker journey, historical first-hop, current-candidate survivor, root-managed VPS, and onboarding pass. Managed-auth historical preservation passes; its future canary exposes an independent existing plugin projection blocker, retained as failing proof for the candidate repair owner. No product schema or refusal guard changes.
* test(release): narrow recorded update command before indexing
* fix(ci): align historical upgrade fixture ownership
Keep all seven external transition assertion cases with the first-hop package fixtures: both own the historical upgrade boundary and now share its temporary-state lifecycle. Preserve helper test selection and all assertions while avoiding an additional standalone tooling group.
Provide the new inference and future-package preparation boundaries in isolated service, mobile, and cron bootstrap probes, and verify each updater argument without changing readiness or migration assertions. Register the path-launched assertion CLI as a Knip entry.
Validation: 67 related tests; 70 exact-merge planner and combined-suite tests; five 80-job stress plans; full Knip unused-file scans; scoped checks and P2 review.
* fix(ci): model upgrade preparation in isolated probes
Complete the companion fixture and Knip repairs for the preceding transition-test consolidation. Model the independent inference and future-package preparation producers, verify exact updater argument boundaries, and retain service-readiness, authored-state, and cron migration assertions. Declare the path-launched transition assertion CLI in the existing Knip entrypoint list.
* test(release): preserve fixture registry on managed restart
Capture the fake service manager registry alongside its endpoint paths, then refresh that owned context after preparing the future cohort. Projected native service callers cannot drop or replace it; transient update state still stays outside the managed child.
The executable service regression fails on the original shim and passes with missing and conflicting caller registries. Four focused tests, the changed gate, and P2 review pass. Reuse the existing d928 package for the final managed-upgrade replay.
* test(release): preserve prepared UI identity in future fixtures
Advance synthetic package versions while retaining the opaque build ID already embedded in unchanged Control UI assets. The canonical asset-health regression fails the prior builder as stale, then verifies readiness and unchanged UI bytes for both future fixture sequences. Package-version update admission and managed-restart proof remain unchanged.
* test(release): retain manager-owned fixture policy
Capture the existing automation and offline-channel policy with registry identity when generating the manager shim. Native service clients project these values away; losing them activated retained synthetic channel credentials and correctly failed readiness. Restore only this explicit manager allowlist and keep transient update/compatibility state excluded.
The executable service regression fails on missing manager flags and passes across projected or conflicting caller environments. Four focused tests, the full changed gate, and fresh P2 review pass. Actual corrected managed Docker proof remains pending.
* docs(ci): keep bundle identity with routing guidance
* fix(release): retry transient GraphQL EOF failures
Classify the observed gh unexpected EOF transport failure within the existing five-attempt retry budget. Preserve authentication, invalid-response, inventory and snapshot handling. Exercise recovery, exhaustion and fail-fast errors through the real verifier CLI with a fake gh transport; also modernize the existing disposable-array sort required by touched-test lint.
* test(release): align survivor fixtures with manager preparation
Supply manager environment JSON to directly launched supervisors, and model the independent manager preparation phase in config-parking and companion fixtures. Preserve service readiness, restart, diagnostics, and strict phase-order assertions without changing production behavior or test limits.
* test(release): align upgrades with deferred schema publication
Restore installed-updater survivor and first-hop coverage after shared migration ownership moved into the durable update ledger. Record applied schema content separately from published schema version, retain the genuine missing-chunk negative control, and label external installation as an alternate with no self-update attempt. Preserve production-owned unsupported migration refusals.
* test(release): restore supported packaged upgrade coverage
Restore real 2026.9.2 to 2026.9.3 self-update after ledger-driven schema migration support. Keep retained conversations, private backup, settings, candidate identity, and serving inference proof. Require migrated content independently of publication grace and remove obsolete external-install diversion from cross-OS and installer update smoke. Shared external helper remains separately owned.
* test(release): refresh legacy manager before updater restart
Bind the service manager to the candidate registry before the published baseline updater starts its managed service. Preserve restart verification for successful and recoverable outcomes alongside explicit future targets. Exercise stale registry capture, replacement failures, pre-update identity, and attribution without changing timeouts or caps.
* fix(runtime): require Node builds with lossless SQLite reads
* fix(runtime): preserve upgrades and guard sealed workers
Validate downloaded Node before switching the active runtime alias, reject unsupported sealed-worker runtimes, and keep the Gateway error fixture on a supported Node release. Document the approved ARMv7 and older macOS compatibility losses and decoder fix boundaries.
* test(runtime): use typed process exports in worker fixture
* test(runtime): align installer fixtures without growing test shards
* test(runtime): align release and guest runtime fixtures
* fix(test): canonicalize Windows temp roots for Node 24
Expand Windows short paths before creating test directories and owned child environments. Node 24 filesystem watchers otherwise abort when native long event paths differ from inherited short temporary paths. Preserve explicit custom-root spelling and existing cleanup ownership.
* test(ci): run Windows temp-root regressions in the native lane
* ci: gate published upgrades with legacy operator state
Add a latest-release PR/main tripwire and runtime-resolved weekly and release baselines. Exercise baseline-authored agents, approvals, cron owners, a plugin, and mock agent turns; preserve the narrow unfenced-updater refusal contract.
Related: #140778, #140784, #140886, #140825.
* test: verify native published upgrade lifecycle
Persist JSON-era CLI approval IDs, retain a moving plugin selector, and exercise the baseline updater-owned restart. Query cron owners before repair, reject malformed canonical approvals, and require clean Doctor evidence.
* test: keep survivor installs in disposable runtime state
Retain verification logs and hashes while letting the container own the legacy scenario npm prefix. Other survivor scenarios keep their existing artifact layout.
* test: preserve immutable sorting in CI docs guard
* fix: preserve historical release upgrade qualification
* test: handle absent release scenario output in assertions
* ci: preserve survivor coverage within existing limits
Keep all native operator and schema cases in their existing owner suites. Register the shell observer entry and current inert scenario catalog, and update fixture, request, and routing contracts without changing runner caps or assertions.
* ci: pair expanded upgrade baselines with native operator state
* test: classify upgrade baseline provenance as internal
* ci: preserve docs-only main push exclusions
* test: require successful upgrades from every supported baseline
Remove the temporary 2026.9.2 schema-refusal pass after #141109. Observe applied shared-state content while publication is deferred, retaining candidate, state-preservation, restart, and validated plugin-consent recovery checks.
* test: require legacy agent migration before upgrade probes
Snapshot the seeded roster and legacy session specimens independently of existing SQLite files. Require migrated agent stores, imported history, and preserved recovery artifacts before candidate probes can initialize state.
* test: scope upgrade specimens to session migration storage
Ignore native model catalogs and unrelated per-agent artifacts while preserving strict classification inside the legacy sessions directory and the existing SQLite runtime inventory.
The agents configuration reference was 1,586 lines / 94,062 characters with
47 headings, and one H2 (`Agent defaults`) held 7,858 of its 11,143 words.
That is roughly five times the 20k split threshold. It is now a short index
plus eight domain pages under `docs/gateway/config-agents/`, matching the
`docs/web/control-ui` split and reusing the wording of the
`configuration-reference` split.
Children (all under the 20k threshold):
- `workspace-and-bootstrap` — workspace, cwd, repoRoot, skills, bootstrap
injection, the context budget map, image handling, timezone (11,163 chars)
- `models` — `agents.defaults.model`, fallbacks, media model slots,
`modelSelectionScope` (18,517 chars)
- `runtime-and-cli-backends` — runtime policy, CLI backend selection, GPT-5
personality (5,172 chars)
- `heartbeat-compaction-and-streaming` — heartbeat, system agent, compaction,
context pruning, block streaming, typing indicators (15,129 chars)
- `sandbox` — the `agents.defaults.sandbox` block (10,718 chars)
- `entries-and-multi-agent` — `agents.entries`, `multiAgent` bindings, match
fields, access profiles (10,826 chars)
- `sessions` — `session.*` (9,869 chars)
- `messages-and-talk` — `messages.*` and `talk.*` (11,964 chars)
Anchors. `redirectSource()` in `scripts/lib/docs-redirects.mjs` rejects any
source containing `[?#]`, so a fragment can never reach a path redirect and no
redirect is added; the parent keeps its route at `/gateway/config-agents`.
Every anchor the single page published stays resolvable on the index instead.
All 82 ids were enumerated with `parseDocsDocument`, not a hand-rolled slug:
81 are retained as authored `<a id>` stubs in a "Where each section moved"
list, and `related` stays a published heading, so no id is both stubbed and
published and the parser reports zero collisions. Each stub links to the page
that now owns the heading, and each of those 63 in-set links was asserted to
resolve against the child's own parsed id set. 25 anchored in-repo references
across 18 files are repointed at the owning child; the 18 links to
`#agent-defaults` keep pointing at the index, which is now the map of the five
pages that H2 became.
Losslessness. Each child body is byte-identical to its original line range,
with two exceptions: 34 headings are raised one level so each child has a
top-level structure, and one cross-reference the split orphaned
(`agents.entries.cwd` -> "Working directory") is repointed from
`/gateway/config-agents#agents.defaults.cwd` to the workspace page. No prose
was rewritten. Fenced blocks 80 before / 80 after. Table rows 29 before /
29 after, unchanged per table: context budget ownership map 6, model selection
scope 6, runtime policy 10, response prefix 7.
`src/config/talk-defaults.test.ts` hard-coded `docs/gateway/config-agents.md`
as the Talk defaults parity source and now reads
`docs/gateway/config-agents/messages-and-talk.md`. The eight children are
registered in `docs/docs.json` as a nested group, mirroring Control UI.
zh-CN glossary entries were added for the eight new titles; those translations
are unreviewed.
Closes audit findings: r3-0297, r3-0299
The github-copilot connection-bound-ids live suite returned green when
OPENCLAW_LIVE_TEST=1 was set but no Copilot token could be resolved from env or
the auth profile, so a Testbox run could report a passing live test that never
contacted the provider. The music and video generation live sweeps had the same
shape: an unfiltered run with zero attempts warned and passed.
Use the Vitest test-context skip(reason) in the Copilot test, and thread the
context skip through the music/video sweep summary helpers so unfiltered
zero-attempt runs show as skipped in the reporter with the skipped entries.
Filtered zero-attempt runs still throw; a present-but-invalid token still fails.
Unit coverage asserts the skip callback contract for both sweep helpers, and
docs/help/testing-live.md records the rule that a live suite without
credentials must skip visibly or fail, never pass green.
The configuration reference was 2,077 lines / 146,192 characters with 55
headings, roughly seven times the 20k split threshold and well past the
40-heading split signal. It is now a short index parent plus nine
domain pages. No prose was rewritten: every moved H2 section body is
byte-identical to its source.
Children (all moved verbatim from configuration-reference.md):
- gateway/config-runtime (1,117 words): worktreeRoot, Models, Discovery,
Update, ACP, Wizard, Bridge (legacy, removed)
- gateway/config-extensions (2,221 words): MCP, Skills, Plugins,
Canvas widget presenter
- gateway/config-browser-ui-desktop (2,012 words): Browser, UI, Desktop
- gateway/config-gateway (3,600 words): Gateway, incl. OpenAI-compatible
endpoints, multi-instance isolation, gateway.tls, gateway.reload
- gateway/config-cloud-workers (1,495 words): Cloud worker environments
- gateway/config-hooks (3,273 words): Hooks, incl. HTTP contract, agent
payload, session policy, mapping, retries and fan-out, Gmail
- gateway/config-secrets-env (1,042 words): Environment, Secrets,
Auth storage, Config includes ($include)
- gateway/config-observability (1,190 words): Audit, Logging,
Diagnostics, Telemetry
- gateway/config-automation (958 words): Automations (cron), Media model
template variables
Anchor preservation. docs.json redirects match on pathname only - 0 of
the 281 existing redirects carry a fragment in `source`, and fragments
never reach a redirect matcher - so no path redirect is added and the
parent keeps its route. Instead every anchor stays resolvable on
/gateway/configuration-reference: all 32 original H2 headings remain as
one-line pointer sections, and all 23 original H3 anchors are retained
as authored <a id> stubs (the pattern already used in
docs/help/faq-first-run.md). All 46 anchored in-repo references across
27 files are rewritten to the child page that now owns the heading;
docs-link-audit --anchors reports 8,717 links checked, 0 broken.
Losslessness (parent + nine children vs. the old single file):
- code fences: 40 -> 40
- MDX components: 4 -> 4
- distinct config keys documented: 490 -> 511 (0 lost, 21 gained from
the new index text)
- words: 16,756 -> 17,542 (+786, all new index and lede text)
- headings: 55 -> 83 (55 original + 27 pointer H2 + 1 index H2)
- parent page: 2,077 -> 226 lines, 146,192 -> 9,701 chars, 55 -> 33
headings
src/docs/cloud-workers-config.test.ts pinned the cloudWorkers examples
to configuration-reference.md by path; it now names
config-cloud-workers.md. Twelve zh-CN glossary entries were added for
the new and existing "Configuration - <domain>" page titles.
Closes audit findings: r3-0286, r3-1510
Walks the new reader path end to end and fixes each step where the docs
sent the reader somewhere the previous step had not prepared.
Landing page (docs/index.md): the Quick start installed with npm and then
ran `openclaw onboard --install-daemon`, which selects the classic wizard
(onboarding-overview.md), so the landing reader never saw the Quick start
and Custom setup choice that Getting Started narrates. It now uses the
installer script that install/index.md calls "Recommended", names which
wizard that opens, and adds the missing Gateway service step. The mobile
hub card "Get started" pointed at `/`, and the "Channels" card pointed at
one channel page while promising the catalog; both now point at the hub
they describe.
Install page (docs/install/index.md): "Verify the install" ended the page
with no route onward. Adds a next-step card group to Getting started and
to the channel hub. docs/install/node.md pointed "installer script" at the
alternative-methods anchor instead of the installer script section.
Getting Started (docs/start/getting-started.md): Step 2 left the Gateway
in the foreground and buried "Ctrl+C, then `openclaw gateway install`" in
prose, while Steps 3 and 4 assumed a running background Gateway. That is
now its own numbered step between onboarding and verification.
Channel hub (docs/channels/index.md): the "fastest setup is usually
Telegram" guidance sat at the bottom, below the catalog and a long group
introductions explanation, and the page carried no `openclaw channels add`
example. Both now sit above the 31-entry catalog.
Telegram (docs/channels/telegram.md): quick setup showed a JSON5 block
with no file path and no `openclaw channels add`, told the reader to run
`openclaw pairing list` without first sending the bot a message, and
started a second foreground Gateway that conflicts with the service the
Getting Started path installs. Step 2 now names `~/.openclaw/openclaw.json`
and leads with `openclaw channels add --channel telegram --token <token>`.
Restart and pairing are separate steps, the restart uses
`openclaw gateway restart` with `openclaw gateway` named as the
no-service case, and pairing starts by messaging the bot. The opening
line now says what the page is for.
macOS onboarding (docs/start/onboarding.md): the first three steps were
images with empty alt text. Each now says what dialog appears and which
button to click, and read_when addresses first-run readers rather than
the people implementing the flow.
First-run FAQ (docs/help/faq-first-run.md): the install answer ran the
installer and then onboarding again, the "what does onboarding do" answer
described the classic 8-step wizard as the default, install and
onboarding answers sat below heartbeat and exec-approval answers, the
provider-add command disagreed with the wizard pages, and the "I am
stuck" link landed on a heading rather than the accordion.
Also trims duplicated Quick start prose from onboarding-overview.md,
names the WhatsApp plugin prerequisite in openclaw.md, repairs the
localhost dashboard link and the install entries in hubs.md, cuts the
19-item "Start here" list in docs-directory.md, and splits the
single-instruction sentences named by the STE findings in telegram.md,
why-openclaw.md, teams.md, setup.md, and wizard-cli-automation.md.
Closes audit findings: r3-0936, r3-0937, r3-0938, r3-0939, r3-0940,
r3-0941, r3-0942, r3-0943, r3-0944, r3-0945, r3-0946, r3-0947, r3-0948,
r3-0949, r3-0950, r3-0951, r3-0952, r3-0953, r3-0954, r3-0955, r3-0956,
r3-0958, r3-0959, r3-0962, r3-0963, r3-0964, r3-0965
Docs governance and publish-hygiene pass over docs/AGENTS.md and the
pages its rules cover.
- Replace every `~/Projects` operator path in docs/ with a neutral
placeholder. 38 occurrences across 13 pages, including the private
repo path `~/Projects/manager/skills`. `docs/AGENTS.md` forbids local
paths, and its own Internal Docs bullet named one.
- Record the placeholder convention in the Published Link Rules bullet
that bans local paths.
- Document the ClawHub docs source in Source Ownership: this repo holds
no `/clawhub/**` page sources even though `docs/docs.json` lists them,
so a local preview and `pnpm docs:check-links` report those routes as
missing until `OPENCLAW_DOCS_SYNC_CLAWHUB_REPO` points at a ClawHub
checkout.
- Drop the Strict-STE hard violation rate on `docs/AGENTS.md` from 18
hard (3.4 per 100 words) to 0 by splitting semicolon sentences and
sentences over 20 words, and by naming the actor in passive
sentences. No rule changes meaning.
- Convert the Maturity Scorecard paragraph to a bulleted list, matching
every other section.
- `docs/prose.md`: name v2026.8.1 as the release that removed OpenProse,
and explain the `--agent codex` flag and the third-party skills CLI.
- Link `/prose` from `tools/skills` and `tools/slash-commands`.
- `docs/docs_map.md`: correct the summary to describe the stub, and drop
the H1 that repeated the frontmatter title.
Closes audit findings: r3-0734, r3-0736, r3-0907, r3-0910, r3-2277,
r3-2278, r3-2285, r3-2286, r3-2287, r3-2288, r4-clawhub-0001
Partially addresses r3-0906 (private path removed; publish-tree
exclusion left as a follow-up). Not addressed: r3-0909 (generator
change).
## What Problem This Solves
Fixes an issue where Skill Workshop revisions could be sent to an unrelated current chat, skill previews omitted valid Markdown or important procedure steps, and collection reviews lost inherited model-runtime settings when an agent had an empty model map.
The Workshop also carried a separate historical-review completion path used only by standalone tests, even though production already had one durable completion owner.
## Why This Change Was Made
- Remove the current-chat switch and its preference/state plumbing. Revision requests use an eligible original conversation or the existing new-conversation path, with ownership, stale-revision, and recovery checks preserved.
- Use the shared sanitized document renderer in Board and Today. Today exposes the complete skill through a native disclosure instead of guessing English section names and truncating steps. Shared styles contain images, tables, and code; external images stay click-to-open.
- Remove dead new/seen state, an unused revision callback override, duplicated prose styles, and unreachable empty-state code. Applied-history links use the canonical grouped projection and open the correct filter.
- Require prepared historical-review inputs and explicit completion. The existing batch owner handles durable completion and late failures; authoring rules come from the Workshop tool once.
- Preserve global model metadata during cron preparation. The canonical request-parameter resolver now composes global defaults, shared model parameters, per-agent model parameters, then agent-wide parameters. Thinking and fast-mode controls use the same per-agent model scope, including the final cron execution and fallback path. No cron-only parameter merge is retained.
No SQLite schema, new configuration, permissions, transcript-selection limits, or automatic-apply policy changes. Historical proposal records are retained. Separating current skills from proposal history is a follow-up, not part of this PR.
## User Impact
Users get complete, safely rendered skill documents and one predictable revision route. Background collection reviews keep the configured model runtime, including when the agent's model map is empty. Failed or interrupted historical reviews retain the existing recovery behavior.
Production delta: +215 / -704, net **-489 lines**. Tests: +482 / -207, net +275 lines. Docs/config: +6 / -6. Overall: -214 lines across 51 files.
## Evidence
- **177 focused tests passed on commit 9da2656b6e02d9d353f233dcfac2a31e51c84015**, covering Workshop service and Gateway handlers, revision ownership/recovery, historical scan completion, cron settings, shared Markdown documents, page actions, and Chromium layout.
- Captured failing regressions before fixes: wrong-chat routing from the saved preference, omitted unordered lists, incorrect applied-history navigation, remote-image loading, wide-image overflow, omitted Today sections/steps, and lost inherited cron runtime/parameters.
- A real collection automation initially failed with the wrong effective runtime. With the cron fix built and deployed, the same automation completed on the configured OpenAI model using the embedded runtime. It reviewed five skills and edited three; filesystem comparison confirmed no additions or removals.
- That live runtime was built from 6b3d88538e plus the identical cron patch. It proves the cron repair, not deployment of all other refactors. The final committed owner tests cover the combined candidate.
- Additional **360 parameter/sibling tests passed on 79f3292**, followed by **59 focused tests on the source committed as ef534a5**. The final regression failed without the executor repair: a per-agent fallback expected thinking off but received the shared high setting.
- **Served-browser proof with a real Gateway and provider passed.** Baseline with the saved current-chat preference sent a skill revision into an unrelated conversation. The candidate created a Workshop conversation, reused its eligible origin for a second revision, persisted proposal versions 2 and 3, and left the unrelated conversation byte-identical. Full mixed-language headings, all steps, bullets, code and table content remained readable. Apply followed by View history opened the exact applied version.
- That browser proof used candidate UI source 9da2656 with the retained Gateway described above. Later UI edits only reorganize empty-state returns and do not change these nonempty flows. A dev-build update banner remained visible; the browser lacked a Japanese font, although stored and accessibility text remained exact. These limits are retained with the complete private captures, not hidden.
- **Exact-head built cron proof on ef534a5** exercises the public cron command through a real Gateway and a local HTTP provider fixture. It checks the actual outgoing sampling parameters, cross-alias token-limit precedence, per-agent thinking, and fast-mode metadata. This is protocol-fixture proof, not a claim about a real provider model's capabilities.
- Direct lint, formatting, styles, English catalog baseline, and whitespace checks passed. The first ClawSweeper findings prompted the populated-parameter repair and complete served-browser proof. Fresh exact-head ClawSweeper and CI remain the landing gates.
## Review Follow-through
- The requested current-chat preference retirement is intentional and owner-authorized. Both saved values now use the same safe routing; no stored history is deleted.
- Both rank-up moves are addressed: populated model parameters have failing/passing regression and outgoing-request proof, and complete sanitized browser captures are available to the local reviewer. Raw provider transcripts and private fixture metadata are not uploaded publicly.
- Existing Today → Applied History → Today selection/loading failure is recorded for the separately authorized collection-navigation follow-up, which removes that dual-view state.
- Existing session-list thinking metadata can display the shared default even when a cron request correctly uses a per-agent model setting. The wire proof distinguishes that pre-existing display gap from actual execution; no accurate-display claim is made here.
AI-assisted.
Co-authored-by: Ayaan Zaidi <hi@obviy.us>