docs/help/testing.md was 79,946 characters and mixed how-to, reference, contributor conventions, and roadmap content in one page. It is now a short index over six pages, one per reader job: - help/testing/suites: quick start, the suite reference, which suite to run, the live-test pointer, docs sanity, and offline regressions. - help/testing/live-workflows: live provider debugging through the Docker and Parallels lanes. - help/testing/docker: the Docker "works in Linux" runners, their weighted scheduler, the lane catalog, and their env vars. - help/testing/qa-runners: the qa-lab command surface, the shared Convex credential contract, and adding a channel to QA. - help/testing/contracts: plugin and channel contract tests. - help/testing/writing-tests: temp-directory rules, the agent reliability eval gaps, and how to add a regression. Anchor strategy: per-anchor routes are impossible because redirectSource() rejects any source containing [?#]. Instead every one of the 49 ids the old page published stays alive on the index as an authored <a id="..." /> stub in "Where each section moved". Ids were computed with parseDocsDocument, not a slug approximation, so punctuated headings keep both emitted forms (for example docker-runners-(optional-%22works-in-linux%22-checks) and docker-runners-optional-works-in-linux-checks). The index still publishes `related` itself, so that id is deliberately not stubbed and no duplicate authored/canonical ID is raised. Losslessness, asserted mechanically rather than by eye: all 14 original section bodies are character-identical after the move (0 lost, 0 changed), and the page lede is byte-identical. Word count 9,632 -> 9,632, code fences 18 -> 18, links 13 -> 13, table rows 0 -> 0. All 751 inline code spans and fenced blocks compare as an identical set, so every command in the guide is unchanged. The only body delta is six trailing newlines removed by scripts/format-docs.mts. Verified independently of docs-link-audit, which a split makes uninformative because it rewrites the repo's own links: the 49 pre-split ids were enumerated from HEAD, the post-split index and every child were re-parsed, and each id was asserted to resolve. 0 unresolved, 0 stub targets that miss their child, 0 collisions. Prose findings are deliberately left alone and deferred: r3-0398, r3-0399, r3-0400, r3-0401, r3-1686, r3-1687, r3-1688, r3-1689, r3-1690. Closes audit findings: r3-0397
7 KiB
| summary | read_when | title | |||
|---|---|---|---|---|---|
| Index of the OpenClaw testing kit, one page per reader job |
|
Testing |
OpenClaw has three Vitest suites (unit/integration, e2e, live) plus Docker runners. This page covers what each suite covers, which command to run for a given workflow, how live tests discover credentials, and how to add regressions for real-world provider/model bugs.
**QA stack (qa-lab, qa-channel, live transport lanes)** is documented separately:- QA overview - architecture, command surface, scenario authoring, and the Matrix live lane.
- Maturity scorecard - how release QA evidence supports stability and LTS decisions.
- QA channel - the synthetic transport plugin used by repo-backed scenarios.
This page covers the regular test suites and Docker/Parallels runners. QA-specific runners below lists the concrete qa invocations and points back at the references above.
This page is an index. The testing kit is documented on six pages, one per reader job. Open the page that matches your task.
| Page | Read it when |
|---|---|
| Test suites and commands | You need to pick a suite, a command, or the offline regression checks. |
| Live and Docker/Parallels workflows | You are debugging a real provider or model through a live Docker or Parallels lane. |
| Docker test runners | You want the Docker "works in Linux" lanes, their scheduler, and their env vars. |
| QA-specific runners | You are running a QA Lab lane or need the shared Convex credential contract. |
| Contract tests | You changed a channel, provider, or plugin-sdk surface. |
| Writing and adding tests | You are writing a test, a regression, or a reliability eval. |
Where each section moved
Every section heading from the previous single-page version keeps its anchor
here, so an existing link such as /help/testing#qa-specific-runners still
resolves. Each entry points at the page that now holds the content.
- Quick start
- Test Temp Directories
- Live and Docker/Parallels workflows
- QA-specific runners
- Shared Telegram credentials via Convex (v1)
- Adding a channel to QA
- Test suites (what runs where)
- Unit / integration (default)
- Projects, shards, and scoped lanes
- Embedded runner coverage
- Vitest pool and isolation defaults
- Fast local iteration
- Perf debugging
- Stability (gateway)
- E2E (repo aggregate)
- E2E (gateway smoke)
- E2E (Control UI mocked browser)
- E2E: OpenShell backend smoke
- Live (real providers + real models)
- Which suite should I run?
- Live (network-touching) tests
- Docker runners (optional "works in Linux" checks)
- Docs sanity
- Offline regression (CI-safe)
- Agent reliability evals (skills)
- Contract tests (plugin and channel shape)
- Commands
- Channel contracts
- Provider contracts
- When to run
- Adding regressions (guidance)