openclaw/docs/reference/test/local.md
Peter Steinberger 8c2996e013
feat(ui): team mode shows every agent and its sessions in the sidebar (#141476)
* feat(ui): add roster-first Agents home

Add the /agents roster with activity, main-chat previews, identity cards, and direct chat navigation. Keep agent configuration at /settings/agents and reuse the existing creation flow. Include bounded session refreshes, synthetic fixtures, focused navigation and page tests, and Control UI documentation.

* fix(ui): align agents e2e navigation with roster home

Retarget settings behavior tests to /settings/agents and include Agents in default sidebar ordering expectations. Preserve existing permission, avatar, persistence, and navigation assertions. Reproduced six failures before repair; all 22 affected and sibling browser tests now pass.

* perf(ui): defer agent page copy and route setup

Register Agents home copy with the lazy view while keeping sidebar labels eager and preserving the complete source catalog. Defer agent page render adapters and the settings roster loader through the existing lazy route pattern, keeping loader logic and displayed copy unchanged.

Reduce startup gzip by 274 bytes on the branch and 275 bytes in the local CI merge-tree replay without changing performance budgets.

* feat(ui): link the agent switcher to the Agents home

Add an All agents action above New agent using the existing menu navigation and dismissal flow. Keep the Agents heading unchanged, document the entry point, and cover roster navigation and action ordering.

* fix(ui): restore eager agent settings route loading

* feat(ui): add optional sidebar agent roster

Add a browser-scoped sidebarAgentsMode preference with the existing chip as
its default. The agent menu toggles a compact roster with working activity,
last-active time, existing unread counts, and links to all agents or creation.

Share identity, activity ordering, previews, and bounded session loading with
the Agents home. Keep agent switching and selected-agent session scope on
the existing sidebar path, and defer roster rendering to protect startup.

Document the mode and include focused unit, mocked browser, and fixture proof.

* fix(ui): share roster activity between sidebar and Agents home

* feat(ui): group sidebar sessions by agent in roster mode

Show every selectable agent's pinned and recent sessions under collapsible,
working-first headers, with per-agent new-session and main-chat links.
Share bounded session activity with Agents home, preserve current-session
visibility and row ownership, and retain the context chip and list filters.
Document the 300-row window and browser-local collapse preferences.

Refs #141476.

* refactor(ui): break sidebar session navigation import cycle

* feat(ui): team mode hides Home, adds a "+" agent switcher, and scopes pages to all agents

Make each agent header its canonical main-chat entry, keep collapse separate,
and offer ordered agent selection from the sidebar New session controls.
Default shared page scope to all agents when team mode starts, preserve manual
filters during navigation, and restore the prior scope when it ends.

Show shared avatar/name chips on agent-owned lists. Keep Memory, Model
providers, and Skill Workshop single-agent, with concrete navigation selection.
Reconcile cold saved scopes and apply Memory route ownership before requests.
Document the mode and validate sidebar, scope, list, and browser behavior.

Refs #141476.

* fix(ui): persist the pre-team agent scope per gateway

Remember the pre-team page scope in gateway-scoped browser preferences, including an explicit all-agents selection. Restore it after chat navigation, reloads, and gateway switching, then clear it when team mode ends.

Validate with 424 focused unit/component tests, the sidebar browser scenario, UI typechecking, i18n, lint, and architecture checks. The full changed-file gate reproduces existing TS2352 fixture errors in the Model providers and Usage route tests at the starting commit.

Refs #141476.

* fix(ui): type route tests against the team-mode page contracts

* fix(ui): expose agent ids on identity chips and assert them in e2e

Include stable owner IDs in accessible chip labels and tooltips. Preserve same-session-id ownership proof with two identically named agents, including an empty current session, and identify row owners through the shared chip across list tests.

Validate the real-Gateway Usage scenario, sidebar and Usage tests, i18n, docs, and changed checks. Refs #141476.

* feat(ui): team mode replaces the agent chip with a workspace header

Show the configured Gateway name or OpenClaw with the product mark in team
mode. Keep chat context with the open session, preserve the header controls,
and limit the workspace menu to Show one agent, Agent settings, and Help.
Restore keyboard focus when switching between the workspace and agent chip.

Reuse the existing help submenu and header styles, document the mode, and
verify workspace identity, menu contents, context changes, and chip restoration.
Validation includes 376 unit/component tests, all 31 selected browser tests,
i18n, changed-file checks, architecture, and production build/performance.

Refs #141476.

* feat(ui): agent-first team sidebar rows and indicators

Integrate the reviewed team sidebar layout, trailing activity and attention indicators, nested row geometry, and default face avatars. Preserve shared roster lifecycle and identity updates. Refs #141476.

* fix(ui): keep team activity on stable compact session rows

Keep configured team order and move main conversations into session rows, leaving header summaries only for collapsed groups. Keep titles on one line, child state in the right column, Online collapsed initially, and conversation actions clearly labeled.

Move default face artwork into the lazy roster module and remove the obsolete shared preview path and indicator export. Preserve normal session visibility filters when restoring the canonical global stream. Verify desktop and touch geometry, nested rows, identity precedence, filters, and same-id Usage ownership. Refs #141476.

* improve(ui): keep agent creation in Settings

Remove the New agent shortcut and command handler from sidebar menus, keeping conversation creation and agent switching focused on existing agents. Preserve administrator-gated creation in Settings at zero, one, and multiple agents, and the Agents home action. Update menu coverage and the Settings and team-mode documentation. Refs #141476.

Co-authored-by: hannesrudolph <49103247+hannesrudolph@users.noreply.github.com>

* fix(ui): deduplicate trailing sidebar status indicators

Keep collapsed descendant summaries in the trailing state column and suppress status kinds already represented by the parent, preserving distinct queued work and attention. Fill the global conversation menu avatar with the default face. Cover parent/child status overlap and wait for menu animation before measuring geometry. Refs #141476.

* test(ui): keep roster activity fixture updates immutable

Use Object.assign for row copies inside the activity-order regression to satisfy the repository map-copy lint rule without changing fixture behavior.

* fix(ui): share one default agent avatar across surfaces

Share image, identity emoji, and deterministic SVG face rendering across
Control UI agent cards, menus, chips, chat, owners, and participants.
Keep people avatars on their existing profile and initials path.

Select seven crisp silhouettes and ten hues from the stable agent id,
load the artwork lazily, and retire pending faces when identity emoji
arrive. Preserve authenticated image ownership and error fallback.
Resolve full-message transcript artwork with the same effective agent id
as its replies. Replace the tiny mock images with crisp synthetic artwork.

Refs #141476.

* fix(ui): polish team-mode session rows and group actions

Give each agent group New conversation, Open main chat, All sessions,
and Collapse others actions. Keep the global team filter without the
redundant Sessions heading, See all link, or section-level creation menu.

Reserve aligned child-count, unread, and state slots for team rows and
collapsed groups. Prioritize input and error attention, preserve each
expanded row's own activity, and use plain 12px nested carets. Keep
chip-mode controls and leading activity unchanged.

Document the team actions and indicators. Cover keyboard, touch, shared
agent scope, collapse persistence, fixed geometry, and conflict ownership.
Update the stale shared-avatar assertion to the current textAvatar field.

Refs #141476.

* test(ui): follow the New conversation label in e2e

Update the new-session transition, session ownership, navigation, and
message-action tooltip flows to the reviewed sidebar accessible name.
Retain keyboard coverage of the chip-mode global creation link.

Refs #141476.

* fix(ui): keep team-mode titles and names readable

Move team session filters into the sidebar header toolbar and remove the
empty filter row. Size trailing indicators to present state and let quiet
session titles use the remaining row width. Share the collapsed summary
and header-action area so hover and keyboard focus preserve agent names.

Keep chip mode, agent-owned menus, plain carets, nested guides, and row
heights. Preserve the workspace label beside four equally sized controls.
Update sidebar docs and cover 258px title/name geometry, touch, keyboard,
and filter behavior. Shrink the assertion baseline after removing a cast.

Refs #141476.

* fix(ui): keep the product mark for the system agent avatar

Resolve reserved system identities to the mounted, build-versioned OpenClaw
favicon in the shared avatar renderer, without image or generated-face fallback.
Pass the custodian's canonical agent ID instead of threading a page-owned image.
Keep the chat image contract and accessible name while removing its extra wrapper.

Preserve ordinary agent image, emoji, and generated-face behavior. Cover system
identities, image errors, and mounted assets, and document the product-mark rule.

Refs #141476.

* fix(ui): keep configured image avatars beside chat replies

Restore the chat image classes and accessible name for configured avatars,
including images fetched through the authenticated workspace avatar path.
Reuse the existing image slot with shared emoji and generated-face fallbacks,
and keep reserved system agents on the OpenClaw product mark.

Cover image sources, fallback behavior, saved and streaming replies, and
forwarded messages. Document configured workspace images beside replies.

Refs #141476.

* Revert "improve(ui): keep agent creation in Settings"

This reverts commit 0d7a78db4f4cfab2cbee00515b38c30ab316780b.

* docs(ui): preserve current sidebar menu wording

Keep the New conversation labels and no-filter agent switcher description accurate after restoring the New agent shortcut.

* fix(ui): reconcile sidebar roster integration checks

Align avatar recovery tests introduced by main with the shared generated-face
fallback and nested image markup, preserving authenticated loading, revision
recovery, and stale-error coverage. Restore the Usage loader's narrow Gateway
and agent-selection input contract after its lazy-loader extraction.

Validation: reproduce the four failing avatar assertions before the fix;
52 avatar and Usage route tests pass afterward. Independent review is clean.

* fix(ui): align sidebar fixtures and identity e2e with main

---------

Co-authored-by: hannesrudolph <49103247+hannesrudolph@users.noreply.github.com>
2026-09-12 00:50:49 -07:00

18 KiB
Raw Blame History

summary title read_when
Routine local test order, the core test commands, and the local PR gate Run tests locally
You are running or fixing tests on your own machine
You need the local land and gate command list

Routine local order

  1. pnpm test:changed for changed-scope Vitest proof.
  2. pnpm test <path-or-filter> for one file, directory, or explicit target.
  3. pnpm test only when you intentionally need the full local Vitest suite.

The project runner prints wrapper usage for a sole --help or -h request. Compound requests, including --help --no-help, follow native Vitest option semantics.

An existing UI directory target stays scoped to that directory, including when combined with explicit E2E test files. Tests retain their owning shared, isolated, or browser lane. UI source/support-file targets that need whole-lane coverage (such as shared styles or setup files) still use that broader fallback; use a directory or explicit test files when you want a bounded run.

When a one-shot routed run targets only explicit test-file paths, each selected Vitest invocation must discover at least one test file. Excluding every selected file fails even when the lane normally permits empty runs. To allow that outcome intentionally, use pnpm test <test-file-path> -- --passWithNoTests. Use --passWithNoTests=false to require nonempty discovery explicitly. Broader selectors and source-derived selections retain their lane defaults.

An explicit --config run through scripts/run-vitest.mjs keeps its stricter named-file policy and does not permit empty named-file runs. Plugin --allow-no-tests and --allow-empty-after-exclude controls are unchanged.

Codex and other linked/sparse worktrees can run local tests and checks. Tooling requires prepared dependencies; it does not create an implicit link to the primary checkout. Keep any existing borrowed install unchanged while another task uses it. A source install's pnpm:devPreinstall check refuses borrowed links at the checkout-root module directories and their .pnpm directories before normal dependency reconciliation. It does not inspect every workspace package's dependencies or lock paths against concurrent replacement; --ignore-scripts skips it, and alternate pnpm directory settings are not exhaustively validated. When dependencies are ready, use the normal commands above. To avoid pnpm's package-manager preflight against an existing shared install, use these direct Node harnesses:

  • Bounded focused proof with ready dependencies: node scripts/run-vitest.mjs <path-or-filter>.
  • Changed typecheck/lint/guard proof: node scripts/check-changed.mjs.

For Control UI route tests, run node scripts/run-tsgo-core-test-shards.mjs ui to check fixture types; node scripts/run-tsgo.mjs -p tsconfig.ui.json checks production UI code and excludes tests. Type route fixtures against the loader's required capabilities instead of asserting a partial fixture as the full application context. Keep real selection capabilities in lifecycle tests so agent scope changes and subscription cleanup follow the application behavior.

For remote-environment proof, invoke node scripts/crabbox-wrapper.mjs directly. Avoid local pnpm crabbox:run in linked worktrees because pnpm may reconcile dependencies before the remote wrapper starts.

Core commands

Run the test toolchain on Node 24.16+ or Node 26.1+, matching the packaged runtime floor. Older Node bindings can truncate SQLite TEXT values at embedded NUL characters.

The test toolchain pins stable Vitest 5.0.0, including its browser and coverage packages. Use describe(name, { concurrent: false }, callback) for ordered suites. Await asynchronous assertions, keep vi.mock/vi.hoisted at module scope, and perform actions whose mock calls you assert inside the test. OpenClaw sets clearMocks: false, so setup and beforeAll calls are preserved. Clear or reset each assertion's owned mock actions explicitly as needed. Name patterns spanning suites use suite > test; native JSON retains its space-joined fullName, so evidence readers match ancestorTitles plus title.

Filesystem transform caching uses test.fsModuleCache and test.fsModuleCachePath; the existing OPENCLAW_VITEST_FS_MODULE_CACHE and OPENCLAW_VITEST_FS_MODULE_CACHE_PATH controls retain their ownership and disable behavior. Cache-key plugins use defineCacheKeyGenerator. Inline projects inherit root configuration in Vitest 5, including concatenated setup and include arrays. The four UI E2E resource projects declare extends: false because each supplies its complete inventory and setup.

Maintained JavaScript tooling wrappers and root package commands load TypeScript through scripts/tsx.mjs, using tsx's ESM entry. This preserves native loading of compiled ESM plugins and their import-only dependencies, including when loaded through require(). Source TypeScript imports and tsconfig path aliases remain available.

These launchers retain tsx's in-process transform cache and Node's module cache. They skip tsx's shared disk cache before the loader starts, and child tooling inherits that policy. This cache policy does not clean existing temporary directories, Node or Vitest caches, or other global caches. Standalone pnpm ui:build keeps native startup and applies the same preload to its post-build validators; it does not require TSX_DISABLE_CACHE in the invoking shell. Raw external tsx and node --import tsx invocations outside these launchers are unchanged.

Parallel project runs on macOS and Linux reuse filesystem transforms within exclusive worker slots, with separate directories for each Vitest configuration. A slot stays owned through preflight, retries, and verified child/group completion; uncertain cleanup retires it. Explicit isolated cache paths, serial and watch runs, and Windows retain their existing cache ownership. Concurrent invocations still need separate cache roots.

Control UI builds report size budgets without enforcing them. Run pnpm ui:check-performance after a build to enforce absolute budgets, or pnpm ui:check-performance:base <base-commit-sha> to build and compare both revisions with the same toolchain. See Control UI size budgets.

Source tests and subprocess builds

Non-watch runs through pnpm test or scripts/run-vitest.mjs keep Vitest tests and runtime parents on TypeScript. Importing a declared subprocess entrypoint compiles the fixed test entry set and its workspace dependencies into one fresh invocation directory under .artifacts/vitest-workers/.

The declared application entries run as plain Node JavaScript without a TypeScript loader: SQLite read-only snapshots, database verification, Tailscale route ownership, the service relay, its POSIX and Windows anchors, the memory plugin's KNN child, session transcript archive and reconciliation workers, and managed GitHub credential resolution. The same generation also compiles the fake-backend TUI fixture's four runtime roots together: the real TUI, embedded reply producer, reply metadata reader, and outbound normalizer. Shared chunks preserve their module and WeakMap identity. Generated TUI fixtures remain .mts files: Node launches them with --import tsx for their own syntax, while Bun handles that syntax natively without the Node loader. Only their runtime imports change. Existing package build entry paths and Vitest source parents stay unchanged. The CLI fork-recovery regression also compiles the real CLI entry and its concurrent rebind's session accessor and binding helper together. Both processes use the same runtime graph while retaining the durable-write race and process-exit assertions. Doctor process output tests with bundled plugins disabled reuse that compiled CLI inside one lazily created package fixture per test run, keeping real UI checks on fixture-owned assets and each scenario’s state separate. Standalone and watch runs use live source inside the same fixture.

Isolated Doctor config scripts also share the prepared config-flow, health-writer, and install-index modules. Each case still starts a fresh process with separate state; standalone and watch runs resolve the original TypeScript entrypoints.

The model-catalog and session model-context workers also use this compiled generation. Model-catalog workers still belong to their prepared model generations; context reads retain their serial worker pool. Plugin source/built selection remains independent of worker compilation. Other worker-thread entries and arbitrary source CLI fixtures remain outside this declared set.

The session-title and child-link retention tests declare their title-reader, session-utils, and listing roots in this same generation. Each fresh heap-measurement child runs their JavaScript without spending its execution deadline on TypeScript imports.

Automatic-triage process fixtures share this generation for admission, failure handling, execution, process identity, and respawn checks. Compilation finishes before readiness deadlines begin, so children load prepared JavaScript. The detached helper uses the same sealed lease runtime as the installed package.

Preparation is lazy across both projects and shards. Config imports, listing tests, and tiny tests that do not import these declarations do not load the subprocess compiler or compile workers. A shard that needs a declaration requests the outer runner's single build through its existing Node IPC channel during module collection, before fixture hooks and readiness deadlines. Every finite invocation that needs a declaration pays for this fixed entry set; preparation timing is reported separately from child execution. The runner starts one short-lived native Node or Bun compiler child and joins it before returning the verified manifest to borrowers. The compiler module graph lives in that child, not the long-lived runner or Vitest worker. No shard can select a different build graph or adopt another invocation's output. The outer runner retains the generation until child close and process-group cleanup finish, then verifies it before reporting success. Verification reads every recorded input and output with bounded asynchronous I/O, keeping the runner responsive during large shutdown scans. The invocation owner verifies each borrower's preparation before replying and verifies again after all borrowers close; Vitest does not repeat these scans inside each shard or during its concurrent pool shutdown. Standalone Vitest and watch runs retain source execution: compilation, verification, and artifact deletion require the repository runner's ownership. A lost owner or failed build fails the run. Disposal cancels pending compilation and joins it, every borrower, and outstanding preparation requests before asynchronously removing the directory. Signal handlers remain active through removal, even when large generations take time to delete. Borrower completion does not wait for compilation, so an early child exit can reach that cancellation path. An uncertain compiler or borrower join retains the generation and fails the run. Abnormal termination can also leave an unused directory; later runs never adopt it.

Every preparation compiles current source; checkout dist/ is neither an input nor a fallback. Build errors, missing artifacts, and changes to recorded build inputs fail the run. Compilation includes the native subprocess fixtures before they impose resource limits. Third-party dependencies remain external except for the always-bundled OpenClaw packages. fs-safe remains external so its native loader resolves the optional platform package from fs-safe's own dependency scope, including nested pnpm installs. Compiled workers use that same installed package; they do not copy native binaries. The default stays off, and the existing off/auto/require opt-ins retain their behavior. Sealed portable worker bundles use guarded JavaScript only and explicitly disable native loading.

Watch mode deliberately keeps the existing live-source path, including tsx for Node subprocesses and native TypeScript handling for Bun. It creates no prepared generation, so a new child launch reads current source rather than reusing a compiled snapshot. Existing Vitest watch dependency tracking still determines when tests rerun.

Test wrapper runs end with a short [test] passed|failed|skipped ... in ... summary; Vitest's own duration line stays the per-shard detail.

A failed invocation ends with one [test] FAILED (exit N) line after child processes, cleanup, and report publication settle. Direct run-vitest.mts calls use [vitest] instead. Nested runners retain their diagnostics and exit status; the top-level CLI owns the final failure line. Successful runs emit no failure trailer.

Command What it does
pnpm test Explicit file/directory targets route through scoped Vitest lanes. Untargeted runs are full-suite proof: fixed shard groups expand to leaf configs for local parallel execution, with the expected shard fanout printed before starting. The extension group always expands to per-extension shard configs instead of one giant root-project process.
pnpm test:changed Cheap smart changed-test run: precise targets from direct test edits, sibling *.test.ts files, explicit source mappings, and the local import graph. Broad/config/package changes are skipped unless they map to precise tests.
OPENCLAW_TEST_CHANGED_BROAD=1 pnpm test:changed Explicit broad changed-test run; use when a test harness/config/package edit should fall back to Vitest's broader changed-test behavior.
pnpm test:force Frees the configured OpenClaw gateway port (default 18789), then runs the full suite with an isolated gateway port so server tests do not collide with a running instance.
pnpm test:coverage Emits an informational V8 coverage report for the default unit lane (vitest.unit.config.ts); no coverage thresholds are enforced.
pnpm test:coverage:changed Unit coverage only for files changed since origin/main.
pnpm changed:lanes Shows the architectural lanes triggered by the diff against origin/main.
pnpm check:changed Runs the local changed formatting/typecheck/lint/guard plan, including targeted Vitest owner tests for selected paths. Use pnpm test:changed or pnpm test <target> for additional test proof matching the touched contract.

pnpm check:changed also runs the mobile protocol-event coverage guard when changes affect the gateway event catalog or constants, scanned mobile sources, coverage declarations, or the guard, its execution helpers, and its routing. All-lane checks include it too. Every gateway event must have a handler or an explicitly approved non-consumption declaration for each mobile client. To run only this guard, use pnpm check:protocol-coverage.

For native app changes, pnpm check:changed uses platform scope to select lint: Android selects pnpm android:lint (the Gradle ktlint checks), while Apple app changes retain Swift lint. Android-only changes do not select Swift lint or its missing-tool notice. Android framework/resource lint and runtime tests remain separate checks; Kotlin lint does not replace them.

Remote filesystem fixtures that execute GNU stat and readlink run locally only on Linux. The shared leading-@ file-tool scenario also runs against a portable remote-only bridge on every platform. Native Python helper coverage remains separate, including macOS; these fixture gates do not restrict the SSH backend's Gateway host.

Tests that discover real bundled provider runtimes declare that prerequisite in scripts/lib/vitest-build-prerequisites.mts, including Telegram sticker-model selection. Local runners and CI prepare those artifacts before admitting workers.

Local PR gate

For local PR land/gate checks, run:

  • pnpm check:changed
  • pnpm check
  • pnpm check:test-types
  • pnpm build
  • pnpm test
  • pnpm check:docs

If pnpm test flakes on a loaded host, rerun once before treating it as a regression, then isolate with pnpm test <path/to/test>. For memory-constrained hosts:

  • OPENCLAW_VITEST_MAX_WORKERS=1 pnpm test
  • OPENCLAW_VITEST_FS_MODULE_CACHE_PATH=/tmp/openclaw-vitest-cache pnpm test:changed