openclaw/docs/help/testing/live-workflows.md
Peter Steinberger ed5937fbd9
Some checks are pending
ClawSweeper Dispatch / dispatch (push) Waiting to run
CodeQL / Security High (actions) (push) Waiting to run
CodeQL / Security High (channel-runtime-boundary) (push) Waiting to run
CodeQL / Security High (core-auth-secrets) (push) Waiting to run
CodeQL / Security High (mcp-process-tool-boundary) (push) Waiting to run
CodeQL / Security High (network-ssrf-boundary) (push) Waiting to run
CodeQL / Security High (plugin-trust-boundary) (push) Waiting to run
CodeQL / Security High (process-exec-boundary) (push) Waiting to run
Control UI Locale Refresh / resolve-base (push) Waiting to run
Control UI Locale Refresh / Verify generated PR App permissions (push) Blocked by required conditions
Control UI Locale Refresh / Refresh (push) Blocked by required conditions
Control UI Locale Refresh / Commit control UI locale refresh (push) Blocked by required conditions
Docs Sync Publish Repo / sync-publish-repo (push) Waiting to run
Docs / docs (push) Waiting to run
Native App Locale Refresh / Refresh native th (push) Blocked by required conditions
Native App Locale Refresh / Refresh native tr (push) Blocked by required conditions
Native App Locale Refresh / Refresh native uk (push) Blocked by required conditions
Native App Locale Refresh / Refresh native vi (push) Blocked by required conditions
Native App Locale Refresh / Refresh native zh-CN (push) Blocked by required conditions
Native App Locale Refresh / Refresh native zh-TW (push) Blocked by required conditions
Native App Locale Refresh / Commit native locale refresh (push) Blocked by required conditions
Native App Locale Refresh / Refresh native nl (push) Blocked by required conditions
Native App Locale Refresh / Refresh native pl (push) Blocked by required conditions
Native App Locale Refresh / Refresh native pt-BR (push) Blocked by required conditions
Native App Locale Refresh / Refresh native ru (push) Blocked by required conditions
Native App Locale Refresh / Refresh native sv (push) Blocked by required conditions
Native App Locale Refresh / resolve-base (push) Waiting to run
Native App Locale Refresh / Verify generated PR App permissions (push) Blocked by required conditions
Native App Locale Refresh / Refresh native ar (push) Blocked by required conditions
Native App Locale Refresh / Refresh native de (push) Blocked by required conditions
Native App Locale Refresh / Refresh native es (push) Blocked by required conditions
Native App Locale Refresh / Refresh native fa (push) Blocked by required conditions
Native App Locale Refresh / Refresh native fr (push) Blocked by required conditions
Native App Locale Refresh / Refresh native hi (push) Blocked by required conditions
Native App Locale Refresh / Refresh native id (push) Blocked by required conditions
Native App Locale Refresh / Refresh native it (push) Blocked by required conditions
Native App Locale Refresh / Refresh native ja-JP (push) Blocked by required conditions
Native App Locale Refresh / Refresh native ko (push) Blocked by required conditions
Node Runtime Conformance / TypeScript contracts (push) Waiting to run
Node Runtime Conformance / Rust workspace (push) Waiting to run
OpenClaw Stable Main Closeout / Resolve stable release closeout inputs (push) Waiting to run
OpenClaw Stable Main Closeout / Verify stable main closeout (push) Blocked by required conditions
Plugin Init Scaffold Validation / Validate provider scaffold (push) Waiting to run
Plugin NPM Release / verify_plugins_npm (push) Blocked by required conditions
Plugin NPM Release / preview_plugins_npm (push) Waiting to run
Plugin NPM Release / Validate release publish approval (push) Blocked by required conditions
Plugin NPM Release / preview_plugin_pack (push) Blocked by required conditions
Plugin NPM Release / Preflight plugin npm package () (push) Blocked by required conditions
Plugin NPM Release / Seal prepared plugin npm release (push) Blocked by required conditions
Plugin NPM Release / Trusted publisher OIDC exchange (push) Blocked by required conditions
Plugin NPM Release / Publish plugin npm package () (push) Blocked by required conditions
Vitest Cache Warm / dependencies (push) Waiting to run
Vitest Cache Warm / warm (push) Waiting to run
Workflow Sanity / no-tabs (push) Waiting to run
Workflow Sanity / actionlint (push) Waiting to run
Workflow Sanity / generated-doc-baselines (push) Waiting to run
improve: profile concurrent Gateway sessions with live OpenAI (#154845)
* improve: profile live Gateway concurrency

* fix: align Gateway benchmark with CI contracts
2026-09-21 13:33:16 +00:00

8.2 KiB
Raw Blame History

summary title read_when
Live provider debugging lanes plus the Docker and Parallels smokes that support them Live and Docker/Parallels workflows
You are debugging a real provider or model
You need a live Docker or Parallels lane

Live and Docker/Parallels workflows

When debugging real providers/models (requires real creds):

  • Live suite (models + gateway tool/image probes): pnpm test:live
  • Target one live file quietly: pnpm test:live -- src/agents/models.profiles.live.test.ts
  • Progress-card refresh: OPENCLAW_LIVE_TEST=1 pnpm test:live -- src/gateway/gateway-progress-refresh.live.test.ts
    • Requires OPENAI_API_KEY and uses openai/gpt-5.6-luna with isolated Gateway state.
    • Completes an earlier turn, then refreshes during a second turn while a command remains held. The original parent must update the card and retain its final reply. A later idle refresh must update the card without adding chat messages or resuming pending work.
  • Live subagent handoff stress: OPENCLAW_LIVE_TEST=1 OPENCLAW_LIVE_SUBAGENT_STRESS=1 pnpm test:live -- src/agents/subagents/announce/subagent-yield-resume.live.test.ts
    • Requires OPENAI_API_KEY and defaults to openai/gpt-5.6-luna; select another OpenAI model with OPENCLAW_LIVE_SUBAGENT_E2E_MODEL.
    • Pins the OpenClaw agent harness and uses isolated Gateway state, synthetic files, and externally held HTTP responses. It checks concurrent hidden-result fanout, status-only interrogation of a waiting child tree, operator resume with preserved task identity and idempotent replay, child timeout, a live HTTP 503 retrieval failure, and cancellation racing an in-flight result after its owner claims the run. Parent reports are checked against held requests, child execution, task delivery state, and hidden results. The 503 case checks that a completed agent turn does not imply a successful retrieval.
    • Each case incrementally preserves bounded runtime facts, plus final assistant replies, in .artifacts/qa-e2e/subagent-challenges-*/evidence.json, including failed runs. Set OPENCLAW_LIVE_SUBAGENT_EVIDENCE_DIR to change the output directory. This live lane does not simulate cold process loss; restart and lost-acceptance ownership are covered by subagent-orphan-recovery.restart-integration.test.ts.
    • Defaults to two batches of three children. Set OPENCLAW_LIVE_SUBAGENT_STRESS_BATCHES (1–5) and OPENCLAW_LIVE_SUBAGENT_STRESS_CHILDREN (1–6) to change the bounded workload.
  • Live Gateway concurrency: dispatch OpenClaw Performance on main with mode=gateway-concurrency and live_openai_candidate=true. It runs 96 real turns across 32 agents with 1,000 seeded sessions, concurrent session activity, and load-phase CPU profiles. Dreaming is disabled while ordinary indexing and recaps remain enabled. Results stay in Actions artifacts. See Gateway concurrency benchmark.
  • Runtime performance reports: dispatch OpenClaw Performance with live_openai_candidate=true for a real openai/gpt-5.6-luna agent turn or deep_profile=true for Kova CPU/heap/trace artifacts. Daily scheduled runs publish mock-provider, deep-profile, and GPT-5.6 Luna lane reports to openclaw/clawgrit-reports from a separate artifact-consuming publisher job; missing or invalid publisher authentication fails scheduled and profile=release runs. Manual non-release dispatches keep the GitHub artifacts and treat report publication as advisory. The mock-provider report also includes source-level gateway boot, memory, plugin-pressure, repeated fake-model hello-loop, and CLI startup numbers.
  • Docker live model sweep: pnpm test:docker:live-models
    • Each selected model runs a text turn plus a small file-read-style probe. Models whose metadata advertises image input also run a tiny image turn. Disable the extra probes with OPENCLAW_LIVE_MODEL_FILE_PROBE=0 or OPENCLAW_LIVE_MODEL_IMAGE_PROBE=0 when isolating provider failures.
    • CI coverage: daily OpenClaw Scheduled Live And E2E Checks and manual OpenClaw Release Checks both call the reusable live/E2E workflow with include_live_suites: true, which includes Docker live model matrix jobs sharded by provider.
    • For focused CI reruns, dispatch OpenClaw Live And E2E Checks (Reusable) with include_live_suites: true and live_models_only: true.
    • Add new high-signal provider secrets to scripts/ci-hydrate-live-auth.sh plus .github/workflows/openclaw-live-and-e2e-checks-reusable.yml and its scheduled/release callers.
  • Native Codex bound-chat smoke: pnpm test:docker:live-codex-bind
    • Runs a Docker live lane against the Codex app-server path, binds a synthetic Slack DM with /codex bind, exercises /codex fast and /codex permissions, then verifies a plain reply and an image attachment route through the native plugin binding instead of ACP.
  • Codex app-server harness smoke: pnpm test:docker:live-codex-harness
    • Runs gateway agent turns through the plugin-owned Codex app-server harness, verifies /codex status and /codex models, and by default exercises image, cron MCP, sub-agent, and Guardian probes. Disable the sub-agent probe with OPENCLAW_LIVE_CODEX_HARNESS_SUBAGENT_PROBE=0 when isolating other failures. For a focused sub-agent check, disable the other probes: OPENCLAW_LIVE_CODEX_HARNESS_IMAGE_PROBE=0 OPENCLAW_LIVE_CODEX_HARNESS_MCP_PROBE=0 OPENCLAW_LIVE_CODEX_HARNESS_GUARDIAN_PROBE=0 OPENCLAW_LIVE_CODEX_HARNESS_SUBAGENT_PROBE=1 pnpm test:docker:live-codex-harness. This exits after the sub-agent probe unless OPENCLAW_LIVE_CODEX_HARNESS_SUBAGENT_ONLY=0 is set.
  • Codex on-demand install smoke: pnpm test:docker:codex-on-demand
    • Installs the packaged OpenClaw tarball in Docker, runs OpenAI API-key onboarding, and verifies the Codex plugin plus @openai/codex dependency were downloaded into the managed npm project root on demand.
  • Codex npm-plugin live package smoke: pnpm test:docker:live-codex-npm-plugin
    • Installs the candidate OpenClaw package and exact Codex plugin into Docker, then uses a real OpenAI key for CLI preflight and same-session turns.
    • Its zero-retry medium-thinking follow-through turn must send progress, keep working through randomized workspace reads and an exact artifact write, then send completion. A progress-only terminal turn fails the lane.
  • Live plugin tool dependency smoke: pnpm test:docker:live-plugin-tool
    • Packs a fixture plugin with a real slugify dependency, installs it through npm-pack:, verifies the dependency under the managed npm project root, then asks a live OpenAI model to call the plugin tool and return the hidden slug.
  • OpenClaw rescue command smoke: pnpm test:live:system-agent-rescue-channel
    • Opt-in belt-and-suspenders check for the message-channel rescue command surface. Exercises /openclaw status, queues a persistent model change, replies /openclaw yes, and verifies the audit/config write path.
  • OpenClaw first-run Docker smoke: pnpm test:docker:system-agent-first-run
    • Starts from an empty OpenClaw state dir and first proves the packaged openclaw setup CLI fails closed without inference. It then tests and activates fake Claude through the packaged activation module. Only afterward does a fuzzy packaged CLI request reach the planner and resolve to typed setup, followed by one-shot model, agent, Discord config, and SecretRef operations. It validates config and audit entries. This is supporting gate/operation evidence, not an interactive onboarding or OpenClaw agent/tool/approval proof. The same lane is exposed in QA Lab by pnpm openclaw qa suite --scenario system-agent-ring-zero-setup.
  • Moonshot/Kimi cost smoke: with MOONSHOT_API_KEY set, run openclaw models list --provider moonshot --json, then run an isolated openclaw agent --local --session-id live-kimi-cost --message 'Reply exactly: KIMI_LIVE_OK' --thinking off --json against moonshot/kimi-k2.6. Verify the JSON reports Moonshot/K2.6 and the assistant transcript stores normalized usage.cost.
When you only need one failing case, prefer narrowing live tests via the [allowlist env vars](/help/testing/docker).