openclaw/qa
Peter Steinberger 4f0541eb46
fix: prevent delegated work from stopping silently (#162567)
* fix: prevent delegated work from stopping silently

* fix: preserve frozen bootstrap metadata and refresh task prompts

* test: retain a foreign manager in frozen hydration proof

* test: exercise Corepack bootstrap warnings during hydration

* test: align completion fixtures with required private replies

Verify meaningful private outcomes, retained completion authority after inline waits, and one real write with no recovery replay. Start the browser fixture completion deadline when its held model is released.

* fix(node): recheck upload cancellation after opening snapshots

* test: group subagent typechecks with session ownership
2026-10-01 18:39:01 +00:00
..
convex-credential-broker chore(deps): refresh dependencies through September 19 cutoff (#159401) 2026-09-27 19:33:08 -07:00
scenarios fix: prevent delegated work from stopping silently (#162567) 2026-10-01 18:39:01 +00:00
frontier-harness-plan.md
maturity-scores.yaml refactor: remove Tasks and TaskFlow runtime (#159179) 2026-09-27 10:40:29 -07:00
README.md fix(qa): await completed memory replies before assertions and reset (#160929) 2026-09-29 01:52:38 -05:00
scenarios.md

QA Scenarios

Seed QA assets for the private qa-lab extension.

Files:

  • scenarios/index.yaml - canonical QA scenario pack, kickoff mission, and operator identity.
  • scenarios/<theme>/*.yaml - one runnable scenario per YAML file.
  • frontier-harness-plan.md - big-model bakeoff and tuning loop for harness work.
  • convex-credential-broker/ - standalone Convex v1 lease broker for pooled live credentials.

Key workflow:

  • qa suite is the executable frontier subset / regression loop.
  • qa manual is the scoped personality and style probe after the executable subset is green.
  • qa coverage prints the scenario coverage inventory from scenario YAML.

Operator workflows:

  • Use the openclaw-qa-testing skill for QA Lab live lanes, Convex credential pool operations, and WhatsApp live credential setup/replacement.

Keep this folder in git. Add new scenarios here before wiring them into automation.

QA-channel flows that assert final replies use waitForCompletedQaReply with the message returned by sendInbound. It waits for the channel's processing acknowledgment before reading the retained reply. QA-channel reset also waits for pending inbound turns before clearing observations.

Generated-media scenarios count attachment deliveries separately from text progress and check the saved bytes plus the persisted completion reply.

Confined repository checkpoint commands

scripts/qa/repository-checkpoint-admission.ts adapts the product-owned checkpoint and publication command planners to the final campaign Git launch boundary. It executes only the canonical bare checkpoint initialization, checkpoint Git-directory query, and read-only publication config probes. All other commands exit 126 before spawning Git.

The launcher supplies freshly validated checkpointRoot (from the repository workspace store), nodeRoot (from the current placement), campaignRoot, cwd, argv (including git), and env. Paths must already exist, be canonical, and remain inside the isolated campaign. The adapter does not discover sessions, read state databases, authorize transport, or manage checkpoint contents. Those responsibilities stay with their owners. Do not turn a denied command into a generic Git write allowance.

Pass that JSON object on stdin to node --import ./scripts/tsx.mjs scripts/qa/repository-checkpoint-admission.ts. The launcher validates those current roots and exact canonical argv immediately before spawning Git, sanitizes its environment, and forwards Git's output/status. It uses the campaign launcher's PATH, not the requested command's PATH. Delegate the process to this entry rather than treating its exit status as permission to spawn Git again. Freeze it with the campaign tooling; preserve the enclosing sandbox, freshly validated current-owner facts, and execution evidence. Historical artifact shims that consumed the old Boolean interface must migrate to this delegated launch; they are not maintained campaign entry points.

Script checkout and artifact roots

The maintained test-file runner expands ${repoRoot} to the selected checkout and ${outputDir} to the scenario's artifact directory. Script scenarios can pass these as distinct arguments without embedding a campaign path or SHA.

Holding a mock provider continuation

Private QA Lab flows using mock-openai can call env.mock.holdNextContinuation(sessionId, signal) before starting a turn. Pass the flow's existing signal; only one hold may be active per provider. Await hold.reached for { cursor, sessionId, toolOutputCallId }, then use hold.release() to send the planned response or hold.cancel() to discard it. The checkpoint identifies the next tool continuation for the exact transport session, after request recording and before any response bytes. Compaction and other sessions do not consume it. The cursor indexes the existing /debug/requests log for this provider lifetime; it is not a durable Gateway run or receipt ID. Collect those identities separately through their authoritative interfaces.

A hold captures one request. Reconnected/replacement requests proceed normally; cancel the old hold in the scenario's finally after interrupting its caller. Scenario cancellation and provider shutdown also settle pending holds. This control is absent from model tools, Gateway RPC and ordinary plugin SDK exports.