openclaw/qa
Ayaan Zaidi a71e5566d7
fix(approvals): let OpenClaw change approvals complete in chat (#157947)
## What Problem This Solves

Fixes: when an agent asks OpenClaw to change config, restart, or manage plugins/channels/agents, the approval never reaches chats without a native approval card, and `/approve <id>` rejects it. The agent's tool call then blocks until the 10-minute expiry, and users are sent to the Control UI or a terminal to approve a change they asked for in chat.

## User Impact

User impact: every approval an agent needs can be completed in the chat where it was requested.
- Native approval cards (Telegram, Discord, Slack, WhatsApp, Signal, Matrix, Teams, iMessage, Google Chat) keep owning their chats.
- Every other requesting chat now gets the change summary with **Allow once** / **Deny** buttons where the channel renders them, plus a `/approve <id> allow-once|deny` line.
- `/approve` now resolves OpenClaw change approvals, not only exec and plugin approvals.

No config options, protocol, or schema changes.

## Why This Change Was Made

Delegated OpenClaw changes ("system-agent" approvals) already had native cards (#134670). Two gaps remained outside those cards:

- **Request delivery.** The request was created with delivery turned off, so the shared approval forwarder never posted a fallback. The forwarder now has a system-agent strategy that always targets the requesting chat and is suppressed whenever that chat's native card handles the request. It uses the same typed-button payload and resolved/expired messages as exec and plugin approvals. The system-agent owner publishes the outcome of a decision (applied, or denied) once; approval publication only adds expiry and cancellation, so a denial produces one chat update. The fallback answers only the live messaging chat that made the request: terminal and Webchat requests never fall back to the session's saved chat, and expiry is reported once, from the Gateway's recorded expiry rather than a local chat timer, so a change approved just before the deadline and applied after it reports its applied outcome instead of a false expiry.
- **`/approve`.** The command probed only exec and plugin approvals. It now also resolves system-agent approvals through the canonical `approval.resolve`, after confirming the id is in `openclaw.approval.list`. Canonical resolution records a kind mismatch as a deny, so `/approve` confirms the owner first instead of probing. On channels with their own approver settings (Telegram, Discord, Slack, …), the Gateway checks the sender as the reviewer, the same check as native buttons. Everywhere else only a configured owner (`commands.ownerAllowFrom`) can approve an OpenClaw change: `/approve` sends the sender as the reviewer, and the Gateway checks owner custody against the current config inside the approval store's final decision guard. An ordinary command-authorized sender is refused, and an owner removed after sending `/approve` cannot complete the decision.

Unchanged: approval authority stays bound to the requesting run; free-text "yes" never approves; Full Access still auto-applies; Control UI and the apps can still decide.

Agent-facing text now says what happens: for runs from messaging channels, the `openclaw` tool description says the change waits for approval in this chat (buttons or `/approve`). Webchat and terminal runs, which the chat fallback cannot reach, are told to approve in the Control UI or OpenClaw apps. The gateway-only prompt line points to `openclaw` and `/restart` instead of "ask human".

## Evidence

Real Telegram (Test Server userbot, Convex-leased credentials, fresh Gateway, mock provider), restricted `exec` mode, agent asks OpenClaw to `set logging.level "info"`:

| Run | Setup | What the user saw | Tap | Result |
|---|---|---|---|---|
| A | tester is owner (native cards on) | 🔒 native approval card, no `/approve` text | **Allow Once** | `answerCallbackQuery` + 2× `editMessageText`: "approved. Applying" → "approved and applied"; final reply delivered |
| C | tester is owner, `channels.telegram.execApprovals.enabled: false` | fallback message with change summary, `/approve <id> allow-once\|deny`, and **Allow Once** button | **Allow Once** | message edited to "approved. Applying…", then "approved and applied" posted; Gateway config on disk has `logging.level: "info"`; final reply delivered |

Telegram Web screenshots from the same leased test account (cropped to the conversation):

| | Pending | After **Allow Once** |
|---|---|---|
| Native card (A) | ![Native Telegram approval card with Allow Once and Deny buttons](https://github.com/user-attachments/assets/5d3de39a-339c-4b7b-80d5-84f5d46cc575) | ![Native card edited to approved and applied, followed by the final reply](https://github.com/user-attachments/assets/3e993ffe-3f85-4941-9f31-0904719ea608) |
| Fallback message (C) | ![Fallback approval message with change summary, /approve line, and Allow Once and Deny buttons](https://github.com/user-attachments/assets/ca78d61c-627e-47b0-9505-66fa7162e540) | ![Fallback message resolved: approved and applying, then approved and applied, then the final reply](https://github.com/user-attachments/assets/fb4b5e28-074b-4f69-bbcb-053398b94c25) |

With no owner configured, the owner-only `openclaw` tool isn't exposed, so no change is proposed (config unchanged). That's the existing design; DM pairing sets the first owner.

qa-channel (no native approval cards), new scenario `system-agent-chat-approval`: the agent's delegated change posts the approval in the requesting conversation. A command-authorized non-owner (`commands.allowFrom` includes them, `ownerAllowFrom` does not) sends `/approve <id> allow-once` and gets "❌ Only the owner can approve OpenClaw changes in this chat."; `logging.level` is unchanged and the approval stays pending. The owner's `/approve <id> allow-once` then resolves it; `logging.level` becomes `info`; exactly one final reply. On `origin/main` the same scenario times out waiting for the approval in the chat.

Tests: `/approve` (resolves a pending OpenClaw change canonically; never submits a canonical decision for an id owned by another kind; on a channel without approver settings, a non-owner and a revoked owner never reach the canonical decision), channel custody (without approver settings only a configured owner holds custody, and only for OpenClaw changes), `approval.resolve` (custody revoked between the request and the final write leaves the approval pending), approval publication (allowed/denied changes leave the chat outcome to the system-agent owner; expired/cancelled are published once), forwarder (requesting chat gets the prompt and outcome without `approvals.*` config; a running native card suppresses it; terminal and Webchat requests never reach the saved session chat; a change applied after the deadline reports its outcome with no false expiry; the recorded expiry is reported once), plus existing gateway approval and system-agent suites. The new owner-custody and single-publisher tests fail with the fix reverted. Shared approval and forwarder fixtures moved to sibling `*.test-support.ts` modules so the test files stay under the line cap. Wall time for the touched suites: `commands-approve` + `approval-publication` + `exec-approval-forwarder` + `system-agent-approval` + `approval` ran in 59.5s across 3 Vitest shards. `pnpm tsgo:core` clean.

Co-authored-by: Ayaan Zaidi <hi@obviy.us>
2026-09-25 15:05:57 +05:30
..
convex-credential-broker chore(deps): refresh dependencies with a seven-day cutoff (#157238) 2026-09-25 02:38:45 +00:00
scenarios fix(approvals): let OpenClaw change approvals complete in chat (#157947) 2026-09-25 15:05:57 +05:30
frontier-harness-plan.md
maturity-scores.yaml docs: update maturity scorecard (#153186) 2026-09-20 10:41:45 -07:00
README.md
scenarios.md

QA Scenarios

Seed QA assets for the private qa-lab extension.

Files:

  • scenarios/index.yaml - canonical QA scenario pack, kickoff mission, and operator identity.
  • scenarios/<theme>/*.yaml - one runnable scenario per YAML file.
  • frontier-harness-plan.md - big-model bakeoff and tuning loop for harness work.
  • convex-credential-broker/ - standalone Convex v1 lease broker for pooled live credentials.

Key workflow:

  • qa suite is the executable frontier subset / regression loop.
  • qa manual is the scoped personality and style probe after the executable subset is green.
  • qa coverage prints the scenario coverage inventory from scenario YAML.

Operator workflows:

  • Use the openclaw-qa-testing skill for QA Lab live lanes, Convex credential pool operations, and WhatsApp live credential setup/replacement.

Keep this folder in git. Add new scenarios here before wiring them into automation.