Commit graph

1045 commits

Author SHA1 Message Date
ytahdn
19fc52aa93
feat(daemon): add stateless generation SSE (#6947)
* feat(daemon): add stateless generation SSE

* test(integration): expect session generation capability

* fix(daemon): address generation review findings

* fix(daemon): harden generation regressions

* fix(daemon): preserve generation error events

---------

Co-authored-by: ytahdn <ytahdn@gmail.com>
Co-authored-by: Shaojin Wen <shaojin.wensj@alibaba-inc.com>
2026-07-16 00:00:08 +00:00
qqqys
859095bc98
feat(channels): support natural memory references (#6952)
Some checks are pending
E2E Tests / E2E Test (Linux) - sandbox:docker (push) Waiting to run
E2E Tests / E2E Test (Linux) - sandbox:none (push) Waiting to run
E2E Tests / E2E Test - macOS (push) Waiting to run
E2E Tests / cron-interactive E2E (nightly) (push) Waiting to run
E2E Tests / web-shell Browser Regression (push) Waiting to run
* feat(channels): plan targeted memory intents

* fix(channels): harden targeted memory planner

* fix(channels): bound serialized memory manifest

* feat(channels): resolve natural memory references

* fix(channels): harden memory intent routing

* docs(channels): explain natural memory references

* fix(channels): type resolved memory intent dispatch

* fix(channels): route Chinese preference filters
2026-07-15 17:46:23 +00:00
顾盼
53ce8a2a40
fix(cua-driver): harden MCP tool reliability (#6968)
* fix(cua-driver): harden MCP tool reliability

* fix(cua-driver): address reliability review findings

* fix(cua-driver): keep config snapshots coherent

* fix(cua-driver): reject empty AX container trees

* chore(cua-driver): prepare v0.7.2 release
2026-07-15 15:40:31 +00:00
BaboBen
f5bdba724e
fix(wecom): prevent requireMention from disabling group chat (#6948)
* fix(wecom): trust group callback mention scope

Fixes #6939

* docs(wecom): clarify mention-scoped behavior
2026-07-15 15:38:21 +00:00
jinye
7a1b182cd1
feat(cli): Add archived session export (#6911)
* feat(cli): add archived session export

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* codex: address PR review feedback (#6911)

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* codex: address PR review feedback (#6911)

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* codex: address PR review feedback (#6911)

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* codex: address PR review feedback (#6911)

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

---------

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>
2026-07-15 12:21:07 +00:00
ytahdn
14993b1cf3
feat(daemon): add immutable session source metadata (#6932)
* feat(daemon): add session source metadata

* test(daemon): update baseline capabilities

---------

Co-authored-by: ytahdn <ytahdn@gmail.com>
2026-07-15 07:35:41 +00:00
jinye
9cb09f4e2e
fix(core): Classify shell timeouts as tool errors (#6864)
* fix(core): classify shell timeouts as tool errors

Refs #6863

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* codex: address PR review feedback (#6864)

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* fix(core): Report no output for sed timeouts

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

---------

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>
Co-authored-by: qwen-code-dev-bot <qwen-code-dev-bot@users.noreply.github.com>
2026-07-15 05:34:39 +00:00
jinye
ca5019968a
fix(web-shell): harden non-primary archive actions (#6912)
Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>
2026-07-15 04:57:07 +00:00
morluto
41157205c1
fix(config): reject fractional session and tool-call limits (#6920)
* fix(config): reject fractional session and tool-call limits

* fix(config): validate persisted session turn limits
2026-07-15 04:56:52 +00:00
pomelo
389a1f9ceb
feat(cli): change default approval mode from default to auto (#6899)
* feat(cli): change default approval mode from default to auto

The default approval mode required manual confirmation for every tool
call, producing dozens of confirmation prompts per task. Auto mode uses
a three-layer filter (workspace edits, read-only allowlist, LLM
classifier) to auto-approve safe operations while still guarding risky
ones.

Untrusted folders are still forced to default mode for safety.

Closes #6898

* fix(cli): keep manual approval in safe and bare modes

Restricted modes (safe/bare) strip permissions, allowlists, MCP servers
and hooks to provide a maximally restrictive session. The new AUTO
default fallback was silently downgrading them to the LLM classifier,
contradicting their lockdown intent. Restore DEFAULT (manual approval)
for these modes while keeping AUTO as the default for normal sessions.

Explicit --approval-mode and --yolo flags still take effect, since they
are resolved before the fallback.

* chore(cli): regenerate settings schema for auto default

Regenerate the VS Code settings schema so the tools.approvalMode default
matches the new auto value (fixes the "settings schema is up-to-date" CI
check). Also add coverage for the serve-mode approval fallback when no
approval mode is configured.

* test(cli): update SettingsDialog snapshots for auto approval default

The settings schema now defaults tools.approvalMode to auto, so the
SettingsDialog renders "Auto" instead of "Ask permissions" for the Tool
Approval Mode field. Regenerate the affected snapshots (10 updated).

* test(core): pin DEFAULT baseline in agent-override tests

These tests exercise createApprovalModeOverride isolation and the
DEFAULT→AUTO rule strip/restore transitions, so they implicitly relied
on the Config constructor defaulting to DEFAULT. Now that the default is
AUTO, pin the baseline explicitly so the tests no longer depend on the
constructor default.

---------

Co-authored-by: pomelo.lcw <pomelo.lcw@alibaba-inc.com>
2026-07-15 04:55:32 +00:00
jinye
4f4387cf57
feat(core): add PDF vision bridge fallback (#6846)
* feat(core): add PDF vision bridge fallback

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* codex: address PR review feedback (#6846)

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* codex: address PR review feedback (#6846)

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* codex: fix CI failure on PR #6846

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* codex: address PR review feedback (#6846)

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* codex: address PR review feedback (#6846)

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* fix(cli): harden vision bridge output handling

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* fix(cli): correct export sanitizer test typing

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* fix(core): disclose selected vision endpoint before egress

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

---------

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>
2026-07-15 03:42:14 +00:00
BaboBen
4b802ca5f9
feat(channels): support DingTalk webhook delivery to direct messages (#6891)
* docs(channels): design DingTalk webhook DM delivery

* docs(channels): translate DingTalk webhook DM design

* feat(channels): support DingTalk webhook direct messages

* fix(channels): isolate DingTalk webhook DM targets

* fix(channels): handle DingTalk direct delivery failures

* fix(channels): reject malformed DingTalk responses
2026-07-15 03:07:30 +00:00
callmeYe
441006b0e1
feat(scripts): add local PR verification gate (#6873)
* feat(scripts): add settings schema check mode

* feat(scripts): add local PR verification runner

* fix(scripts): harden local PR verification

* docs: document local PR verification gate

* fix(scripts): isolate local verification tools

* fix(scripts): scope PR formatting checks

* fix(scripts): skip symlinked PR paths

* fix(scripts): preserve verification gate integrity

* fix(scripts): canonicalize verification temp paths

* fix(scripts): stabilize local PR verification

* fix(scripts): clear built-in test credentials

* fix(scripts): enforce isolated test environment

* fix(scripts): serialize local verification tests

* fix(scripts): address PR verification review

* fix(scripts): preserve review git wrapper environment

* refactor(scripts): avoid step helper shadowing

* fix(scripts): distinguish forwarded child signals

* fix(scripts): preserve relayed signal exit codes
2026-07-15 00:58:17 +00:00
qqqys
fb0239eab4
feat(channels): add structured channel memory management (#6860)
* feat(core): add structured channel memory document

* fix(core): preserve legacy channel memory whitespace

* feat(core): structure channel memory storage

* fix(core): close channel memory races

* test(core): cover channel memory read failures

* feat(channels): parse channel memory item intents

* test(channels): cover memory intent precedence

* feat(channels): manage structured channel memory

* fix(channels): address structured memory review

* feat(cli): wire structured channel memory

* docs(channels): document structured channel memory

* fix(core): harden channel memory persistence

* fix(core): preserve channel memory uniqueness

* fix(channels): prioritize channel memory item updates
2026-07-15 00:44:15 +00:00
ranzhenyu
cf42ab6b7e
feat(acp): expose tool-call preparation lifecycle (#6819)
* feat(acp): expose tool-call preparation lifecycle

Why:
ACP clients receive no signal while providers stream tool arguments, making long calls appear stalled and delaying tool-identity policy decisions.

What:
- attach transient preparation metadata for Anthropic and OpenAI-compatible streams
- emit correlated ACP pending, execution, and discarded lifecycle updates
- preserve normalized call IDs across partial chunks and provider ID reuse
- clear abandoned retry calls and keep cleanup failures from terminating healthy retry/fallback streams
- deduplicate suppressed preparations and protect completed remapped parser buffers
- cover multi-tool Anthropic streams, ID reservation, TodoWrite suppression, retry cleanup, cancellation, and stream failure

Impact:
The metadata is additive and consumed only by ACP. It exposes no partial arguments, is not persisted to conversation history, and does not move permissions, hooks, scheduling, or execution ahead of complete function calls.

Tests:
- Core provider and stream suites: 649 passed
- ACP lifecycle suites: 316 passed
- npm run build
- npm run typecheck
- npm run lint:ci
- changed-file Prettier and git diff checks

Refs: #6775

* fix(acp): stabilize tool preparation lifecycle updates

Why:
- ACP cleanup failures must not convert a successful model stream into a failed prompt.
- A prepared tool call must be updated in place when execution starts instead of creating a second card.

What:
- Preserve the primary stream outcome when preparation cleanup fails and remove duplicate message display finalization.
- Track prepared call IDs so execution starts use tool_call_update, guard empty preparation metadata, and cover late stable IDs.

Impact:
- Ordinary tool calls keep their existing tool_call start frame.
- Streaming parser production behavior is unchanged.

* Update packages/cli/src/acp-integration/session/Session.test.ts

overrides参数在createPreparationResponse中被声明但从未使用——所有11个调用点仅传递callId且toolName. 该as GenerateContentResponse强制类型转换会绕过对始终为空对象的结构化类型检查。

Co-authored-by: qwen-code-ci-bot <qwen-code-ci@service.alibaba.com>

* fix(acp): harden preparation lifecycle handling

Why:
- A malformed test helper prevented the preparation lifecycle suite from compiling.
- Duplicate preparing frames and state cleanup need direct regression coverage.

What:
- Repair the preparation response helper and isolate cleanup warning assertions.
- Suppress duplicate preparing frames and cover terminal cleanup plus missing tool call IDs.

Impact:
- Normal preparation and execution transitions remain unchanged.
- Repeated preparation frames for the same call ID are now ignored.

---------

Co-authored-by: qwen-code-ci-bot <qwen-code-ci@service.alibaba.com>
2026-07-15 00:40:02 +00:00
易良
d7e2892a7c
fix(cli): avoid updating active CLI processes (#6874)
* fix(cli): avoid updating active processes

* fix(cli): close update relaunch gaps

* test(cli): fix standalone update source path

* fix(cli): reset deferred update per relaunch
2026-07-15 00:33:17 +00:00
jinye
c538bd70d2
feat(core): emit liveness heartbeats for silent foreground shell commands (#6876)
* feat(core): emit liveness heartbeats for silent foreground shell commands

Silent foreground commands previously produced no events between spawn
and settle, so ACP gateways and stream-json consumers could not tell a
long-running command from a dead session. The shell tool now emits a
structured ShellProgressData through the existing updateOutput channel
whenever no display update has fired for tools.shell.heartbeatIntervalMs
(default 10s, 0 disables). Heartbeats carry liveness stats only - never
command output - and never enter model context.

Consumers: the ACP session forwards heartbeats as meta-only
tool_call_update frames (gated so a tick racing the settle path cannot
regress status after completion) and records heartbeat span attributes;
stream-json forwards them as tool_progress events behind
includePartialMessages; the TUI scheduler, React hook, and subagent
runtime ignore them so live output views are not replaced by stats
objects.

* docs(design): add silent command heartbeat design doc

* fix(acp): keep tool_call_update heartbeats from breaking in-repo consumers

Codex review of the heartbeat change found that in-repo ACP consumers
did not tolerate the new meta-only in_progress frames. A full sweep of
tool_call_update consumers found three that mishandled them, each now
guarded with a regression test:

- The desktop agent converted every tool_call_update into a terminal
  tool_result, so the first heartbeat would prematurely complete the
  command with an empty result. It now skips in_progress updates.
- DaemonChannelBridge requires kind on tool_call_update and flagged the
  kind-less heartbeat as a malformed-protocol error every interval. It
  now drops kind-less in_progress frames silently.
- The web-shell daemon UI normalizer derived the tool block title from
  _meta.toolName, overwriting the human-readable title on every
  heartbeat. It now drops heartbeat frames outright.

The remaining consumers (VS Code companion, acp-bridge compaction,
session export, daemon TUI adapter) merge updates conditionally and are
heartbeat-safe without changes.

* fix(core): address PR review — heartbeat monotonic gate, guard scope, telemetry

Review round 1 on #6876 (yiliang114, wenshao, chiga0, qwen3.7-max):

- shell.ts: the silent-idle gate now uses the monotonic performance.now()
  clock (via lastOutputPerfTime, falling back to spawn time) instead of the
  Date.now()-based lastUpdateTime, so an NTP step can neither skew the
  payload nor misfire a heartbeat — matching the design doc's monotonic
  commitment. It also keys off actual output arrival rather than the
  throttled display update.
- session-tracing.ts: endToolExecutionSpan now applies caller-supplied
  attributes BEFORE the canonical keys (duration_ms, success, error) so a
  passthrough attribute can never mask the span's own outcome fields.
- desktop qwen-agent.ts: the in_progress drop guard is now scoped to frames
  carrying _meta.shellProgress, matching the daemon bridge and web-shell
  normalizer guards, so a future non-heartbeat in_progress frame is not
  silently swallowed.
- Tests: the desktop regression test now pins result==='done' (previously
  it stayed green even with the guard removed); added a Session.test
  assertion that heartbeat counts reach the tool-execution span attributes.

* fix(acp): align desktop heartbeat guard with normalizer; test kind pass-through

Review round 2 on #6876 (qwen3.7-max via ci-bot):

- The desktop qwen-agent in_progress drop guard was broader than the
  web-shell normalizer's: it dropped any in_progress + shellProgress frame
  regardless of kind, while the normalizer only drops kind-less ones. The
  comment claimed they matched. Added the kind-absent check so the desktop
  guard matches the normalizer exactly — a kind-bearing frame now passes
  through on both platforms (heartbeats emitted by the ACP session never
  carry a kind, so real behavior is unchanged).
- Added pass-through tests on both sides (daemonUi + desktop) asserting an
  in_progress frame WITH a kind normalizes to a tool.update / tool_result
  rather than being dropped, so the load-bearing kind-absent condition is
  no longer only exercised on the drop path.

* fix(channels): scope daemon bridge heartbeat drop to shellProgress frames

Review round 3 on #6876 (qwen3.7-max via ci-bot): the DaemonChannelBridge
heartbeat guard lived in the shared tool_call / tool_call_update case and
dropped ANY kind-less in_progress frame, so a genuinely malformed kind-less
tool_call (status in_progress, no shellProgress) was silently swallowed
instead of reaching emitProtocolError. Gate the drop on _meta.shellProgress
— matching the qwen-agent and web-shell normalizer guards — so real
heartbeats are still dropped while malformed frames are flagged. Added a
regression test for the malformed path.
2026-07-15 00:07:26 +00:00
ytahdn
b59b341a0a
feat(web-shell): add extension management page (#6815)
* feat(daemon): support interactive extension installs

* feat(web-shell): add extension management page

* fix(web-shell): align extension update behavior

* fix(web-shell): polish extension management UI

* fix(extensions): harden interactive operations

* fix(web-shell): address extension review suggestions

* fix(web-shell): refine extension interaction handling

* fix(web-shell): resolve extension operation races

* fix(web-shell): harden extension action admission

* fix(web-shell): surface extension recovery failures

* fix(web-shell): preserve extension card titles

* fix(web-shell): refine extension card layout

* fix(extensions): address operation review findings

* test(extensions): close remaining review gaps

---------

Co-authored-by: 钉萁 <dingqi.jww@alibaba-inc.com>
Co-authored-by: ytahdn <ytahdn@gmail.com>
Co-authored-by: qwen-code-dev-bot <qwen-code-dev-bot@users.noreply.github.com>
Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>
2026-07-14 08:31:13 +00:00
jifeng
079ba35207
feat(web-shell): add selection statistics to markdown tables (#6838)
* feat(web-shell): add selection statistics

* fix(web-shell): clear stale table selections

* fix(web-shell): address selection statistics review
2026-07-14 03:54:33 +00:00
jinye
1f0078c7a2
feat(serve): Add workspace-qualified session export (#6844)
* feat(serve): add workspace-qualified session export

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* codex: address PR review feedback (#6844)

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

---------

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>
Co-authored-by: qwen-code-dev-bot <qwen-code-dev-bot@users.noreply.github.com>
2026-07-14 03:50:06 +00:00
jinye
c7250df8ea
feat(serve): Add workspace-qualified Voice (#6839)
* feat(serve): add workspace-qualified voice

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* fix(serve): harden workspace voice lifecycle

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* codex: address PR review feedback (#6839)

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* codex: address PR review feedback (#6839)

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* test(cli): address workspace Voice review feedback

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* codex: address PR review feedback (#6839)

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* fix(cli): clean up Voice lifecycle resources

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* fix(cli): address Voice review feedback

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

---------

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>
Co-authored-by: qwen-code-dev-bot <qwen-code-dev-bot@users.noreply.github.com>
2026-07-14 03:42:58 +00:00
jinye
fea3ab3854
feat(serve): add extension management v2 (#6825)
* feat(cli): workspace-qualified extensions REST (daemon multi-workspace)

Mirror the daemon extension-management REST surface to per-workspace routes, reusing the Phase 3 runtime resolver and trust gate. Extract a per-workspace extensions controller so the primary workspace shares one install queue, operation history, and status cache across the legacy and workspace-qualified routes. Reads resolve the target runtime only; mutations require a trusted workspace. Advertise a new baseline capability so clients can discover the surface, and add matching SDK client methods.

Refs #6378.

* qwen: address PR review feedback (#6638)

Align the new extensions controller file's copyright year with the other new files added in this change.

* qwen: address PR review feedback (#6638)

Redact credentials from the extension source on the two success-path fan-outs (session refresh and refresh-failure broadcast), matching the operation record and failure broadcast. Document the non-cancellation semantics of the extension timeout wrapper.

* qwen: address PR review feedback (#6638)

Share the queue-full sentinel message via an exported constant so the throw site (controller) and the 429 match site (routes) cannot drift after the module split. Include the bound workspace in the extension operation log prefixes so concurrent per-workspace controllers are distinguishable in stderr.

* feat(cli): add concurrent extension preparation

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* fix(cli): remove redundant extension context build

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* codex: address PR review feedback (#6638)

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* codex: address PR review feedback (#6638)

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* fix(cli): address extension review feedback

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* fix(extensions): address final review feedback

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* fix(extensions): address latest review feedback

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* codex: address PR review feedback (#6638)

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* fix(core): reject links in npm extension archives

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* fix(core): limit npm extension archive downloads

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* fix(extensions): address review follow-ups

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* fix(extensions): address latest review feedback

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* fix(extensions): release rejected operation slots

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* fix(extensions): address operation review feedback

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* fix(extensions): align archive handling contracts

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* fix(extensions): preserve watcher generation state

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* docs(extensions): align management contracts

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* fix(sdk): bound extension operation polls

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* test(core): cover forged prepared commits

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* test(core): assert activation generation increment

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* fix(extensions): close archive and polling gaps

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* fix(cli): retry suppressed extension generations

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* test(cli): cover archive URL extension updates

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* fix(sdk): preserve unbounded operation waits

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* fix(core): share npm redirect download deadline

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* fix(core): preserve extension reload diagnostics

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* fix(core): preserve installed Claude plugin paths

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* fix(cli): return committed activation state

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* fix(cli): preserve extension preparation errors

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* fix(core): validate extension setting env vars

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* fix(cli): target extension reconciliation

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* codex: address PR review feedback (#6638)

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* codex: address PR review feedback (#6638)

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* codex: address PR review feedback (#6638)

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* codex: address PR review feedback (#6638)

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* codex: address PR review feedback (#6638)

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* codex: address PR review feedback (#6638)

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* codex: address PR review feedback (#6638)

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* test(cli): cover resultless legacy commit warnings

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* fix(cli): retain suppressed extension generations

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* fix(cli): record legacy runtime reconciliation

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* fix(cli): validate extension clients by runtime

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* fix(cli): record workspace activation refresh

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* test(cli): stop extension reconcilers after cases

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* fix(cli): resolve global runtimes at reconciliation

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* fix(cli): reconcile newly registered runtimes

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* fix(cli): prevent overlapping runtime reconciliation

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* fix(cli): dispose late runtime apps during shutdown

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* fix(core): keep projection repair best effort

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* fix(core): preserve committed store results

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* fix(core): quarantine corrupt store journals

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* fix(core): harden npm download redirects

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* fix(extensions): address review edge cases

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* fix(extensions): honor cancellation between preparation stages

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* fix(extensions): retry prepared cleanup failures

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* test(extensions): cover committed artifact recovery boundary

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* fix(serve): release extension refresh queue on timeout

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* codex: address PR review feedback (#6638)

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* codex: address PR review feedback (#6638)

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* codex: address PR review feedback (#6638)

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* codex: address PR review feedback (#6638)

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* fix: address extension review feedback

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* codex: address PR review feedback (#6638)

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* codex: address PR review feedback (#6638)

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* codex: address PR review feedback (#6638)

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* codex: address PR review feedback (#6638)

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* fix(core): reconcile extension store compatibility state

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* fix(core): bound npm redirects and isolate extension tests

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* fix(core): make extension uninstall store-authoritative

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* fix(core): defer prepared extension secret mutations

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* fix(core): validate staged extensions before commit

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* fix(core): enforce public extension network policy

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* fix(core): handle extension response failures

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* fix(extensions): surface committed refresh warnings

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* fix(cli): guard timer unref calls

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* fix(extensions): release commit lane after durable writes

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* test(serve): update mutation callback assertions

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* fix(serve): refresh live extension instructions

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* fix(extensions): address latest review feedback

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* fix(extensions): address follow-up review findings

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* fix(extensions): address remaining activation feedback

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* fix(serve): preserve preparation queue status

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* fix(extensions): enforce network request deadlines

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* docs(serve): clarify single-workspace capabilities

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* fix(extensions): guard deferred settings commit

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* fix(extensions): cancel archive extraction

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* fix(extensions): harden refresh recovery

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* fix(serve): serialize extension reconciliation

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* fix(extensions): address post-commit review feedback

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* codex: address PR review feedback (#6825)

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* codex: address PR review feedback (#6825)

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* codex: address PR review feedback (#6825)

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* codex: fix CI failure on PR #6825

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* codex: fix CI failure on PR #6825

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* codex: address PR review feedback (#6825)

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* fix: address critical PR review feedback

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* fix(cli): bound legacy extension update checks

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* fix(acp): deduplicate extension refresh requests

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* codex: address PR review feedback (#6825)

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* codex: address PR review feedback (#6825)

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* codex: address PR review feedback (#6825)

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* codex: address PR review feedback (#6825)

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* fix(sdk): update browser bundle budget after main merge

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

---------

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>
Co-authored-by: qwen-code-dev-bot <qwen-code-dev-bot@users.noreply.github.com>
2026-07-14 03:30:47 +00:00
callmeYe
53468cd8af
feat(daemon): add workspace skill toggle API (#6816)
* feat(daemon): add workspace skill toggle API

* test(daemon): cover skill toggle capability integration

* fix(daemon): harden skill refresh handling

* fix(daemon): improve skill refresh diagnostics

* test(daemon): expand skill toggle coverage

---------

Co-authored-by: qwen-code-dev-bot <qwen-code-dev-bot@users.noreply.github.com>
2026-07-14 02:59:13 +00:00
tanzhenxin
220fba7917
feat(subagents): make Explore inherit the main model by default (#6807) 2026-07-14 01:23:35 +00:00
jinye
9dd8389ebe
fix(serve): Route session continue, language, and artifacts by owner (#6833)
* fix(serve): route session mutations by owner

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* codex: address PR review feedback (#6833)

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

---------

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>
2026-07-13 23:11:50 +00:00
jinye
13c224f5e9
feat(serve): support runtime workspace removal (#6745)
* feat(serve): support runtime workspace removal

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* fix(cli): address workspace removal review feedback

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* test(cli): strengthen workspace removal regressions

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* test(webui): fix timeout assertion lint

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* fix(cli): address workspace removal review feedback

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* fix(cli): update workspace Git test registry

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* fix(daemon): address workspace removal review feedback

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* codex: address PR review feedback (#6745)

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* codex: address PR review feedback (#6745)

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* codex: address PR review feedback (#6745)

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* codex: address PR review feedback (#6745)

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* codex: address PR review feedback (#6745)

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* codex: address PR review feedback (#6745)

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* codex: address PR review feedback (#6745)

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* codex: address PR review feedback (#6745)

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* codex: address PR review feedback (#6745)

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* codex: fix CI failure on PR #6745

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* codex: address PR review feedback (#6745)

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* codex: address PR review feedback (#6745)

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* test(web-shell): cover workspace removal after sidebar rebase

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* codex: address PR review feedback (#6745)

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

---------

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>
2026-07-13 15:43:38 +00:00
jinye
b933b90172
feat(serve): support multi-workspace rewind and shell (#6826)
* feat(serve): support multi-workspace rewind and shell

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* codex: address PR review feedback (#6826)

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* codex: address PR review feedback (#6826)

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

---------

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>
2026-07-13 15:32:44 +00:00
callmeYe
be4c46d085
feat(serve): expose skill installation paths (#6811) 2026-07-13 09:29:09 +00:00
jinye
3d5dd41bc7
fix(serve): route session actions to the owning workspace (#6798)
* fix(serve): route session actions to the owning workspace

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* codex: address PR review feedback (#6798)

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

---------

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>
2026-07-13 08:56:33 +00:00
ytahdn
e6bd1b1c12
feat(web-shell): add session created callback (#6703)
* feat(web-shell): add session created callback

* fix(web-shell): bound session callback setup

* fix(web-shell): serialize session preparation

* fix(web-shell): clarify session preparation diagnostics

* fix(web-shell): handle session preparation races

---------

Co-authored-by: ytahdn <ytahdn@gmail.com>
Co-authored-by: qwen-code-dev-bot <qwen-code-dev-bot@users.noreply.github.com>
Co-authored-by: Shaojin Wen <shaojin.wensj@alibaba-inc.com>
2026-07-13 07:04:42 +00:00
jinye
6d286149d6
feat(serve): Bound persisted transcript pages (#6769)
* feat(serve): bound persisted transcript pages

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* fix(serve): address transcript page review feedback

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

---------

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>
2026-07-13 03:18:11 +00:00
jinye
98f2bb37ec
feat(cli): Add runtime daemon channel control (#6741)
Some checks are pending
E2E Tests / web-shell Browser Regression (push) Waiting to run
E2E Tests / E2E Test (Linux) - sandbox:docker (push) Waiting to run
E2E Tests / E2E Test (Linux) - sandbox:none (push) Waiting to run
E2E Tests / E2E Test - macOS (push) Waiting to run
* feat(cli): add runtime daemon channel control

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* fix(cli): address daemon channel review feedback

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* fix(daemon): align channel control timeout budget

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* fix(cli): preserve serve fast-path import boundary

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* fix(cli): distinguish pending channel generations

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* fix(cli): use workspace env for deferred webhook auth

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* codex: address PR review feedback (#6741)

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* codex: address PR review feedback (#6741)

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* codex: address PR review feedback (#6741)

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* codex: address PR review feedback (#6741)

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* codex: address PR review feedback (#6741)

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* codex: address PR review feedback (#6741)

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* codex: address PR review feedback (#6741)

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

---------

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>
Co-authored-by: Shaojin Wen <shaojin.wensj@alibaba-inc.com>
2026-07-13 02:53:27 +00:00
BaboBen
1dd984509b
fix(feishu): validate credentials before WebSocket startup (#6780)
* docs(channels): design Feishu worker credential validation

* docs(channels): plan Feishu worker credential fix

* fix(feishu): wait for authenticated WebSocket startup

* fix(feishu): close late WebSocket readiness

* fix(feishu): validate credentials before WebSocket startup
2026-07-13 02:10:16 +00:00
dreamWB
460c92102e
feat(web-shell): support custom composer placeholders (#6765) 2026-07-12 12:44:59 +00:00
Shaojin Wen
d1b2a7fb72
feat(review): procedural correctness finders, effort levels, and posting/verify guardrails (#6711)
* feat(review): procedural correctness finders, effort levels, and posting/verify guardrails

Rework the /review skill's finder layer and add precision and cost controls,
informed by dogfooding the skill against real PRs.

Recall:
- Split Agent 1 (Correctness) into three procedural finders defined by how they
  walk the diff — 1a line-by-line (incl. language-pitfall and wrapper-routing
  checks), 1b removed-behavior audit, 1c cross-file tracer — so coverage is
  complementary instead of overlapping. Bump the 3A dimension fan-out to 12
  agents and shift the 3A/3B gate to 3200 diff lines.
- Add Agent 8: up to two diff-specialized finders written per-review when the
  diff concentrates in a domain with a known failure grammar.
- Fold altitude into Code Quality and a quote-the-rule discipline into the
  conventions pass.

Precision:
- Every finding must state a concrete failure scenario (trigger to wrong
  outcome, or concrete cost); findings that can't are dropped at the source,
  and verification re-traces the scenario rather than judging prose.
- Verification checks a finding against the diff's own documented intent: a
  "regression" the diff deliberately changes and documents is a design
  decision, not a defect.

Cost and safety:
- Add --effort low|medium|high: cheap inline passes with no subagents (default
  high for PRs, medium for local changes).
- Hard-gate PR posting: never submit a review unless --comment was passed or the
  user explicitly asked, regardless of verdict.
- Add a substantive-return check for whole-diff agents (invariant, cross-file,
  test-coverage matrix) so a silently whiffing agent is caught like a missing
  chunk receipt.

DESIGN.md records the rationale and dogfooding cases behind each change; user
docs updated with the effort levels and the new agent roster.

* docs(review): fix stale topology numbers flagged in review

- Define H in the 3B pipeline diagram cost annotation (3 invariant agents
  per heavy file).
- Annotate the 40-PR re-gating cost figures with the roster they were
  measured under (22 agents / ~5% at 10 agents; ~34 / ~7% at 12).
- Correct the fork-subagent savings estimate to ~88-92% (~750-950K →
  ~80-88K); the previous range predated the updated totals.
- "None or nine" receipts under 3A is eleven under the 12-agent roster
  (every agent except Build & Test walks the diff).

* fix(review): address review feedback on effort/verify/lightweight edge cases

Criticals from review:
- Apply the --comment→high-effort override only after target disambiguation;
  an ignored --comment (non-PR target) no longer silently forces high.
- The documented-intent gate caps confidence only when the rationale makes the
  harm uncertain; a traced harm that survives the rationale keeps high
  confidence, and rejection is reserved for pure re-descriptions.
- Lightweight cross-repo mode degrades Agents 1a/1b to diff-only and routes
  unverifiable re-establishment claims to low confidence instead of asserting
  them, matching the verifier's limits.

Suggestions:
- Agent 0's empty-scope exit now carries its evidence and the whiff check
  recognizes it, so a legitimate no-linked-issue return is not relaunched.
- Reframe the 3200-line clause as an attention bound (3B is not guaranteed
  cheaper with heavy files or specialized finders).
- Fix call-budget notation: F for findings vs N for chunks; correct the 3B
  budget to rounds × chunks for the reverse audit (~70 calls on the 19-chunk
  example, not ~28-30); state the runtime concurrency cap (10) instead of
  claiming ~1x wall time.
- Preserve the failure scenario through pattern aggregation and posted inline
  comments; extend quality-finding verification to check the named helper
  does what the finding claims.
- Document medium effort's roster (no dedicated security/test-coverage pass);
  qualify the cross-effort scope note (incremental cache is high-only); define
  "lenses" on first use; fix tense and the stale 9-agent line in commands.md;
  clarify the 1b skip condition (no removed/replaced lines).

* fix(review): close 422-relocation verdict hole and lightweight-mode context gap

From review feedback (one human, two model reviews):

- 422 recovery: Criticals relocated into the review body now keep the event at
  REQUEST_CHANGES — the event/body table counted comments only, so a review
  whose blockers were all relocated could submit as APPROVE/no-blockers
  COMMENT with blocker text in the body.
- Lightweight cross-repo mode now runs pr-context (pure GitHub API): Agent 0
  and the open-Critical re-check need the PR body and open threads, which the
  bare gh-pr-diff setup never captured.
- Define --effort value parsing so a non-enum next token (e.g. a PR number) is
  never consumed as the value.
- Add the missing test-coverage-matrix definition section; mark agent counts
  as maxima (1b skipped on no-deletion diffs).
- Sync DESIGN's documented-intent paragraph with the corrected confidence
  policy; align 3A/3B budget headings with the dual-trigger gate and state the
  38-95 reverse-audit range explicitly.
- Docs: dual-trigger diagram labels, attention-bound wording, diff-reading
  lenses phrasing, per-stage-bounded (not fixed-total) cost claim, effort
  table qualifications, failure-scenario in the Step 7 JSON samples.

* docs(review): reconcile verifier rejection rule and close remaining edge notes

- State the Critical-rejection bar once, without the self-contradicting
  "never reject / to reject" phrasing: rejection requires quoting the
  contradicting code, and the floor verdict is confirmed (low confidence)
  when it cannot be quoted.
- Note the one sanctioned exception to the empty REQUEST_CHANGES body:
  unmappable or 422-relocated Criticals.
- Give the callee-direction check a concrete procedure (walk the other
  changed symbols this territory calls, re-read their post-change
  contracts).
- Define lightweight-mode pr-context failure handling: continue diff-only,
  skip Agent 0, open-Critical re-checks become "cannot tell" (no Approve).
- Clarify Agent 8 applies in every mode (it needs only the diff) and that
  Step 6 follow-up tips are high-effort only.

* fix(review): close flag-parse, context-unavailable, and downgrade edge cases

Address the latest review round on the skill text:

- An invalid spaced --effort value is discarded (with the warning) whenever
  another token is the target, so `/review 6711 --effort typo` reviews PR
  6711 instead of leaking `typo` into target disambiguation; the token is
  kept only when it is itself the sole target candidate.
- The lightweight-mode pr-context failure now names a context-unavailable
  state with a defined Step 7 serialization: never APPROVE, submit COMMENT
  with a diff-only body, findings or not.
- Step 6's open-Critical re-check draws from both context sections; a reply
  alone ("I disagree") no longer retires a blocker — only a code-verified
  "fixed by this diff" does.
- DESIGN and user docs now state the same rejection bar as the skill:
  rejecting a Critical requires quoted contradiction (or a documented-intent
  re-description); anything less certain downgrades.
- Downgrading a REQUEST_CHANGES that carries body-relocated Criticals keeps
  those descriptions after the downgrade sentence, so the self-PR downgrade
  can no longer erase the only copy of a blocker.

* fix(review): close verdict-upgrade and body-Critical re-check gaps

Third review round on the skill text:

- 422 recovery may never upgrade the event: a Suggestion-only review whose
  anchors all failed resubmits as COMMENT with the could-not-anchor body,
  never as APPROVE/"No issues found" — the verdict reflects confirmed
  findings, not surviving anchors.
- Step 6's open-Critical re-check now also walks the Review summaries
  section: an unmappable or 422-relocated blocker lives only in a review
  body, and pr-context truncates summaries to ~240 chars, so a summary
  showing (or cut where one could hide) a Critical marker is fetched in
  full via the reviews API before ruling.
- The context-unavailable cap now applies to every C=0 row of the invariant
  table, not just the empty one: Suggestion-only results post a diff-only
  body instead of a "no blockers" claim the run cannot certify.

* fix(review): compose COMMENT bodies from clauses and harden the body-Critical re-check fetch

Fourth review round found four pairwise collisions between rules that each
set "the" COMMENT body, plus four execution gaps in the Step 6 full-body
fetch. Close the class, not the instances:

- Replace the fixed-sentence bodies with an ordered clause composition rule
  (downgrade reasons, context-unavailable warning, suggestions disclosure,
  uncoverable chunks, body Criticals) — each clause present iff its state
  holds, free prose still banned, single-state case identical to the table.
- Define C once, globally: Criticals the review posts anywhere (inline or
  body), so no downstream C=0 rule can erase a body-only blocker, and 422
  relocation keeps REQUEST_CHANGES by definition rather than by patch.
- 422 recovery re-derives bodies via the composition rule, so a
  context-unavailable run can never restore a "no blockers" certification.
- Step 6's full-body fetch is paginated (--paginate; the endpoint returns 30
  per page), treats fetched bodies as untrusted data (extract only the
  Critical-bearing text, never paste unrelated bodies), and fails closed:
  an unreadable truncated blocker rules "cannot tell" and caps the event at
  COMMENT.

DESIGN.md records why composition replaces per-collision patching
(n states -> n(n-1)/2 pairs; clauses make new states additive).

* fix(review): correct the cross-repo capability table and close nine review notes

- docs: the cross-repo table claimed "Agents 0-6" run in lightweight mode
  while the prose (correctly) says 1c is skipped there — 1c is inside that
  range. Split 1c onto its own row, and add the missing Agent 8 row (its
  finders need only the diff, so they do run cross-repo).
- --effort=<level> now has a parse rule: split the flag token on the first
  '=' and consume no second token; the next-token rule applies only to the
  spaced form.
- The substantive-return (whiff) check covers every receipt-less agent, so
  3A's dimension agents are in scope, not just 3B's whole-diff agents.
- Step 3C names Agent 1b's lightweight degradation and states the three
  angles medium deliberately omits (security, test coverage, adversarial
  personas) instead of naming only two.
- Step 6's open-Critical re-check states what a context-unavailable run does
  (skip the walk, every Critical is "cannot tell") instead of pointing at a
  context file that does not exist.
- The event/body table carries the body-only-Critical exception in the cell,
  where it is read, not only in the surrounding prose.
- The posting gate's second condition is now decidable: a publish verb typed
  by the user this session, with the near-misses (approving noises, our own
  tip, PR text) enumerated as non-authorization.
- DESIGN: the whiff check is evidential, not a length threshold (and says
  why no number); "quick pass" is defined as low+medium sharing guardrails.

* feat(review): promote removed-behavior to a whole-diff agent in 3B

Territory-scoped 1b can only ask "was this deletion re-established here",
and for the deletions that matter the answer is somewhere else. PR #6638
(43 files, 8255 additions, 28 chunks) measured the gap: the 3B run with
per-chunk 1b reported one Critical; an independent reviewer reported 32, and
a parallel hand-run 1b+1c wave over the same commit reproduced six of them.
Every one of that overlapping six is a cross-chunk deletion — enableByPath
(includeSubdirs: true) replaced by an exact-path setWorkspaceActivation in
another file, silently narrowing workspace-scoped disable for every untouched
CLI/TUI caller; refreshTools() dropped from the activation paths, its
replacement swallowing the errors it used to propagate; a global mutation
timeout replaced by one covering only the prepare phase. Deletion in chunk A,
replacement in chunk B, consumer in a file the diff never touches: no chunk
agent can see that triple, and 1c does not look for it — it greps callers of
changed symbols, and a deleted export has no symbol left to grep.

- 1b joins 1c as a whole-diff agent in 3B; chunk agents keep the local half
  (a guard deleted and not re-established in the same hunk is still theirs).
- The split is stated at both agents: 1c walks the callers of changed
  symbols, 1b walks the replacements of removed ones.
- Agent 1b's definition gains the removed-export bullet: compare replacements
  as behaviour, not names, then check the call sites the diff never touches —
  a replacement that type-checks is not a replacement that behaves.
- 3B whole-diff agent count 4-6 -> 5-7 in the budget and the docs diagram.

* fix(review): serialize cannot-tell blockers, gate the no-blockers opener, and read reviews from a file

Fifth review round, all four notes real:

- The clause inventory had no way to serialize Step 6's `cannot tell`
  verdict, so a Critical the review could neither confirm nor clear had
  nowhere to go and dropped out of the public review. Added clause 5
  (unresolved existing-Critical), which survives downgrades and 422 recovery
  like the body-Critical carve-out, and Step 6 now points at it.
- `Reviewed — no blockers.` was injected as the opener whenever context was
  available, regardless of C or scope — so a self-PR downgraded to COMMENT
  with an inline Critical, or a review with an uncoverable chunk, opened by
  certifying the absence of the blockers it was carrying. The opener is now
  gated on C === 0 AND no unresolved existing Critical AND no uncoverable
  chunk AND context available; otherwise it is a plain `Reviewed.`
- The paginated `/reviews` fetch ran through the shell, whose successful
  output is capped at 30 000 chars and split head/tail: a body-only blocker
  in the elided middle passes with exit 0 and the fail-closed branch never
  fires. It is now redirected to a file and paged with read_file, and a body
  read only in part is `cannot tell`, not "no Critical in it" — the same
  lesson as "the diff is a file, not a command".
- The substantive-return gate rejected a bare "No issues found" while the
  agent contract demanded exactly that string. The contract now asks for
  `No issues found — <one line naming what you examined>`, and the relaunch
  is capped at one attempt per agent, with the dimension reported under "Not
  reviewed" if the second return is still bare.

* fix(review): wire whole-diff 1b into the gates, cap the event on unread scope, page reviews as NDJSON

Sixth round. Four Criticals all trace to the two previous commits:

- Whole-diff Agent 1b was declared but never wired in: it was missing from
  the receipt-less roster (so a whiffing 1b passed undetected) and the launch
  contract handed every 3B agent "its own chunk range", which is exactly what
  a cross-chunk pairing agent cannot work from. The payload contract now
  splits by role — chunk agents get one range, every whole-diff agent gets
  the entire chunks[] plan.
- The whiff check ended at "note it as Not reviewed", which left an
  unreviewed Security or removed-behavior lens able to ship an LGTM. It now
  carries an unreviewedDimensions state that forbids Approve, caps the event
  at COMMENT, and is serialized in the body next to uncoverable chunks.
- Clause 5 put an undecidable existing Critical in the body while event
  selection still chose APPROVE from the C/S table — a review approving the
  very blocker it asks the author to confirm. The table now has explicit
  overrides: cannot-tell existing Critical, uncoverable chunk, and unreviewed
  dimension each cap the event at COMMENT (a confirmed Critical still earns
  REQUEST_CHANGES).
- Redirecting `gh api --paginate` to a file does not make it pageable: it
  emits compact JSON, so the file is one 150 KB+ line that read_file
  truncates and offset skips past to EOF. The fetch now filters with --jq to
  marker-bearing bodies and emits line-delimited records that page normally.

Also: the 1b/1c split is by task, not by symbol (1c greps the removed
export's old name and owns caller compatibility; 1b owns the pairing and the
semantic comparison) — the earlier "no symbol left to grep" claim understated
1c and risked dropping its removed-symbol pass. Plus stale arithmetic from the
larger roster (+4 -> +5, worked example 26-28 -> 27-29), the receipt count's
missing Agent 8, "0 LLM calls" -> "0 subagent calls", the fixed "12 parallel
tasks" -> its real range, and 1c's callee procedure no longer speaking of a
"territory" it does not have.

* fix(review): select body Criticals offline, propagate unreviewed dimensions, close the no-findings bypass

Seventh (final self-review) round. The three Criticals all attack the newest
machinery:

- The NDJSON fetch filtered on a literal [Critical] marker, but a body-only
  blocker is not guaranteed to carry it (a real emitted review on this repo
  does not) — the filter discarded exactly what the re-check exists to
  recover. The --jq now keeps every nonempty body and selection happens
  offline after reading records whole; clauses 5 and 7 additionally mandate
  the marker on everything we serialize, so our own output stays
  self-identifying.
- unreviewedDimensions stopped at the event cap: Step 6's Not-reviewed
  section only listed uncoverable chunks (a non-posting run hid the missing
  lens entirely), and the body invariant made the required disclosure
  illegal on a REQUEST_CHANGES. The section now lists both, and the
  not-reviewed clause is the second sanctioned REQUEST_CHANGES body
  exception — a confirmed Critical must not squeeze out the disclosure of
  what was never read.
- The no-confirmed-findings branch still said "APPROVE by default",
  special-casing only presubmit and context-unavailable — bypassing the
  cannot-tell/uncoverable/unreviewed caps added one commit earlier. The
  branch now runs the same machinery as every submission: table with
  overrides, then downgrades, then composition; the hard-coded LGTM example
  applies only with no cap state present.

Plus the round's consistency notes: cross-file trace marked same-repo-only
in the docs' medium row; +5 -> +4 in the crossover arithmetic (Build & Test
reads no diff) so "crosses twelve about there" is true at 3200; DESIGN's
whole-diff enumeration gains 1b; budget total widened to the honest 15-21
row-sum; fork-subagent math redone at 52K/agent; the payload paragraph
names the invariant agents' third payload class; consumer-direction grep
patterns get Python/Go forms; 3C medium states 1a's lightweight degradation
and scopes the grep permission; the aggregated-format shorthand carries
Failure scenario and Severity; the exactly-one-sentence rule forward-
references the composition rule; and the Step 7 comment template embeds the
failure-scenario shape it was already demanding in prose.

* feat(review): sink argument parsing into a tested parse-args subcommand

The --comment/--effort grammar and target disambiguation were ~400 words of
prose in SKILL.md that the model re-simulated on every run; three separate
parsing bugs shipped that way (the spaced form consuming a flag as its
value, the --effort=<level> form left undefined, and an invalid value token
surviving into target disambiguation). Each is now a table-driven test case.

qwen review parse-args '<raw args>' emits a JSON verdict: classified target
(pr-number / pr-url with owner+repo+number extracted / file / local),
resolved effort with its source (explicit / default / forced-by-comment),
comment.requested vs comment.effective, verbatim warnings, and leftover
tokens the parser refuses to guess about. The skill's Step 1 shrinks to
"run the parser, use the verdict verbatim", and the target branches key off
target.type instead of hand-classifying tokens.

* feat(review): sink event selection and body composition into compose-review

The Step 7 machine — the C/S table, three event-capping overrides, the
seven-clause body composition, and the presubmit downgrade carve-outs — was
restated across four places in SKILL.md, and keeping the restatements in
sync by hand produced five shipped bugs (four Critical), all one shape: a
downstream branch not updated when an upstream rule gained a new state.

qwen review compose-review reads a state JSON (inline/body Critical and
Suggestion counts, discarded anchors, cannot-tell existing Criticals,
uncoverable chunks, unreviewed dimensions, context-unavailable, presubmit
flags, model id) and returns {event, body, baseEvent, cappedBy, downgraded}
for verbatim submission. The truth-table tests pin every previously shipped
bug as a named case: caps forbid APPROVE but never soften a REQUEST_CHANGES;
discarded Suggestions still count toward S so a 422 resubmit can never
upgrade to LGTM; a self-PR downgrade keeps body Criticals after the
downgrade sentence; the no-blockers opener appears only when certifiable;
every disclosure survives every stacking. Writing the tests immediately
caught one more instance of the class (all-discarded -> S=0 -> APPROVE).

SKILL.md's Step 7 shrinks to gathering the state and using the output
verbatim; the 422 recovery becomes "re-run compose-review with updated
counts"; the no-findings branch is the same call with zero counts; the
posting gate (judgment, not bookkeeping) stays prose.

* feat(review): render review bodies in full, quarantine replied Criticals, raise the gh buffer

The Step 6 body-fetch instruction was rewritten five times in four review
rounds (missing pagination -> shell truncation -> unpageable single-line
JSON -> a marker filter that discarded markerless blockers -> offline
selection) — the signature of a download program written in English. This
ends the chain at its root, in pr-context itself:

- Review bodies render in full under "Review summaries" instead of
  240-char snippets: an unmappable or 422-relocated blocker lives only
  there, and a snippet once hid one from the re-check. A body past the 8000
  cap ends by naming its review id, so the tail stays fetchable as a single
  object; a body read in part is `cannot tell`, not "no Critical in it".
- Replied Critical threads are quarantined into their own "Replied
  Criticals" section, rendered before the settled threads, instead of
  sinking into "Already discussed" — a reply alone ("I disagree") never
  retires a blocker, and marker-matching in this direction is fail-safe: a
  forged marker can only add a thread to the re-check list, never hide one.
- The gh wrapper's maxBuffer rises from Node's 1 MiB default to 64 MiB,
  closing the ENOBUFS that killed pr-context and presubmit mid-review on a
  comment-heavy 43-file PR.

SKILL.md's NDJSON fetch block is deleted: the re-check reads the context
file's three finding-bearing sections under its untrusted-data preamble,
with one residual single-object fetch for capped bodies. Verified against
this PR's own 100+-comment history: the markerless body-Critical review
that motivated the last rewrite now renders whole, and the fetch survives
without ENOBUFS.

DESIGN.md records the sinking rationale for all three subcommand changes;
the user docs note that parsing and the event/body decision are now pinned
by unit tests rather than prompt text.

* test(review): register parse-args and compose-review in the exact-list assertion

The parent-command test pins the exact subcommand roster; the two new
subcommands landed without updating it, which is precisely the drift the
assertion exists to catch — it caught it in CI, one directory above where
the new tests were run locally.

* fix(review): carry every disclosure on REQUEST_CHANGES and select blockers semantically

Review round on the new subcommands, plus the prompt notes it surfaced:

- compose-review's REQUEST_CHANGES branch dropped the context-unavailable
  clause entirely and gated the not-reviewed disclosure on other parts being
  present — an RC with only an uncoverable chunk disclosed nothing. Every
  clause whose state holds now appears on every event (a confirmed blocker
  must not squeeze out the trust warning or the unread-scope disclosure);
  four new tests pin it.
- Step 6 selects blockers semantically, not by the literal [Critical]
  marker: legacy body-only blockers were emitted markerless, and a marker
  filter once discarded exactly such a review.
- The same-repo pr-context failure now sets context-unavailable like the
  lightweight path (the guard's "lightweight" narrowing is removed) — a
  same-repo run that lost the context file must not behave as if it had
  read it.
- Step 5's dry-round return aligns with the agent contract (receipt-bearing
  "No issues found — <what it re-examined>"), ending the contradiction where
  a compliant reverse auditor would be flagged as whiffing.
- Consumer-direction grep forms for Python/Go are call sites now, with the
  declaration forms explicitly labeled as callee lookup.
- The 15-19 totals left downstream (docs table, DESIGN heading and cost
  row) move to the honest 15-21 / 13-20.

* fix(review): stdin transport for parse-args, validated compose input, full-body re-check context

Round 9 of review-the-review on this PR: 19 unique findings across three
reviews, each verified against source before fixing.

parse-args:
- The documented positional invocation broke on any flag-first raw string
  (`qwen review parse-args '--effort low'` -> "Unknown argument") and the
  `--` form silently returned a wrong local/default verdict. The raw
  string now travels on stdin (`--stdin`; SKILL.md pipes a quoted
  heredoc, immune to leading dashes, quotes, and $(...)); positional +
  --stdin and post-`--` smuggling are refused loudly. Wiring-level tests
  drive the real yargs command, pinning the strict-mode rejection that
  pure-function tests could not see.
- PR URL identity hardened: the number must end its path segment
  (/pull/42oops is refused, never PR 42), owner/repo restricted to
  GitHub's name charset (keeps shell metacharacters out of derived
  values), scheme matched case-insensitively, url canonicalized
  (lowercase scheme/host, query/fragment dropped) with a new host field;
  near-miss URLs are warned about and reported in extraTokens, never
  guessed into a file path or PR number. Step 1 remote matching now
  requires host AND owner/repo.
- Repeated --effort warnings state what is actually in effect (last valid
  occurrence / --comment forcing / the default), composed after
  resolution; previously a later typo claimed the default while an
  earlier valid effort stayed active.

compose-review:
- Input validated at the boundary: absent counts default to 0; malformed
  values throw typed errors naming the field. Previously
  {bodyCriticals:["x"], modelId} made undefined+1=NaN, failed both event
  comparisons, and returned APPROVE over the only blocker.
- "Suggestions are inline." keys off suggestionsInline, not s: an
  all-discarded 422 recovery no longer claims inline suggestions while
  the discarded sentence says the opposite (s still decides the event).
- canCertify requires !downgraded: a downgraded Approve opens with the
  neutral "Reviewed." instead of certifying "no blockers" two clauses
  after naming failing CI.
- unreviewedDimensions entries may carry their own reason after an
  em-dash and render verbatim (used by Agent 0's fetch failure below).

pr-context:
- Replied-Critical root bodies render in full (shared capBody; a cut
  names the comment id and the exact fetch); reply snippets name their
  comment id when cut. The Step 6 re-check no longer rules on
  silently-truncated claims, and the fail-closed "read in part = cannot
  tell" rule can actually fire for this section.
- The LGTM filter matches the exact canonical template, anchored to the
  whole body: a legacy body opening with the LGTM line but carrying a
  relocated blocker below it is shown instead of dropped.
- classifyInlineThreads() extracted: buildMarkdown and the stdout count
  use the same walk, so the count cannot diverge from the file.

SKILL.md:
- Step 6 re-check scope: every comment-bearing section, including
  "Already discussed" (inline threads and issue-level comments) — the
  quarantine keys on the literal marker, a floor not a ceiling, so
  unmarked blockers settle there; the false "holds only non-Critical
  threads" parenthetical is gone. The residual long-body fetch redirects
  to a file (shell output truncates at 30k) and is read paged.
- Step 5 reverse audit: dry = zero new findings WITH the evidence-bearing
  receipt; the substantive-return check runs after every round (one
  relaunch); a twice-whiffed agent's round is never dry.
- Step 3: Agent 7 added to both whiff-check rosters (evidence = commands
  run + outcomes; build-and-test recorded in unreviewedDimensions on the
  second whiff). Agent 0's linked-issue fetch failure is fail-closed
  after one retry via a self-explained unreviewedDimensions entry.
- Step 8: a fail-closed run (unreviewed dimensions, uncoverable chunks,
  context-unavailable) must not advance the incremental cache — caching
  it would exempt the disclosed-unreviewed scope from every future run.
- Counting truthfulness: "Twelve agents all reading the same diff" is
  eleven (every 3A agent except Build & Test walks the chunk plan); fixed
  in the 3B rationale, the diff-capture section, and the user docs.

review.ts: demandCommand message names plan-diff, with a test that the
message stays in sync with the registered roster.

* fix(review): nested-safe stdin guard, validated presubmit, refetchable snippets everywhere

Round 10: 12 findings, all verified before fixing. The headline is
self-inflicted: the round-9 post-`--` guard read argv._ as
['parse-args', ...extras], but the real CLI nests the command, so argv._
is ['review', 'parse-args'] and the guard rejected every real
invocation — while the wiring tests, which register the command
top-level, stayed green. Reproduced against the built CLI before
fixing.

parse-args:
- The smuggle guard skips the command-path prefix in argv._; new wiring
  tests go through the real parent `review` command (nested stdin
  invocation + nested post-`--` refusal).
- --effort values match case-insensitively (`--effort High` is not a
  file target named High); the verdict keeps the lowercase form.
- Single-dash tokens are unknown flags, never target candidates
  (`/review -c 6711` reviewed a nonexistent file `-c` and demoted the
  PR number to extraTokens).

compose-review:
- presubmit and contextUnavailable get the same boundary validation as
  the counts: boolean flags reject stringified "false" (truthy — it
  flipped an inline-Critical RC to COMMENT and published the diff-only
  warning on runs that fetched context fine), downgradeReasons rejects
  scalars with the field name (was a raw .join TypeError), presubmit
  rejects non-objects.
- Certification is gated on what presubmit PERMITS, not on whether it
  changed the event: a Suggestion-only review is already COMMENT, so
  failing CI flipped nothing and the body still certified "no
  blockers". Either downgrade flag now suppresses the certifying
  opener.

pr-context:
- Every truncating render carries an exact refetch ref: open-root
  snippets, settled replied threads (roots and replies), and
  issue-level comments (their own issues/comments endpoint). The
  Step 6 semantic re-check reads these sections, and a markerless
  blocker past the 240-char cut was invisible with no way back.
- Refs are copy-runnable: buildMarkdown threads owner/repo and PR
  number into every ref, so emitted commands carry real values.
  `gh api` substitutes only {owner}/{repo} — from the CURRENT repo,
  wrong in cross-repo mode — and passes {n} through literally.

SKILL.md:
- Step 1: the raw argument string travels via write_file to
  .qwen/tmp/qwen-review-args-input.txt and stdin redirection. A quoted
  heredoc disables expansion but not delimiter recognition, so a raw
  string containing the delimiter line would end the heredoc early and
  execute the rest as shell. Step 9 removes the file.
- Step 1: remote matching is structural segment equality (host AND
  owner/repo, .git stripped, case-insensitive) — substring "contains"
  let shao/qwen-code match a wenshao/qwen-code remote. Non-github.com
  hosts must carry GH_HOST on every gh call for the PR.
- Step 5: a twice-whiffed reverse-audit scope is tracked, cleared only
  by a later substantive audit, and fed into unreviewedDimensions as a
  self-explained entry when the loop ends — terminal prose alone let a
  capped run approve with an audit that never ran.
- Step 6: snippet cuts carry their own filled-in fetch note; ruling on
  a cut prefix is the fail-closed violation.
- Step 7: the stale hand-derivation bullets (event table, empty-RC-body
  rule, one-line COMMENT inventory) are replaced with descriptions of
  what compose-review guarantees; the sanity check is byte equality
  with the subcommand's output; the last-resort 422 branch re-runs
  compose-review instead of hand-building "the one-line body".
- Step 8: the fail-closed cache rule includes cannotTellCriticals — a
  cached SHA plus the same-SHA shortcut would skip the very re-check
  that must re-rule on an undecided blocker.
MSG2
git log --oneline -1; git push origin feat/review-procedural-finders-effort 2>&1 | tail -2

* feat(review): deterministic overlap disposal, --host routing, machine-readable completion line

Three changes measured out of the first six-PR dogfood batch, not
predicted from review comments.

Overlap disposal (SKILL.md Step 7): presubmit's overlap report used to
end in "list the overlaps to the user, ask whether to proceed" — 2 of 6
batch runs stalled on an improvised interactive question (fatal for a
headless run) while the other 4 proceeded. An overlap is a duplicate by
the Exclusion Criteria; the rule is now drop the overlapping finding,
adjust the counts handed to compose-review (a dropped finding never
flips the verdict), note "already reported at <path>:<line>" in the
terminal, and continue without asking. Zero findings left after
dropping is still not a question — compose-review handles the shape.

--host routing (lib/gh.ts + fetch-pr/pr-context/presubmit): the
round-10 GH_HOST-by-prose rule required the model to remember a prefix
on every call; a forgotten one silently reads from and posts to
github.com's same-named owner/repo. The three gh-calling subcommands
now accept --host and thread it through setGhHost()/ghEnv(), so every
wrapped gh call carries GH_HOST in code; hostname input is
charset-validated. SKILL.md keeps the prose prefix only for the gh
commands the orchestrating model runs directly (Agent 0's fetches,
Step 6's residual body fetch, Step 7's submission).

Completion line (SKILL.md Step 9): three different ad-hoc completion
phrasings across one batch each needed their own driver regex. Every
run now ends with exactly one line, `Review complete: <target> —
<disposition>`, with a closed disposition grammar covering posted
events, unposted verdicts, and quick passes — detectable with a single
^Review complete: match.

Tests: gh host-state unit tests (inherit-by-default, GH_HOST extension,
host:port, charset rejection), presubmit handler --host threading (set
and reset), builder registration checks for fetch-pr and pr-context.
2026-07-12 11:33:52 +00:00
jinye
60fcc8cbce
fix: Make chat recording failures durable and visible (#6743)
* fix(core): Stop chat recording after write failure

Keep the canonical JSONL write chain rejected after the first asynchronous failure so queued descendants are skipped and flush reports the original error consistently.

Cover sticky failures across ordinary, strict, parent-session, ACP close, rename, branch, and rewind paths.

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* fix: Surface chat recording failures

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* fix(core): Await custom title persistence

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* codex: address PR review feedback (#6743)

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* docs(cli): clarify degraded branch behavior

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* fix(acp): align recording state entry type

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* fix: address PR review feedback (#6743)

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* fix: address durability review findings (#6743)

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* fix(core): allow artifact migration after recording failure

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* docs(daemon): correct UI event counts

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

---------

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>
Co-authored-by: Shaojin Wen <shaojin.wensj@alibaba-inc.com>
2026-07-12 10:52:26 +00:00
samuelhsin
7468e75e3d
feat(web-shell): support custom Hex session group colors (#6752)
* feat(web-shell): support custom Hex session group colors

* docs(web-shell): add custom group color screenshot

* fix(web-shell): address custom Hex color review feedback

* fix(web-shell): validate group presets against daemon catalog and auto-prefix Hex input

Review follow-ups for the custom Hex group color editor:

- Validate the preset branch against the daemon-provided color catalog
  instead of the hardcoded palette, so future preset additions stay
  selectable in the editor.
- Auto-prefix bare values with '#' in the Hex field so pasted bare Hex
  validates, and free text can no longer collide with a preset name and
  silently flip the select out of Custom mode.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(web-shell): cap custom Hex input length

---------

Co-authored-by: qwen-code-dev-bot <qwen-code-dev-bot@users.noreply.github.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-12 10:41:17 +00:00
jinye
d14aca72a6
feat(serve): add workspace persisted transcript reader (#6740)
* feat(serve): add workspace persisted transcript reader

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* refactor(cli): rename replay modules to kebab case

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* codex: address PR review feedback (#6740)

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* test(serve): cover multi-record transcript paging

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

---------

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>
Co-authored-by: Shaojin Wen <shaojin.wensj@alibaba-inc.com>
2026-07-12 10:39:05 +00:00
nas
0579be6ee8
feat(core): add configurable default timeout for foreground shell commands (#6628)
* feat(core): add configurable default timeout for foreground shell commands

Foreground shell commands started by the agent time out after a hardcoded
120s (DEFAULT_FOREGROUND_TIMEOUT_MS). A per-call `timeout` param can raise
that for a single command, but there is no way to change the default for a
project or session, so users repeatedly watch long-running commands fail at
the 2-minute mark.

Add a `tools.shell.defaultTimeoutMs` setting that feeds the existing timeout
resolution. Precedence is now: per-call `timeout` param > setting >
built-in default. When the setting is unset, behavior is unchanged; a value
of 0 disables the timeout, matching the existing per-call semantics.

Fixes #5838

* fix(core): add mock getShellDefaultTimeoutMs + bound defaultTimeoutMs

Address review on #6628:
- Add getShellDefaultTimeoutMs to mock configs in coreToolScheduler.test.ts
  and toAutoClassifierInput.test.ts (ShellTool construction now reads it).
- Add minimum: 0 / maximum: 600000 to the defaultTimeoutMs setting so a
  negative value can't reach AbortSignal.timeout(); regenerate schema.

* chore(core): polish shell defaultTimeoutMs per review

- shell.ts: debug-log the resolved foreground timeout (per-call vs
  configured default vs built-in) for observability
- settingsSchema.ts: use type 'integer' for tools.shell.defaultTimeoutMs
  to match sibling visionBridgeTimeoutMs; regenerate settings.schema.json
- config.test.ts: add loadCliConfig test asserting
  tools.shell.defaultTimeoutMs maps to Config.getShellDefaultTimeoutMs()

* fix(core): validate shell defaultTimeoutMs and fix disabled-timeout hint

Address review on the configurable foreground shell timeout:

- Config: validate shellDefaultTimeoutMs at construction, mirroring
  visionBridgeTimeoutMs, but allow 0 (disables the timeout). Negative,
  fractional, or out-of-range values now coerce to undefined instead of
  reaching AbortSignal.timeout() via a hand-edited settings.json that
  bypasses schema validation.
- settingsSchema: mark tools.shell.defaultTimeoutMs requiresRestart, since
  Config.shellDefaultTimeoutMs is private readonly with no setter, so a
  mid-session change cannot take effect.
- shell: when the timeout is disabled (effectiveTimeout === 0), suppress
  the long-run backgrounding hint instead of firing it on every command
  over ~1s via the longRunThresholdFor floor.
- shell: correct the precedence comment; 0 disables only at the
  settings/default level, as the per-call timeout param rejects <= 0.

Add coverage for negative/fractional coercion to the built-in default and
for 0 disabling the timeout without emitting the spurious hint.

---------

Co-authored-by: Shaojin Wen <shaojin.wensj@alibaba-inc.com>
2026-07-12 06:10:29 +00:00
ever-o
ec43c1e951
feat(web-shell): render composer references in user messages (#6537)
* feat(web-shell): render composer references in user messages

* refactor(web-shell): consolidate composer tag utilities

* fix(web-shell): leave custom references as text

* fix(web-shell): avoid ambiguous reference chips

* fix(web-shell): thread composer tag icons to messages

* feat(web-shell): render user references from annotations

* fix(web-shell): include inline tags in input annotations

* fix(web-shell): remove duplicate composer tag icon option

* fix(web-shell): forward plan prompt annotations

* test(web-shell): cover composer annotation edge cases

* fix(web-shell): forward split pane prompt annotations

* fix(web-shell): guard malformed input annotations

---------

Co-authored-by: zhanghuapeng.zhp <zhanghuapeng.zhp@alibaba-inc.com>
Co-authored-by: 易良 <1204183885@qq.com>
Co-authored-by: Shaojin Wen <shaojin.wensj@alibaba-inc.com>
2026-07-11 17:44:47 +00:00
jinye
51d4ce48db
feat(serve): persist dynamic workspace registrations (#6716)
* feat(serve): persist dynamic workspace registrations

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* codex: address PR review feedback (#6716)

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* codex: address PR review feedback (#6716)

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* codex: address PR review feedback (#6716)

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* codex: address PR review feedback (#6716)

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* codex: address PR review feedback (#6716)

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

---------

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>
2026-07-11 16:49:40 +00:00
易良
263dbba741
fix(core): preserve managed memory during microcompaction (#6714)
* fix(core): preserve managed memory during microcompaction

Refs #6487

* test(core): cover managed memory read errors

---------

Co-authored-by: Shaojin Wen <shaojin.wensj@alibaba-inc.com>
2026-07-11 15:58:54 +00:00
qqqys
58dc985ed8
chore: remove DingTalk planning artifacts (#6722)
Co-authored-by: Shaojin Wen <shaojin.wensj@alibaba-inc.com>
2026-07-11 15:47:05 +00:00
jinye
e403246dc2
feat(serve): Expose read-only untrusted session catalogs (#6717)
* feat(serve): expose read-only untrusted session catalogs

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* refactor(cli): address session catalog review feedback

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* codex: address PR review feedback (#6717)

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* codex: address PR review feedback (#6717)

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* codex: address PR review feedback (#6717)

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

---------

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>
Co-authored-by: qwen-code-dev-bot <qwen-code-dev-bot@users.noreply.github.com>
2026-07-11 15:28:41 +00:00
Nothing Chan
b19ebd8fc6
fix(packaging): bundle clipboard addon in standalone builds (#6708) 2026-07-11 15:18:24 +00:00
Alex Yanchenko
218dec6a6f
feat(hooks): add MessageDisplay hook for mid-turn streaming (#6489)
* feat(hooks): add MessageDisplay hook for mid-turn streaming

Fires repeatedly as the assistant reply streams, before Stop (which only fires once at the end of the turn). Fire-and-forget, cumulative text payload, debounced (~200ms) except for the unconditional final firing. Fires from the single streaming loop in client.ts shared by the terminal UI and ACP paths.

Fixes #6488

* fix(hooks): address MessageDisplay review feedback

- Chain fire-and-forget MessageDisplay requests per message_id instead of
  firing them fully unbounded, so a slow hook command can't pile up
  concurrent processes.
- Gate the final flush on non-empty displayed_text and !signal.aborted,
  matching the adjacent Stop hook's guard.
- Document why the final flush intentionally re-sends the last debounced
  text (is_final itself is new information).
- Simplify the debounce constant's JSDoc to drop the competitor comparison.
- Add tests for the mid-stream debounced flush and the rejected-request
  warn path.

* test(hooks): drain microtasks before asserting on chained MessageDisplay calls

fireMessageDisplayHook now chains per-message_id through a promise (see
previous commit), so the final flush's actual messageBus.request() call
lands a few microtask ticks after the generator itself finishes — the
mid-stream-flush test needs to let that chain settle before asserting.

* fix(hooks): flush MessageDisplay is_final on every for-await exit path

The three early `return turn` paths inside the streaming loop (always-on
loop-detection safety, heuristic loop detection, and the stream Error event)
exited before the final MessageDisplay flush, which only sat after the loop
ended normally. Hook scripts relying on is_final: true to know when to flush
never received it when a turn ended via loop detection or an API error.

Extracts the flush into a shared closure and calls it from all four exits
(the three early returns plus the normal fall-through), instead of only the
one at the bottom of the loop. Adds regression tests for all three previously
missed exits, plus the two guard-coverage tests requested in review (abort
suppresses the flush, a tool-call-only turn with no Content events does not
fire a vacuous empty-text event).

Addresses the outstanding critical review comment and the follow-up test
coverage suggestion on PR #6489.

* fix(hooks): fire MessageDisplay on the ACP surface, coalesce delivery, drain is_final before turn end

Addresses the three findings from the local verification report on #6489:

- ACP/qwen serve (Finding 1): the delivery logic now lives in a shared
  MessageDisplayDispatcher (packages/core), and Session.ts wires it into
  all four raw-stream loops (main prompt, Stop-hook continuation, cron
  tick, background notification) — these surfaces consume GeminiChat's
  stream directly and never enter GeminiClient.sendMessageStream, so
  they need their own fire sites. The daemon no longer advertises an
  event it never emits.

- Slow-hook backlog (Finding 2): the per-message promise chain is
  replaced by coalescing delivery — at most one in-flight request plus
  one pending payload per message; newer flushes overwrite the pending
  slot, which is lossless because displayed_text is cumulative, and
  is_final is sticky. A slow hook now sees fewer, newer payloads instead
  of an ever-growing queue of stale ones.

- Headless is_final drop (Finding 3): finish() resolves only once every
  enqueued payload has actually been delivered, and every exit out of
  the streaming loops awaits it (early returns, normal fall-through,
  and the enclosing finally for uncaught exceptions), so a short-lived
  -p process can no longer exit with the final payload still queued.
  As a consequence, is_final delivery now strictly precedes the Stop
  hook rather than racing it.

Also: the failure log line carries the message_id, finish() is
idempotent, the review-requested tests are added (mid-stream and final
firings share one message_id; isFinal as the sole flush reason), and
hooks.md gains a delivery-semantics contract covering coalescing, the
drain guarantee, no is_final on cancellation, provisional
displayed_text, and multiple messages per tool-using turn.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(hooks): bound MessageDisplay drain wait, fix test gaps flagged in review

finish() now gives up waiting on drain after 5s (MESSAGE_DISPLAY_DRAIN_TIMEOUT_MS) instead of blocking turn teardown for up to the full 60s hook timeout, per the re-verification's S1 finding. Delivery keeps running in the background past the timeout; only the caller's wait is bounded.

Also: add the config.ts bridge test for MessageDisplay field extraction (S5), and add the missing MessageDisplay/InstructionsLoaded entries to acpAgent.test.ts's HookEventName mock (S6).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(hooks): dispatch MessageDisplay is_final alongside stale deliveries, share one drain budget

Round-3 review findings on #6489:

- finish() no longer queues the is_final payload behind an in-flight
  mid-stream delivery: the pending slot's supersession argument applies to
  the in-flight slot too, so the final payload is dispatched immediately,
  alongside the stale delivery if one is still running. is_final is handed
  to the hook the moment the message ends — before Stop — on every surface,
  and can no longer be dropped by a short-lived process exiting with it
  still queued (Finding 1).
- The bounded drain wait is memoized: every finish() call (explicit,
  finally, or concurrent) shares one promise and one timer, so the teardown
  ceiling is MESSAGE_DISPLAY_DRAIN_TIMEOUT_MS itself, not a multiple of it
  (Finding 2).
- hooks.md delivery semantics rewritten to match the shipped behavior,
  including the headless orphaned-hook caveat and the unspecified completion
  order between an overlapped stale execution and the final one (Finding 3).
- The dispatcher mirrors its warnings to console.warn itself (stderr on
  headless/ACP, ink patchConsole in the TUI) in addition to the injected
  debug-file sink, so hitting the drain timeout is visible by default
  (Finding 4).
- A superseded mid-stream delivery that fails after the final was
  dispatched no longer warns; failures during streaming still do.
- New tests: finish() twice while delivery is in flight (the exact
  client.ts sequence), concurrent finish() calls sharing one budget,
  is_final overtaking a held mid-stream delivery, and drain resolving on
  the final delivery alone.

* refactor(core): consolidate MessageDisplay finish() calls, dedupe test spy setup

client.ts: wrap the turn.run() streaming loop in try/finally so messageDisplay.finish() fires once instead of at each of the three early-return sites plus the post-loop path -- matching the pattern the four raw-stream loops in Session.ts already use for the same dispatcher.

message-display-dispatcher.test.ts: centralize the console.warn spy setup/teardown in beforeEach/afterEach instead of five repeated per-test try/finally blocks.

No behavior change: full client.test.ts (246/246) and the message-display-buffer/dispatcher suites (24/24) pass unchanged.

* docs(hooks): clarify MessageDisplay cancellation timing (round-4 nit)

* test(hooks): cover the 3 untested MessageDisplay dispatch sites, fix cancellation doc wording

Adds MessageDisplay is_final coverage for the Stop-hook continuation loop, the in-session cron fire, and the background-notification loop, each with a normal-completion and an abort case. Adds three MessageDisplayDispatcher edge-case tests: a delivery settling just before the drain timeout, an abort arriving after a drain wait has already started, and addChunk called after abort but before finish(). Rewords the cancellation-timing doc bullet to state the actual criterion (abort signal state when finish() runs) rather than an approximation of it.

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: Shaojin Wen <shaojin.wensj@alibaba-inc.com>
2026-07-11 13:24:32 +00:00
jinye
230db17650
feat(cli): group daemon channel workers by workspace (phase 4b) (#6635)
* feat(cli): group daemon channel workers by workspace (phase 4b)

Multi-workspace `qwen serve --channel` now runs one channel worker per owning workspace instead of a single primary-bound worker. Each worker binds to its workspace's directory, daemon-workspace env marker, and effective env overlay. Channels are grouped implicitly by their configured working directory: a channel belongs to the registered workspace its resolved cwd matches, mirroring the worker's own workspace validation. Unknown, ambiguous, or untrusted targets fail fast at startup.

The pidfile and daemon status grow an additive per-workspace worker list while keeping the existing single-worker fields for older readers; single-workspace daemons stay byte-identical to before. `--channel all` stays primary-only.

Refs #6378

* fix(cli): harden multi-workspace channel workers

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* fix(cli): close listener after channel worker startup failure

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* fix(cli): restore grouped channel webhooks

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* fix(cli): mount runtime before channel workers start

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* codex: address PR review feedback (#6635)

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* test(cli): strengthen channel worker edge coverage

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* chore(cli): address channel review suggestions

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

---------

Co-authored-by: qwen-code-dev-bot <qwen-code-dev-bot@users.noreply.github.com>
Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>
Co-authored-by: Shaojin Wen <shaojin.wensj@alibaba-inc.com>
2026-07-11 13:08:02 +00:00
jinye
6df19602c4
fix(cli): Scope session organization mutations by workspace (#6724)
* fix(cli): scope session organization mutations by workspace

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* codex: address PR review feedback (#6724)

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* codex: address PR review feedback (#6724)

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

* codex: address PR review feedback (#6724)

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

---------

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>
2026-07-11 12:25:16 +00:00
Heyang Wang
5f89459e90
feat(core): add unified session recovery planning (#6731)
Centralize resume recovery classification so each entrypoint can share
the same decision about interrupted prompts, dangling tool calls, and
history gaps.

- Add a core SessionRecoveryPlan builder with provider-safe repaired history
- Route headless, stream-json control, and ACP continue paths through it
- Show a TUI resume notice when an interrupted tool turn is detected
- Document the Recovery Service design and cover the core plan with tests

Co-authored-by: heyang.why <heyang.why@alibaba-inc.com>
2026-07-11 11:15:40 +00:00
destire-mio
7279d6f360
feat(cli): add project-scoped prompt stash (#6709)
Co-authored-by: zhuyouwei <zhuyouwei@testin.cn>
2026-07-11 10:36:06 +00:00