mirror of
https://github.com/QwenLM/qwen-code.git
synced 2026-08-13 10:45:27 +00:00
7731 commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
68c7cbc980 | test(cli): add multi-chunk raw paste and passthrough Ctrl+C escape tests (#6506) | ||
|
|
8d18426426
|
Merge branch 'main' into worktree-fix-paste-perf | ||
|
|
e7097d0ef6
|
feat(sdk-java): Add daemon transport (#7463)
* feat(sdk-java): add daemon transport Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com> * codex: address PR review feedback (#7463) Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com> * test(sdk-java): stabilize lifecycle lock test Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com> * codex: address PR review feedback (#7463) Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com> * codex: address PR review feedback (#7463) Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com> * codex: address PR review feedback (#7463) Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com> * fix(cli): preserve prompt cancellation during context propagation Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com> --------- Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com> |
||
|
|
afacef5a38
|
fix(core): preserve disabled reasoning effort (#7541)
Co-authored-by: 易良 <1204183885@qq.com> |
||
|
|
ff46cbe57d
|
fix(core): strip daemon secrets from hook and tool-discovery child env (#7527)
* fix(core): strip daemon secrets from hook and tool-discovery child env Follow-up to #7256, which introduced sanitizeChildEnv and applied it to the shell, monitor, and MCP stdio spawn paths. Three agent-launched child processes were left inheriting the full environment: - hooks/hookRunner.ts spreads process.env into the hook command's env - tools/tool-registry.ts spawns the configured tool-call command and the tool-discovery command with no env option, so both inherit implicitly All three run user- or config-supplied commands on the agent's behalf and have no need for QWEN_SERVER_TOKEN / QWEN_DAEMON_TOKEN, so they are the same credential-exposure gap #6601 describes. Route each through sanitizeChildEnv(process.env); benign inherited env is unchanged. * fix(core): normalize Windows PATH on the tool-registry child env Both tool-registry spawns previously passed no env option, so Node inherited the parent environment natively and Windows resolved its case-insensitive PATH keys itself. Passing env explicitly gives that up: on Windows process.env can carry both Path and PATH, and the child may pick the wrong one. Route them through normalizePathEnvForWindows, matching the shell (shellExecutionService.ts) and MCP stdio (mcp-client.ts) spawn sites that already pair it with sanitizeChildEnv. It returns env untouched off win32, so nothing changes on other platforms. hookRunner.ts is left as-is: it already built an explicit env before this branch, so it does not regain inheritance semantics here. --------- Co-authored-by: Shaojin Wen <shaojin.wensj@alibaba-inc.com> |
||
|
|
baaedfbf7d
|
fix(core): jsonl write([]) should leave an empty file, not a stray newline (#7533)
write() built its content with join('\n') and then appended a trailing
newline. For an empty array the join produces '', so the file ends up as
a single '\n': one byte, no records.
That makes the module's own accessors contradict each other — exists()
tests size > 0 and returns true, while read() skips blank lines and
returns []. Clearing a JSONL file therefore leaves something that reports
as non-empty but has nothing in it.
Terminate each record instead of joining with separators. The output for
a non-empty array is byte-identical; an empty array now writes nothing.
Co-authored-by: Shaojin Wen <shaojin.wensj@alibaba-inc.com>
|
||
|
|
abc9666f77
|
fix(web-shell): prevent React measure detail OOM (#7596)
Co-authored-by: ytahdn <ytahdn@gmail.com> |
||
|
|
1832a45ab8
|
fix(core): reject nested background requests (#7593) | ||
|
|
aa11e8d126
|
fix(web-shell): initialize workspace selector from ID (#7518)
* fix(web-shell): initialize workspace selector from ID * Update packages/web-shell/client/components/WorkspaceSessionProvider.tsx fix: adjust initial workspace directory resolution logic to support embeddings passing only unlocked workspaceCwd Replace condition check from effectiveWorkspaceId to !lockWorkspaceCwd && targetWorkspace, prevent sessions from incorrectly initializing under primary workspace for secondary workspace path inputs Co-authored-by: qwen-code-ci-bot <qwen-code-ci@service.alibaba.com> --------- Co-authored-by: qwen-code-ci-bot <qwen-code-ci@service.alibaba.com> |
||
|
|
fcc250beb5
|
fix(core): record auto-memory index reads in FileReadCache (#7468)
* fix(core): record auto-memory index reads * fix(core): seed memory read cache from read stats |
||
|
|
5cad7fe642
|
feat(autofix): update a stale base when the gate rejects a behind-main fix (#7595)
A fix that fails to build is not always the fix's fault. #7471 stalled when the agent's verification gate failed with "Cannot find module 'update-notifier'" — a dependency main removed in #7515, still imported on a branch 32 commits behind. The loop could not tell a stale-base build failure from a genuine one, so it advanced past the feedback and asked a human to take over. In the gate-rejection branch, before the handoff, compare the checked-out head with main; if it is behind or diverged, update-branch (a CAS on REPORT_HEAD) merges main in and the round retries (sentinel ts keeps the feedback live). It self-limits: after the update the PR is current, so a next-round rejection is no longer "behind" and falls through to the human handoff — a genuine fix failure costs at most one base-update. The round is exempt from the consecutive-failure breaker (not the PR's fault), and every API call is fail-safe. This is the agent-gate sibling of #7554, which only sees PR status checks, never the gate's own build. Co-authored-by: wenshao <wenshao@example.com> |
||
|
|
819cd4ab4a
|
fix(core): retry requests when providers require thinking (#7534)
* fix(core): retry requests that require thinking * test(core): cover required-thinking retry failure * fix(core): retry required thinking for chat templates * fix(core): remove rejected thinking opt-outs on retry * test(core): cover aborted required-thinking retry * test(core): cover required-thinking retry edges * chore: trigger pr synchronization --------- Co-authored-by: Shaojin Wen <shaojin.wensj@alibaba-inc.com> |
||
|
|
77af061e68
|
fix(core): persist usage for tool-only subagent rounds (#7557)
* fix(core): persist usage for tool-only subagent rounds * test(core): cover drop path for empty usageless subagent round --------- Co-authored-by: Claude <noreply@anthropic.com> |
||
|
|
779e10622b
|
perf(cli): Defer ACP telemetry initialization (#7558)
* perf(cli): defer ACP telemetry initialization Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com> * fix(cli): preserve deferred ACP metrics Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com> --------- Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com> |
||
|
|
86103203e4
|
fix(cli): correct queued message display style and ordering (#7381)
* fix(cli): correct queued message display style and ordering
Mid-turn steer messages (user input queued while the model is
responding) had two display bugs:
1. They rendered with notification styling (● icon) instead of
user-input styling (> prefix) because accept() added them to
UI history as MessageType.NOTIFICATION.
2. They appeared below the model's reply because accept() was
only called in the finally block after the entire response
stream completed, appending the user message after all model
response items.
Fix: use MessageType.USER with sentToModel: true for steer
messages, and settle the steer input on the first stream event
(after the user-content push lands but before model-response
events are committed to UI history). Pass steer inputs through
to recursive sendMessageStream calls so all takeSteerInput paths
benefit from early settlement. Add a WeakSet guard to
settleSteerInput for idempotency across recursive invocations.
* test(core): add ordering test for early steer settlement
Verify that accept() is called after the first stream event is
pulled but before subsequent events reach the consumer, pinning
the settle-before-content timing that ensures queued user
messages render above the model's reply.
* fix(cli): use sentToModel: false for steer messages, address review
- Use sentToModel: false instead of true: steer messages are injected
into an existing tool-result turn, not standalone user turns.
sentToModel: true would make isRealUserTurn() count them as real
turns, inflating the rewind turn index.
- Remove unnecessary as HistoryItemWithoutId cast.
- Add post-cleanup assertion in ordering test to verify the WeakSet
guard prevents double-settlement.
* fix(cli): align resumed mid-turn steer display with live session (#7381)
Resume path now renders mid_turn_user_message as MessageType.USER with
sentToModel: false, matching the live-session styling. Add a comment
documenting the intentional sentToModel: false choice.
* fix(cli): exclude steer messages from user-turn filters (#7381)
Steer messages (sentToModel: false) were counted as real user turns by
five downstream consumers that filter on type === 'user' without checking
sentToModel, breaking cancel auto-restore, telemetry turn count, prompt
recall, away-recap thresholds, and resume collapse boundaries.
Add sentToModel !== false guards at each site.
* test(cli): add coverage for sentToModel !== false guards (#7381)
* test(cli): add coverage for sentToModel !== false guard in input-history filter (#7381)
* test(cli): add coverage for sentToModel !== false guard in YOLO turn-count telemetry (#7381)
* fix(cli): restore corrupted docs and classify steer items as synthetic (#7381)
* fix(docs): restore corrupted autogenerated input names in GitHub Action docs (#7381)
* fix(cli): deduplicate findLastUserItemIndex and add steerInput forwarding test (#7381)
* fix(cli): keep code-block copy numbering continuous across steer items (#7381)
* test(core): add Hook continuation steerInput forwarding test
Verify that steerInput is forwarded through the Stop-hook
continuation path and settled early on the first content event
of the continuation turn, matching the existing Steer
continuation coverage.
* fix(cli): sync selection test fixtures with ink FrameCell/ReadonlyFrame types (#7381)
* fix(core): align cron day wildcard semantics (#7464)
Co-authored-by: destire-mio <248462155+destire-mio@users.noreply.github.com>
* feat(core): keep completed background agents resident (#7426)
* feat(core): keep background agents resident
* fix(core): harden background continuation boundaries
* docs(core): move per-spawn cleanup comment to subagentDispose
The comment describing the per-spawn cleanup (which stays undefined on
the fork-resume path) had drifted above the launchModel declaration,
where it no longer applied and could mislead readers. Relocate it to the
subagentDispose assignment in the non-fork branch it actually documents.
* fix(core): close finishing window and release resident on error in background GOAL path
- Non-worktree GOAL completion drained the message queue but never called
registry.beginFinishing(), unlike the worktree path. A send_message racing
the terminal transition could be accepted (status still running,
finishingAgents empty) and then orphaned by complete(). Call beginFinishing()
after the empty drain to reject the racing message instead.
- The completion catch block never reset keepResident, so a throw from
patchAgentMeta/registry.complete left the runtime resident but finalized as
failed — a zombie that cleanupRuntime never reclaimed. Reset keepResident in
the catch so the finally block disposes it.
---------
Co-authored-by: Claude <noreply@anthropic.com>
* ci(autofix): continue environment-specific fixes (#7444)
* ci(autofix): continue environment-specific fixes
* docs(autofix): align verification wording
* docs(autofix): require bundle before integration tests
* docs(autofix): scope surrogate verification rules
* docs(autofix): require focused tests before integration checks
* docs(autofix): clarify review verification guidance
* fix(acp-bridge): close prompt-terminal follow-ups from the PR #7400 self-review (#7453)
* fix(acp-bridge): close prompt-terminal follow-ups from PR #7400 self-review
Keep a removed RUNNING prompt visible to the teardown flush via a removed flag so its terminal still publishes when the session closes before the agent cooperates; gate broadcastTurnError's session turn-state mutation to running prompts; propagate the typed PromptDeadlineExceededError from the pre-dispatch abort check; document the deadline FIFO-release overlap trade-off, the trailing prompt_cancelled after flush, and the result.then/finally ordering invariant; route the dedup log to the debug channel; drop the prompt-deadline re-export that pulled the bridge into a leaf module.
Fixes #7451
* test(acp-bridge): cover promote-then-remove-then-settle duplicate completed guard (#7453)
---------
Co-authored-by: Qwen Code Bot <qwen-code-bot@users.noreply.github.com>
* fix(core): strip Qwen-internal daemon secrets from agent-spawned child env (#7256)
* fix(core): strip Qwen-internal daemon secrets from agent-spawned child env
Shell subprocesses (and the monitor tool and stdio MCP servers) inherited
the full daemon process.env, including QWEN_SERVER_TOKEN (the serve-daemon
bearer credential), so an agent-run command like printenv QWEN_SERVER_TOKEN
could read an internal secret. Add a shared sanitizeChildEnv() that removes
Qwen-internal daemon/server tokens (QWEN_SERVER_TOKEN, QWEN_DAEMON_TOKEN)
before spawning, and apply it at the shell child_process + PTY paths,
monitor.ts, and the mcp-client stdio transport.
The denylist is deliberately narrow: it does NOT strip third-party
credentials (GH_TOKEN, AWS_*, NPM_TOKEN, ...) that real shell workflows
legitimately inherit -- only Qwen-internal secrets. Exported from the
package root so the desktop denylists can consolidate onto it later.
Fixes #6601.
* test(core): cover daemon-secret stripping on monitor and mcp-client spawn sites
* test(core): replace process.env instead of mutating in shell sanitization tests
The file restores process.env by reference in afterEach, so in-place key
mutations leaked into later tests. Use the replacement pattern already used
by setupConflictingPathEnv.
* docs(core): align JSDoc @param names with actual function signatures (#7492)
Fix 6 instances where JSDoc @param tags had drifted from their
corresponding function signatures — parameters were renamed, removed,
or undocumented over time but the doc blocks were not updated.
Closes #7446
* feat(serve): support forced MCP reconnects (#7488)
* feat(serve): support forced MCP reconnects
* test(serve): cover forced MCP reconnect options
---------
Co-authored-by: 克竟 <dingbingzhi.dbz@alibaba-inc.com>
* fix(cli): insert newline on Shift+Enter and stop streaming thinking-block flicker (#7397)
* fix(cli): re-push Kitty keyboard flags onto the alternate screen in VP mode
In VP mode the app renders on the alternate screen (`alternateScreen: true`),
but the Kitty keyboard progressive-enhancement flags were pushed only once at
startup on the main screen. The Kitty spec tracks these flags per screen
buffer, so the alternate screen's stack stays empty and the terminal never
reports modifiers: Shift+Enter arrives as a bare Enter (submit) or, when the
terminal emits an ESC-prefixed variant, as an orphaned Escape that trips the
empty-buffer double-Esc rewind prompt — so Shift+Enter can never insert a
newline in VP mode even on Kitty-capable terminals (e.g. cmux).
Re-push the flags onto the alternate screen right after Ink enters it (Ink
writes the enter-alt-screen sequence synchronously inside render(), so the
push is correctly ordered). Ink discards the alternate screen and its flag
stack on unmount, leaving the startup main-screen push balanced by the
existing disableKittyProtocol() on cleanup.
Generated with AI
Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>
* fix(cli): stabilize streaming thinking block height to stop flicker
The pending "Thinking…" block renders the tail of the reasoning stream in a
content-sized box. As the model emits paragraph separators, a blank line
enters and leaves the tail window (and `trimEnd` drops trailing blanks), so the
visible line count oscillates and the block flickers 2→3→5 rows during
streaming.
Track the tallest height the block has reached for the current thought and
never render fewer rows than that (capped at the streaming window size),
padding at the top so the newest line stays pinned to the bottom. The tracker
resets when streaming ends or when the buffer shrinks (a new thought replaced
it), so height is monotonic within a thought without leaking across thoughts.
Generated with AI
Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>
* fix(cli): decode xterm modifyOtherKeys Shift/Ctrl/Alt+Enter so it inserts a newline
Terminals such as Ghostty report Shift+Enter as the xterm modifyOtherKeys
sequence `ESC [ 27 ; <mods> ; <key> ~` (e.g. `ESC [ 27 ; 2 ; 13 ~`) when the
Kitty keyboard protocol is not negotiated — which is the default, since Kitty
detection does not always succeed. Two bugs kept this from inserting a newline:
1. The CSI-u parser read the leading `27` marker as the key code (matching the
Escape key code 27) instead of the real key code in the third parameter, so
with Kitty enabled Shift+Enter was mistaken for Escape and tripped the
double-Esc rewind prompt.
2. The reassembly path that stitches readline's shredded CSI fragments back
together was gated behind `kittyProtocolEnabled`, so with Kitty disabled the
`ESC [ 27 ; 2 ;` head plus the stray `13~` tail leaked into the composer as
literal text and no newline was inserted.
Decode the third parameter as the real key code for the `27;…~` form, and route
those sequences through the reassembly buffer even when Kitty is off (only the
`ESC [ 27` marker opts in, so keys readline already parses cleanly are
untouched). Shift/Ctrl/Alt+Enter now insert a newline in both VP and non-VP
mode regardless of Kitty negotiation.
Generated with AI
Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>
* fix(cli): anchor VP viewport to the top until a conversation turn exists
On a fresh VP-mode session the virtualized list holds the banner plus startup
notices (tips / MOTD / info), so it is longer than one item. Keying the initial
scroll anchor off list length alone selected scroll-to-end, which pinned the
banner to the bottom of the full-height viewport and left the top half of the
screen blank.
Anchor to the top until there is an actual conversation turn (a user/user_shell
history item or a pending response), then resume scroll-to-end so the latest
output stays in view. Startup notices no longer count as content that forces
bottom alignment.
Generated with AI
Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>
* fix(cli): stabilize streaming thinking window against availableTerminalHeight drift
The grow-only streaming thinking window still flickered because its line cap was
derived from availableTerminalHeight. While a thought streams the terminal keeps
constrainHeight on, so availableTerminalHeight (and the derived maxLines) drifts
up and down as sibling pending content grows, and the grow-only clamp
`min(maxLines, …)` shrank the block whenever it dipped.
Use a constant window height (MAX_STREAMING_THINKING_VISUAL_LINES) for the
pending window instead. The window is only a few lines, so a fixed cap cannot
meaningfully overflow (VP scrolls anyway), and the height stays stable while
still growing monotonically within a thought.
Generated with AI
Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>
* Revert "fix(cli): anchor VP viewport to the top until a conversation turn exists"
This reverts commit
|
||
|
|
9d029835fc
|
test(autofix): single-source the infra-signature list from the workflow (#7565)
* feat(autofix): auto-rerun a check that died on infrastructure, once A failed check can be red because the machine died, not the code — a self-hosted runner losing the server, the disk filling. #7490's E2E failed with "runner lost communication with the server" and went green on a rerun. The scan now reruns such a check's failed jobs automatically. Detection is a conservative annotation whitelist (INFRA_FAILURE_SIGNATURES) — only unambiguous machine failures, never a test-level timeout, which could be a real regression. The one-shot guard is run_attempt, not a marker: a run already retried to attempt 2 and still infra-failing is persistent, so it is left for a human; after a rerun the attempt increments, so the next scan will not rerun it. Every step is fail-safe (any API error → no rerun), it runs only when the PR actually has a failed check, and the gate carries the same review-address carve-out as the other check selectors so the loop never reruns its own runs. This is the transient-infra sibling of #7554 (stale-base): that merges current main when a check is base-inherited; this reruns when a check died on the runner. Neither touches a check that is a genuine failure. Note: rerun-failed-jobs needs the PAT to hold `actions: write`. * fix(autofix): use POSIX ERE groups in infra-failure regex, cover all signatures in tests (#7562) * fix(autofix): also treat a git fetch/clone transport death as infra #6506's checkout died mid-transfer — "fetch-pack: invalid index-pack output" and "RPC failed; curl 92 ... CANCEL" — which then hung the job into the 20m limit. That is infra, not the PR (it only touches a doc), and a re-run made it green. But the infra-signature whitelist did not cover it, so the auto-rerun did not fire and it waited on a human. Add `invalid index-pack output` and `RPC failed` — the two canonical git-transport-death phrases — to INFRA_FAILURE_SIGNATURES. A co-present job-timeout line does not block the match (one matching line classifies the run), and a BARE timeout with no transport signature is still left alone, since it can be a real regression. Both new signatures are pinned in the test's per-signature loop, plus a case on #6506's real composite annotation and a bare-timeout-is-not-rerun guard. * test(autofix): single-source the infra-signature list from the workflow The infra-rerun test re-typed INFRA_FAILURE_SIGNATURES as an inline mirror of the workflow's env value. Two copies that must be hand-synced can drift — the test could keep passing against a stale list while production changed, or vice versa. That is exactly the copy the git-transport follow-up had to remember to update in two places. Extract the list from the workflow source instead, the same extract-from-source idiom the file already uses for NON_BLOCKING_CHECKS, so there is only one copy and drift is impossible. A toContain guard fails loudly if the env is renamed or the regex breaks, rather than letting an empty pattern match every line and silently pass. * fix(autofix): paginate annotations and filter Autofix runs in infra-rerun loop (#7562) --------- Co-authored-by: wenshao <wenshao@example.com> Co-authored-by: Qwen Code Bot <qwen-code-bot@users.noreply.github.com> Co-authored-by: qwen-code-ci-bot <qwen-code-ci-bot@users.noreply.github.com> |
||
|
|
c5dddb5e5d
|
fix(sdk-python): validate max_tool_calls and max_subagent_depth as integers (#7548)
* fix(sdk-python): validate max_tool_calls and max_subagent_depth as integers max_session_turns rejects bools and non-ints before its range check, but the two sibling limits only range-checked, so values their own error messages rule out were accepted and stringified onto the CLI: max_tool_calls=True -> --max-tool-calls True max_tool_calls=2.5 -> --max-tool-calls 2.5 max_subagent_depth=True -> --max-subagent-depth True bool is the sharp edge: it subclasses int, so True passes both isinstance and the range comparison as 1. Extract the bool-aware check the max_session_turns branch already performed into _is_int and apply it to all three. max_session_turns keeps its exact previous behavior; the expression is the same, only named. * fix(sdk-python): say integer in the max_subagent_depth range message A float like 2.5 is between 1 and 100, so the old message described a constraint the value already satisfied and gave no hint that the type was the problem. --------- Co-authored-by: Shaojin Wen <shaojin.wensj@alibaba-inc.com> |
||
|
|
c0fee1b609
|
feat(web-shell): add workspace agent management (#7572)
* feat(web-shell): add workspace agent management * test(web-shell): remove unstable extensions page tests * fix(agents): address management review feedback --------- Co-authored-by: ytahdn <ytahdn@gmail.com> |
||
|
|
74a786dac3
|
fix(web-shell): isolate slash command plugin pages (#7581)
Co-authored-by: ytahdn <ytahdn@gmail.com> |
||
|
|
93af84f248
|
fix(cli): surface unhandled rejections and render errors instead of swallowing them (#7406)
* fix(cli): surface unhandled rejections and render errors instead of swallowing them Unhandled promise rejections were emitted to AppEvent.LogError which has no production listener, making crashes completely invisible in debug logs and stderr. Additionally, the main App tree had no top-level ErrorBoundary, so React render errors were caught only by Ink's internal boundary which silently exits via exitPromise.catch(noop). - Write unhandled rejection details to the debug logger so they appear in ~/.qwen/debug/<session>.txt - Wrap the interactive UI tree in an ErrorBoundary that logs fatal render errors and shows a fallback message instead of silent exit * fix(cli): address review feedback — normalize non-Error throws and add exit path - Normalize non-Error thrown values (strings, null, etc.) to Error instances in ErrorBoundary before logging/rendering, preventing secondary TypeErrors from sanitizeTerminalText(undefined) - Schedule a graceful exit (5s delay) from the fatal render error fallback so the session does not hang under the Kitty keyboard protocol where Ctrl+C is a keypress, not SIGINT * test(cli): add test for non-Error thrown value normalization in ErrorBoundary --------- Co-authored-by: 秦奇 <gary.gq@alibaba-inc.com> |
||
|
|
e58f995fae
|
fix(core): avoid empty transcript history pages (#7582)
Co-authored-by: 钉萁 <dingqi.jww@alibaba-inc.com> |
||
|
|
3b3a593600
|
Fix(cli): use npm view for update check instead of update-notifier (#7515) (#7528)
* fix(cli): use npm view for update check instead of update-notifier (#7515) * fix(cli): use npm view for update check instead of update-notifier (#7515) * fix(cli): accept array-wrapped npm view output in update check (#7515) npm 11+ prints `npm view <pkg> dist-tags.<tag> --json` as ["0.20.1"] instead of "0.20.1", so the strict string check re-broke the update check with "Invalid npm latest version response". Accept both shapes. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * chore(cli): drop update-notifier dependency and dead install-detection code (#7515) Version checking now goes through npm view for every install type, so: - replace the update-notifier UpdateInfo type import with a local interface and remove update-notifier / @types/update-notifier from dependencies - remove isGlobalNpmInstallation and looksLikeNpmPackagePath, which no longer have any production callers, along with their tests Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs(cli): fix npm version in array-output comment (npm 12+, not 11+) --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> Co-authored-by: Shaojin Wen <shaojin.wensj@alibaba-inc.com> |
||
|
|
b7b5ed00e8
|
fix(cli): prevent monitor turns after task_stop (#7573)
Some checks are pending
E2E Tests / E2E Test (Linux) - sandbox:docker (push) Waiting to run
E2E Tests / E2E Test (Linux) - sandbox:none (push) Waiting to run
E2E Tests / E2E Test - macOS (push) Waiting to run
E2E Tests / cron-interactive E2E (nightly) (push) Waiting to run
E2E Tests / web-shell Browser Regression (push) Waiting to run
SDK Python / Classify PR (push) Waiting to run
SDK Python / SDK Python (3.10) (push) Blocked by required conditions
SDK Python / SDK Python (3.11) (push) Blocked by required conditions
SDK Python / SDK Python (3.12) (push) Blocked by required conditions
|
||
|
|
545e924e35
|
fix(serve): avoid TOCTOU race dropping live sessions from list response (#7556)
* Initial plan * fix(serve): avoid TOCTOU race dropping live sessions from list response --------- Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com> Co-authored-by: 易良 <1204183885@qq.com> |
||
|
|
3617397e5f
|
feat(core): Align GenAI telemetry with ARMS (#7536)
* feat(core): align GenAI telemetry with ARMS Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com> * fix(core): remove estimated token usage splits Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com> * fix(core): address GenAI telemetry review feedback Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com> --------- Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com> Co-authored-by: Shaojin Wen <shaojin.wensj@alibaba-inc.com> |
||
|
|
947b632140
|
fix(serve): detect stale SSE cursors across daemon restarts via epoch token; preserve turn attribution and surface compaction failures in replay (#7458)
* fix(daemon): epoch-token restart detection, compaction attribution, and degraded-snapshot signaling (DAEMON-001/007/008) * fix(acp-bridge): field-level turn attribution merge and replayDegraded bridge test (#7458) * fix(serve): skip bus epoch lookup for virtual subagent SSE streams (#7458) The REST SSE route looked up the bus epoch for every session id, but virtual subagent sessions ride their own bus and their compound ids are not in the bridge's byId map, so the lookup threw and aborted the subscription — breaking subagent event streams. Skip the lookup for the virtual path and degrade a torn-down real session to a headerless stream (mirrors the /acp route). Also bumps the daemon browser SDK bundle budget (167KB -> 168KB) for the epoch fields and declares eventEpoch on DaemonSession so the create/attach path drops its inline type cast. * fix(serve): stamp eventEpoch on accepted continuations and surface replayDegraded in the SDK (#7458) Address three review suggestions: - POST /session/:id/continue now returns eventEpoch alongside lastEventId, mirroring the prompt 202 envelope so continuation-seeded SSE cursors detect daemon restarts (DAEMON-001) - DaemonSessionClient exposes replayDegraded from the load response so SDK consumers can prefer the full transcript over a degraded snapshot - add /acp dispatch-level regression test for the degraded-snapshot stderr breadcrumb (fires only when snapshot.degraded is set) * test(cli): fix load-reply race in the degraded-breadcrumb transport test Await each session/load reply frame before opening the session stream so the GET cannot race conn.ownSession() into a 403; addresses the review Critical on the deg-0 arm. * fix(serve): allow and expose X-Qwen-Event-Epoch in CORS headers Cross-origin SSE clients must send the epoch header through preflight and read it from the response, or stale-cursor detection (DAEMON-001) is silently disabled for every CORS client. --------- Co-authored-by: qwen-code-bot <qwen-code-bot@users.noreply.github.com> Co-authored-by: qwen-code-dev-bot <qwen-code-dev-bot@users.noreply.github.com> Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com> Co-authored-by: Qwen Autofix <qwen-autofix[bot]@users.noreply.github.com> |
||
|
|
d9f7e1fbe1
|
feat(autofix): auto-rerun a check that died on infrastructure, once (#7562)
* feat(autofix): auto-rerun a check that died on infrastructure, once A failed check can be red because the machine died, not the code — a self-hosted runner losing the server, the disk filling. #7490's E2E failed with "runner lost communication with the server" and went green on a rerun. The scan now reruns such a check's failed jobs automatically. Detection is a conservative annotation whitelist (INFRA_FAILURE_SIGNATURES) — only unambiguous machine failures, never a test-level timeout, which could be a real regression. The one-shot guard is run_attempt, not a marker: a run already retried to attempt 2 and still infra-failing is persistent, so it is left for a human; after a rerun the attempt increments, so the next scan will not rerun it. Every step is fail-safe (any API error → no rerun), it runs only when the PR actually has a failed check, and the gate carries the same review-address carve-out as the other check selectors so the loop never reruns its own runs. This is the transient-infra sibling of #7554 (stale-base): that merges current main when a check is base-inherited; this reruns when a check died on the runner. Neither touches a check that is a genuine failure. Note: rerun-failed-jobs needs the PAT to hold `actions: write`. * fix(autofix): use POSIX ERE groups in infra-failure regex, cover all signatures in tests (#7562) * fix(autofix): also treat a git fetch/clone transport death as infra #6506's checkout died mid-transfer — "fetch-pack: invalid index-pack output" and "RPC failed; curl 92 ... CANCEL" — which then hung the job into the 20m limit. That is infra, not the PR (it only touches a doc), and a re-run made it green. But the infra-signature whitelist did not cover it, so the auto-rerun did not fire and it waited on a human. Add `invalid index-pack output` and `RPC failed` — the two canonical git-transport-death phrases — to INFRA_FAILURE_SIGNATURES. A co-present job-timeout line does not block the match (one matching line classifies the run), and a BARE timeout with no transport signature is still left alone, since it can be a real regression. Both new signatures are pinned in the test's per-signature loop, plus a case on #6506's real composite annotation and a bare-timeout-is-not-rerun guard. * fix(autofix): paginate annotations and filter Autofix runs in infra-rerun loop (#7562) --------- Co-authored-by: wenshao <wenshao@example.com> Co-authored-by: Qwen Code Bot <qwen-code-bot@users.noreply.github.com> Co-authored-by: qwen-code-ci-bot <qwen-code-ci-bot@users.noreply.github.com> |
||
|
|
8b13c86742
|
feat(cli): post the review body bilingually when the PR description is Chinese (#7564)
When the PR author writes Chinese, the posted /review body was
English-only. fetch-pr now records whether the PR description contains
Han characters (prDescriptionHasHan, detected from the same gh pr view
call and stamped into the plan report), and compose-review renders the
body bilingually off that flag: the English body leads, the complete
Chinese version rides collapsed in a <details><summary>中文说明</summary>
block, and the model footer stays outside the fold. The signal is the
CLI's own — the caller cannot toggle the register of a certified body —
and a local plan has no field, so nothing changes for terminal-only
reviews.
Every deterministic body fragment carries an en/zh pair end to end:
compose-review's clause templates and describeChunkGap phrases, the
coverage disclosures (reasons, publicLabel role subjects via a new
publicLabelZh, the path-free unread-brief reason) and the Step 4/5 gap
texts including the combined same-shape sentence. Fragments with no
deterministic translation — model-written findings, caller echoes,
interpolated errors — ride verbatim in both halves. verificationGaps now
returns structural {subject, reason, subjectZh, reasonZh} entries, which
also removes compose-review's last recover-the-boundary-from-prose parse.
SKILL.md instructs the same format for the model-authored inline
comments: English finding first (marker and suggestion block stay in the
English half — tooling filters on them), full Chinese translation
collapsed beneath, footer last.
Co-authored-by: verify <verify@local>
|
||
|
|
80784e645c
|
fix(autofix): make the review-address report wrapper lines bilingual (#7569)
The agent's address-summary.md / no-action.md already ends with a collapsed Chinese translation, but the workflow-appended wrapper lines around it — the "Addressed/Reviewed the latest feedback" lead-in, the "Base-conflict check" line, and the "Re-review when you have a moment" footer — were English-only and sat outside that block. So the posted comment was only half translated, unlike the takeover-ack comments (full collapsed Chinese block) and the "model/模型" sign-off in this same report (already inline-bilingual). Give each wrapper line an inline Chinese translation, matching the model/模型 idiom. The English halves are preserved verbatim — the streak-reset detector globs on "Addressed the latest review feedback" and "no changes needed", and a test extracts these lines — so behaviour is unchanged and old English-only comments still match. A new test pins each English-Chinese pair so a future reword that drops the Chinese fails. The terminal handoff/failure comment is left English-only for now (SKILL.md keeps it so by design); that is a separate change. Co-authored-by: wenshao <wenshao@example.com> |
||
|
|
01ee3bf921
|
fix(feishu): await stream cancels in media download teardown (#7465)
* fix(feishu): await stream cancels in media download teardown downloadMedia left two reject paths' stream teardown unawaited: - the oversize-stream path called reader.cancel() without awaiting, so a cancel error during teardown became an unhandled rejection (fatal under Node's default --unhandled-rejections=throw); - the Content-Length reject path returned without cancelling resp.body, leaving the connection pinned until GC. Both were already fixed for the sibling DingTalk downloader in #7361 (which was itself modelled on this Feishu code), so this brings Feishu to parity. Adds a regression test that pins the reader.cancel() await via a rejecting cancel, plus an assertion that the Content-Length path releases the body. * test(feishu): cover a rejecting body.cancel() on the Content-Length path Mirrors the existing reader.cancel() teardown test for the other reject path, per review feedback. Removing the await on resp.body?.cancel() flips execution onto the 'rejected: size ... exceeds' branch and the test fails. |
||
|
|
b436855a40
|
feat(core): propagate trusted daemon invocation context (#7279)
* feat(core): propagate trusted daemon invocation context * test(cli): update ACP startup expectation * refactor(core): centralize ACP capability env key * test(cli): update worktree ACP core mock * test(integration): run daemon context smoke on PRs * test(ci): update no-AK smoke expectation * test(core): cover invocation context isolation * fix(cli): compare ACP capability safely * fix(docs): restore GitHub action input names * fix(core): sanitize private ACP capability from child env * fix(core): reuse private ACP capability env constant * test(cli): cover malformed trusted invocation context * test(acp-bridge): assert exact child environment --------- Co-authored-by: qwen-code-dev-bot <qwen-code-dev-bot@users.noreply.github.com> Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com> Co-authored-by: 易良 <1204183885@qq.com> |
||
|
|
22cd9c7d7a
|
fix(web-shell): sync background agent status (#7561)
* fix(web-shell): sync background agent status * fix(web-shell): harden background agent reconciliation --------- Co-authored-by: ytahdn <ytahdn@gmail.com> |
||
|
|
5176ed0fa3
|
fix(sdk-python): require canonical form in validate_session_id (#7532)
uuid.UUID() accepts several non-canonical spellings — braced
{...}, urn:uuid:..., and dash-less hex — so validate_session_id let them
through after the RFC 4122 variant check. The value is then forwarded to
the CLI verbatim as --session-id/--resume, producing a malformed session
id downstream rather than a clear error at the SDK boundary.
Reject anything whose canonical form differs from the input. Case is
deliberately not part of the comparison: UUID() lowercases, and an
all-uppercase spelling is still valid canonical input.
Co-authored-by: Shaojin Wen <shaojin.wensj@alibaba-inc.com>
|
||
|
|
0a16b89bb4
|
feat(serve): persist workspace channel configuration (#7514)
* feat(serve): persist workspace channel configuration * fix(serve): harden channel settings snapshots * fix(serve): validate startup channel names * fix(serve): reserve all channel name --------- Co-authored-by: Shaojin Wen <shaojin.wensj@alibaba-inc.com> |
||
|
|
228fdd830c
|
fix(web-shell): include managed id in artifact open requests (#7570)
Co-authored-by: 钉萁 <dingqi.jww@alibaba-inc.com> |
||
|
|
3cc16e2723
|
ci: matrix ECS runner update + sudo install + repository_dispatch trigger (#7513)
* ci: matrix ECS runner update with sudo install - Use matrix strategy (ecs-update-sg, ecs-update-64c) to update both physical ECS hosts in parallel (fail-fast: false). - Always use sudo npm install -g so the package lands in /usr/local (system-wide PATH) instead of the runner user's home directory. - Move concurrency to job level (matrix context not available at workflow level per actionlint). - Add repository_dispatch trigger for release-driven updates. - Register new runner labels in actionlint.yaml. * fix(ci): use dispatch version for runner update --------- Co-authored-by: Shaojin Wen <shaojin.wensj@alibaba-inc.com> |
||
|
|
dc74279103
|
feat(serve): add workspace-level generation (#7552)
* feat(serve): add workspace-level generation * docs(serve): document workspace generation capability * fix(serve): align workspace generation contracts --------- Co-authored-by: ytahdn <ytahdn@gmail.com> |
||
|
|
5df71c8fd1
|
fix(autofix): retry an agent timeout instead of advancing past its feedback (#7563)
A timeout evaluated NOTHING — the agent ran out of budget before finishing, so nothing was committed and the feedback is unaddressed. It was treated as an evaluated verdict (real ts, watermark advances), which strands that feedback: the next scan sees "nothing new" and never retries. Observed on #7471 (round 13/100), a heavily-reviewed 1871-line PR: rounds 11 and 13 timed out, but round 12 pushed — so a timeout is transient far more often than not, and advancing past it left the round-13 feedback unhandled. run-agent.mjs now drops an `agent-timeout` signal on result.timedOut, and the handoff routes it like a pre-verdict crash: sentinel ts (feedback stays live) and a retry, with a headline that names the real fix at the cap (split the PR or raise the budget). A PR that PERSISTENTLY times out is bounded by the round cap and the consecutive-failure cap, so this cannot loop forever — it just stops treating a one-off budget blip as a verdict. The loop guard stays terminal (a tool-call loop is a real defect, not a budget blip). An API error still routes to its own model-key handoff; the timeout signal is written only when NOT an API error. Co-authored-by: wenshao <wenshao@example.com> |
||
|
|
135723a1f7
|
fix(cli): keep role codenames and brief paths out of the posted review body (#7560)
The posted body still carried two operator registers #7550 left in place: roster role subjects rendered their internal codenames ("Agent 1c: Cross-file tracer", "Test coverage matrix (whole-diff)"), and an unread brief's disclosure interpolated its filesystem path. And when verify and the reverse audit failed the same way, the body said it twice, in two near-identical sentences. - Every Brief now carries a publicLabel — the dimension said as what it checks ("the cross-file consistency pass") — and coverage's structural disclosures carry it as publicSubject beside the internal subject, plus a path-free publicReason for unread briefs. The internal label and the path stay on stderr, where they are the selector an operator acts on; every dedup and certification check still keys on the internal subject. - compose-review renders the public fields and groups by the reason the body PRINTS, so two unread briefs share one path-free sentence instead of repeating it per role. - verificationGaps merges verify and reverse-audit failures of the same delivery shape into one sentence with both subjects and both consequences; mixed shapes keep their precise per-role texts, and the per-role rebuild commands stay on stderr either way. Co-authored-by: verify <verify@local> |
||
|
|
51cf26e30a
|
fix(autofix): retry a skipped-Prepare instead of stranding the PR terminal (#7490)
* fix(autofix): retry a skipped-Prepare instead of stranding the PR terminal A base/infra failure BEFORE the agent runs was misread as an agent crash and terminated the PR forever. When an early step fails — installing or building the trusted base, checkout, node setup — the `Prepare branch and feedback` step is skipped, so NEWEST is empty, and the report step's "crashed before reading feedback" branch fired: MARK_ROUND=MAX_ROUNDS, terminal, scan skips it on every future tick. Observed: a web-shell TypeScript break on `main` failed `Install dependencies and build` (which builds the trusted base) across a whole scan batch, and SIX healthy PRs were stranded terminal at round=100 in one run — including ones at round 9 and 11 that had nothing to do with the break. `round=100` there is a terminal sentinel, not 100 attempts. NEWEST-empty now splits on steps.prepare.outcome: - 'skipped' (an earlier step failed, the agent never ran) is infra/base and transient: retry with a sentinel ts so the feedback stays live, incrementing the round so a PERSISTENTLY broken base is still bounded and stops at the cap (recoverable with /retry). - 'success'/'failure' (Prepare ran, no feedback produced) is a genuine pre-read agent crash: unchanged terminal behaviour. This is the reverse of the asymmetry #7482 addresses: that bounds a crash AFTER reading that retried forever; this stops a transient failure BEFORE reading from going terminal after one. * docs(autofix): note a pre-Prepare cancel also retries intentionally (#7490) * fix(autofix): also retry a cancelled/empty prepare outcome, not just skipped A previous review comment on this PR noted that a job cancelled before Prepare should retry too. It was right about the intent but the code did not do it: `steps.prepare.outcome` is 'cancelled' for a cancel and '' for a job that stopped before Prepare entered the step context — both DISTINCT from 'skipped', so `== 'skipped'` sent them to the terminal branch, the same over-termination this PR exists to fix. Match on "not a real Prepare run" (`!= 'success' && != 'failure'`) instead, so skipped, cancelled, and empty all retry; only a Prepare that actually ran to a verdict (success/failure) with no feedback stays terminal — the genuine pre-read agent crash. Test extended to drive the cancelled and empty cases (retry) and both real-run outcomes (terminal); mutation-verified that reverting to `== 'skipped'` reddens the cancelled case. * test(autofix): update the pre-read-crash case for the broadened retry The prior commit broadened NEWEST-empty retry to skipped/cancelled/empty but left the older 'replays the handoff decision' test asserting the old terminal behaviour for an unset PREPARE_OUTCOME (which now retries). That test's terminal cases now set PREPARE_OUTCOME=success/failure explicitly — the only outcomes that still terminate — so it exercises the genuine pre-read agent crash rather than the infra/cancel path. * test(autofix): anchor the skipped-Prepare extraction past the CONSEC block CI reddened `retries a skipped-Prepare` after main's consecutive-failure cap (#7482) merged into this branch: that block was inserted between this decision block and the report `{`, and it calls `gh api`. The test's `{`-anchored regex over-captured through it, so the extracted script ran the unstubbed `gh api` and failed. Anchor the end on the same `# Consecutive-failure` comment the sibling gate-crash test already uses, so the extraction stops at this decision block's own closing `fi`. * fix(autofix): exempt skipped-Prepare from the consecutive-failure breaker A broken base build skips Prepare, producing no API error file — so the consecutive-failure breaker ran on the new retry path and, after 5 scans, re-introduced the exact mass-stranding this PR exists to prevent. Exempt pre-agent infra failures (skipped/cancelled/empty outcome) from the breaker, mirroring the transient 429/5xx exemption: same failure class (not the PR's fault, self-heals, hits the whole batch). The round cap + sentinel-ts /retry recovery already bounds a persistently broken base. Also trim "checkout" from the retry headlines (checkout failures do not land in this branch) and hoist the duplicated MARK_TS assignment. * fix(autofix): reset the consecutive-failure streak on prior infra-failure markers The streak walker counted prior infra-failure headlines ("AutoFix could not start —…") as failures, inflating the consecutive-failure count on subsequent rounds. A PR with 3 real agent failures, then 3 rounds of base-build infra failures, then 1 more real failure would trip the cap-5 breaker even though only 4 rounds were the PR's fault. Add the two infra-failure headline patterns as reset strings in the streak walker, alongside the existing push and no-op resets. The genuine agent-crash headline ("AutoFix could not start evaluation —…") is deliberately excluded — it is a real failure and must still count. * fix(autofix): clarify infra-failure headlines and else-branch comment (#7490) Address review nits: the retry headline now mentions cancelled runs, the cap headline says 'reached the round cap' instead of overstating 'could not start for N rounds', the else-branch comment says 'prepare itself crashed' instead of 'agent crash', and the streak-reset pattern is simplified now that both infra headlines share the same prefix. --------- Co-authored-by: wenshao <wenshao@example.com> Co-authored-by: qwen-code-dev-bot <qwen-code-dev-bot@users.noreply.github.com> Co-authored-by: qwen-code-ci-bot <qwen-code-ci-bot@users.noreply.github.com> |
||
|
|
e07ebdcd87
|
fix(cli): say review coverage gaps in the author's units, not chunk ids (#7550)
The posted review body rendered coverage disclosures with the run's own bookkeeping as subjects: bare chunk ids, unsorted, one per subject. On a run that certified nothing (PR #7268) the body enumerated all 49 chunk ids across two sentences while opening with "Reviewed. Suggestions are inline." — the opener certified the exact thing every following sentence took back, and nothing on the PR page maps a chunk id to code. Three changes, all render-time — the structural entries, the caps, the caller-echo dedup and the stderr remediation still key on chunk ids, which is where the id is the selector a reader can act on: - Coverage now returns the plan's chunk→files table (DiffChunk.files was already in the plan JSON; the coverage type slice dropped it). - compose-review renders chunk gaps through describeChunkGap: every planned chunk collapses to "the entire diff", a narrow gap with known files names the files, and anything wider is counted against the plan's total. Applied to the receipt sentence, the uncoverable sentence (bare CLI entries only — caller-authored entries render verbatim) and the grouped per-cause sentences. - The COMMENT opener may no longer say "Reviewed." over a disclosure set that denies it: when no chunk is both covered and undisclosed — or no chunk universe could be read at all — it opens with a zero-certified warning instead. A rewritten launch demonstrably read its chunk, so coverage alone is not the test; certified is covered with no disclosure against it. Co-authored-by: verify <verify@local> |
||
|
|
95fc7ca152
|
feat(web-shell): add renderChatHeader slot for custom session header (#7553) | ||
|
|
fee601ae77
|
feat(web-shell): add selective shadow DOM isolation (#7551)
Co-authored-by: 钉萁 <dingqi.jww@alibaba-inc.com> |
||
|
|
14f1f2bb36
|
fix(ci): don't let one failing scenario sink the whole visual preview (#7511)
The web-shell visuals render runs every screenshot and flow in a single `test:e2e:visuals`, and that step had no `continue-on-error`, while the compose and upload steps had no `if: always()`. So one failing or timing-out scenario failed the job, the artifact was never uploaded, and the publish workflow had nothing to post — the entire preview vanished even when every other scenario passed and its PNG was already on disk. A flow (a long multi-click sequence) is the most fragile scenario kind, so the fragile one silently takes down the deterministic screenshots. PR #7498 hit exactly this: 29 scenarios passed, one new channel-management flow timed out, and the PR got no preview and no comment at all. Make the after-capture step `continue-on-error` so the passing captures survive and the later steps still compose and upload them. The publish job only runs on a `success` conclusion, so the job must stay green — but a masked failure must not read as a clean preview. Ship the step's real `.outcome` (which continue-on-error does NOT mask, unlike `.conclusion`) to the publisher as `render-status.txt`, and have the comment builder use it: an empty preview whose render failed says "one or more scenarios failed to render" and is explicitly NOT the reassuring green check or the coverage-gap prompt (both imply the render ran); a partial preview is labelled partial above the shots that did render. A missing status file (older run) defaults to complete, so this only ever adds a warning, never suppresses a real preview. The failing scenario still needs fixing — it's now surfaced in the comment rather than by silently deleting everyone else's preview. Co-authored-by: wenshao <wenshao@example.com> |
||
|
|
83b97ec79e
|
fix(cli): open the actual serve fallback port (#7501)
* fix(cli): open actual serve fallback port * test(cli): match serve URL to fallback listener * docs(cli): clarify serve listen error handling --------- Co-authored-by: Shaojin Wen <shaojin.wensj@alibaba-inc.com> |
||
|
|
b3e1662416
|
fix(vscode): use file picker image paths for vision input (#7493)
* fix(vscode): use image paths from file picker * fix(vscode): keep image picker paths raw * fix(vscode): resolve image picker paths on submit * fix(vscode): send picked images as vision context * fix(vscode): encode prompt image file URIs * fix(vscode): address image path review comments * test(vscode): cover image file reference edge cases |
||
|
|
7c73768fa5
|
perf(startup): lazy-load Google GenAI SDK on first use (#7512)
* perf(startup): lazy-load Google GenAI SDK on first use Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com> * codex: address PR review feedback (#7512) Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com> * codex: address PR review feedback (#7512) Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com> --------- Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com> |
||
|
|
32c491fc6e
|
test(core): stub the registry methods agent.ts actually calls (#7538)
The shared stubRegistry in agent.test.ts was missing six methods that
agent.ts reaches: bridgeApprovalEvents, getQueuedCount,
registerResidentAgent, restartCompletedAgent, unregisterResidentAgent and
waitForMessages.
That is not a benign omission. The background body wraps its work in a
try/catch that routes any throw into registry.fail(), so a missing method
never surfaces as 'not a function' — it silently converts a successful
run into a failed one. On the GOAL completion path
unregisterResidentAgent is called immediately before complete(), so the
TypeError replaced the completion entirely:
registry.fail('fork-...', 'registry2.unregisterResidentAgent is not a
function', ...)
That is what broke 'runs a non-interactive fork through the background
registry' on main. #7460 added the registry.complete assertion, which
exposed the incomplete stub — before it, nothing checked whether the
background body finished successfully and the TypeError was swallowed.
Stub all six with their real return shapes (unregisterResidentAgent
returns boolean, bridgeApprovalEvents returns the unsubscribe callback
agent.ts later invokes, waitForMessages resolves to a list) and assert
registry.fail was not called before asserting completion, so a future
gap reports the actual error instead of 'complete: 0 calls'.
|
||
|
|
d064bd7dcf
|
feat(cli): preserve semantic text when copying VP selections (#7286)
Some checks are pending
E2E Tests / E2E Test (Linux) - sandbox:docker (push) Waiting to run
E2E Tests / E2E Test (Linux) - sandbox:none (push) Waiting to run
E2E Tests / E2E Test - macOS (push) Waiting to run
E2E Tests / cron-interactive E2E (nightly) (push) Waiting to run
E2E Tests / web-shell Browser Regression (push) Waiting to run
* docs(cli): define semantic copy fidelity scope * docs(cli): address semantic frame review gaps * docs(cli): preserve soft-wrap source separators * feat(cli): preserve semantic selection copy * fix(cli): address semantic copy review findings * fix(cli): preserve clipped semantic boundaries * fix(cli): limit separator carrier joiner to visible width in wrap metadata The greedy /\s+/ match in wrapTextWithMetadata could capture more source whitespace than the separator carrier row actually consumed (e.g. a tab following a space), causing duplicated whitespace in semantic copy. Limit the match to visibleLine.length characters and add a mixed space/tab regression test. --------- Co-authored-by: 秦奇 <gary.gq@alibaba-inc.com> |
||
|
|
ce40e33a65
|
fix(ci): autofix route checks existing labels on non-trigger label events (#7481)
* fix(ci): autofix route checks existing labels on non-trigger label events When triage adds multiple labels in sequence, per-issue concurrency cancels earlier runs. If the last label is not a trigger label (e.g. scope/build-system), the surviving run skips the issue phase even though the issue already has autofix/approved + status/ready-for-agent. Before ignoring a non-trigger label event, check ISSUE_LABELS_JSON for both required labels. If present and the issue is open, proceed with the issue phase. Trust was already established when the trigger labels were applied (both require triage+ permission). * fix(ci): require trusted sender for label fallback |