qwen-code

mirror of https://github.com/QwenLM/qwen-code.git synced 2026-04-29 20:20:57 +00:00

Author	SHA1	Message	Date
jinye.djy	7b628d7003	fix(core): shallow-clone savedApiKeySource to avoid mutation risk Copy the ConfigSource object before applyResolvedModelDefaults runs, so a future refactor that mutates source objects in place won't break the save/restore logic. Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>	2026-04-26 00:44:41 +08:00
jinye.djy	0e3cb0a753	fix(core): use baseUrl source as hasBeenApplied signal for provider change detection Replace `apiKeyEnvKey !== undefined` guard with `baseUrl source === 'modelProviders'` to reliably detect whether applyResolvedModelDefaults has been called before. This fixes two edge cases: 1. No-envKey models: hot-reload changing baseUrl was undetected because apiKeyEnvKey remained undefined. Now baseUrl source is checked. 2. Startup with envKey but omitted baseUrl: undefined !== default URL could falsely trigger isProviderChanged. Now skipped at startup since baseUrl source is not yet 'modelProviders'. Updates hot-reload test fixtures to simulate post-apply state (baseUrl source as 'modelProviders') and adds no-envKey hot-reload test. Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>	2026-04-25 23:50:09 +08:00
jinye.djy	d3563960ca	fix(core): detect provider config hot-reload in isUnchanged check When a model provider config is hot-reloaded (e.g. via Coding Plan update) changing envKey or baseUrl while keeping the same model id, the save/restore logic must not preserve the old apiKey. Extend the isUnchanged guard to compare apiKeyEnvKey and baseUrl against the resolved model, but only after applyResolvedModelDefaults has run at least once (apiKeyEnvKey !== undefined). On first startup call these fields are still unset, so the check is skipped to preserve the settings/cli-sourced key correctly. Adds two hot-reload tests (envKey change and baseUrl change). Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>	2026-04-25 17:06:48 +08:00
jinye.djy	5467075a6e	fix(core): replace non-null assertion with truthiness guard and add cold-start test - Replace `savedApiKeySource!` with a truthiness guard for safer source restoration - Add test for cold-start scenario (previousAuthType undefined) to verify no key preservation occurs on first syncAfterAuthRefresh - Fix stale "short-circuit" comment in programmatic key test Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>	2026-04-25 00:18:44 +08:00
jinye.djy	9f470f50d5	fix(core): move apiKey preservation from applyResolvedModelDefaults to syncAfterAuthRefresh The previous fallback logic inside applyResolvedModelDefaults could leak a settings/cli-sourced apiKey to a different provider when switching models within the same authType (e.g. dashscope → openai). This is a credential safety issue because the two providers may have different baseUrls. Move the save/restore logic to syncAfterAuthRefresh Step 1, guarded by an `isUnchanged` check (same authType AND same modelId). This ensures: - Restart scenario: apiKey preserved (same model, no change) - Cross-provider switch: apiKey cleared (different modelId) Also adds two cross-provider switch tests (settings-sourced and CLI-sourced) per review feedback. Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>	2026-04-25 00:07:41 +08:00
jinye.djy	b58f171f56	fix(core): also preserve CLI-sourced apiKey during syncAfterAuthRefresh Address review feedback: keys passed via CLI flags (e.g. --openaiApiKey) were dropped on restart because source kind 'cli' was not in the fallback allowlist. Add 'cli' to the condition and a regression test. Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>	2026-04-24 11:30:45 +08:00
jinye.djy	3a1087f7a0	fix(core): preserve settings-sourced apiKey when registry model envKey is absent (#3417 ) On restart, `applyResolvedModelDefaults` unconditionally cleared the apiKey resolved from `settings.security.auth.apiKey` (layer 4 fallback) and only read from `process.env[model.envKey]`. When the provider-specific env var was absent (e.g. key stored only in settings), the correctly resolved key was discarded, causing a 401 error. Now capture the previously-resolved apiKey before clearing and fall back to it when `process.env[model.envKey]` is empty, but only for safe source kinds (`settings` and general `env` without `via.modelProviders`). Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>	2026-04-21 14:43:15 +08:00
Shaojin Wen	afbb5e71db	fix(cli): rework session recap rendering and add blur threshold setting (#3482 ) * feat(cli): make recap away-threshold configurable The 5-minute blur threshold was hard-coded. Confirmed from Claude Code's own binary (v2.1.113) that 5 minutes is their default as well (and that they shift to 60 minutes when 1h prompt-cache is active) — so the default stays, but expose it as `general.sessionRecapAway ThresholdMinutes` for users who briefly alt-tab often and don't want recaps piling up, or who want to lower it for testing. Non-positive / unset values fall back to the 5-minute default, so dropping the key has the same behavior as before. * fix(core): align recap prompt with Claude Code (1-2 sentences, ≤40 words) The earlier "exactly one sentence, 80-char cap" was an over-correction to a single in-the-moment ask. Going back to it: the natural shape of "current task + next action" is two clauses, and forcing them into a single sentence either crams them with a semicolon or drops the next action entirely on complex sessions. Adopt Claude Code's prompt verbatim (extracted from the v2.1.113 binary): "under 40 words, 1-2 plain sentences, no markdown. Lead with the overall goal and current task, then the one next action. Skip root-cause narrative, fix internals, secondary to-dos, and em-dash tangents." Add a Chinese-budget note (~80 chars) and keep the <recap>...</recap> wrapping that protects against reasoning-model preambles leaking into the UI. The sticky banner already re-measures controls height when the recap toggles, so a 2-line render lays out cleanly. Sweep "one-line" out of user-facing copy (settings description, slash-command description, feature docs, design doc) so the documentation matches the new shape. * fix(cli): restore "one-line" in user-facing recap copy Verified from the Claude Code v2.1.113 binary that the slash-command description IS literally "Generate a one-line session recap now" even though the underlying prompt allows 1-2 sentences. Claude Code is deliberately setting a tighter user expectation than the prompt guarantees, which keeps the surface feel "glanceable". Mirror that asymmetry: keep the prompt at 1-2 sentences (the previous commit) for behavioral parity, but put "one-line" back in the user- visible copy (slash-command description, settings description, user docs). Internal design doc keeps the accurate "1-2 sentence" wording. * fix(cli): render recap inline in history to match Claude Code Earlier I read the user's complaint that the recap "scrolled away" as "the recap should be sticky above the input box," and built a sticky banner accordingly. Disassembly of the Claude Code v2.1.113 binary shows the actual behavior is the opposite: their away_summary is a plain `type:"system", subtype:"away_summary"` message dispatched through the standard message renderer (no Static, no anchor, no flexbox pinning) — it scrolls with the conversation like every other system message. Tear out the sticky-banner machinery so recap matches that: - Recap is back in the `HistoryItemWithoutId` union and `addItem`'d into history (both from `/recap` and from auto-trigger), so it serializes into session saves and behaves like every other history item — no special clear paths, no resume-wrapper, no layout-effect re-measure dance. - `useAwaySummary` takes `addItem` again instead of a setter callback. - `AwayRecapMessage` renders the way Claude Code does: a 2-column gutter with `※`, then bold "recap: " and italic content, all in dim color. Drop the prior `StatusMessage`-shaped layout that fused prefix and label into "※ recap:". - Remove the AppContainer plumbing, the slashCommandProcessor state, the UIStateContext fields, the DefaultAppLayout / ScreenReader placement blocks, the test-utils mocks, and the noninteractive stub. Restore `useResumeCommand.handleResume` to a void return since callers no longer need the success boolean. Sweep the design doc so the architecture diagram, files table, and hook deps reflect the inline-history flow. * fix(cli): dedupe back-to-back auto-recaps with no new user turns between Two consecutive blur cycles, each over the threshold but with no new user activity in between, would each fire their own auto-recap and add two near-duplicate entries to history (same task, slightly different wording from temperature-driven LLM variance). Reported case: leaving the terminal twice while a /review of one PR was still on screen produced two recaps both about that same review. Add a `shouldFireRecap` gate before kicking off the LLM call: - Need at least 3 user messages in history total (don't fire on a near-empty session). - If a previous away_recap is already in history, need at least 2 new user messages since that one before another can fire. Same shape as Claude Code's `Ic1` gate (`Sc1=3`, `Rc1=2`). Read history through a ref so this isn't in the effect's deps and the effect doesn't re-run on every message. * fix(cli): type useResumeCommand.handleResume as Promise<void> Per gemini review on #3482: the interface declared this as `() => void` but the implementation is `async` and returns `Promise<void>`. The mismatch silently lost the chainable promise — tests had to launder it through `as unknown as Promise<void> \| undefined` just to await. Tighten the interface to `Promise<void>` and drop the cast in the "closes the dialog immediately" test. * fix(cli): persist auto-fired recap to chat recording so /resume keeps it Per yiliang114 review on #3482: the manual `/recap` path persists across `/resume` because the slash-command processor records every output history item via `chatRecorder.recordSlashCommand({ phase: 'result', outputHistoryItems })`, but the auto path called `addItem` directly and bypassed that recorder. The result was an asymmetry where users who triggered recap manually saw it after `/resume`, while users whose recap fired automatically lost it. Mirror the manual recording from useAwaySummary's `.then` callback — record only the `result` phase (not invocation, since we don't want a fake `> /recap` user line replayed) with the away-recap item as the single output. Wrapped in try/catch because recap is best-effort and must never surface a failure to the user. Add useAwaySummary.test.ts covering: - the recording path is taken on a successful auto-trigger - the dedup gate (`shouldFireRecap`) suppresses the LLM call entirely, including the recording, when no new user turns happened since the last recap * fix(cli): cast recap item via spread to satisfy strict tsc --build CI's `tsc --build` (stricter than local `tsc --noEmit`) rejected the direct `item as Record<string, unknown>` cast: HistoryItemAwayRecap's literal `type: 'away_recap'` field doesn't overlap with `unknown`, TS2352. Use the `{ ...item } as Record<string, unknown>` spread pattern that the rest of the codebase (arenaCommand, slashCommandProcessor's serializer) already uses for the same SlashCommandRecordPayload field.	2026-04-21 14:39:13 +08:00
tanzhenxin	b27cb81bb7	feat(cli): attribute /stats rows to the originating subagent (#3229 ) * feat(cli): attribute /stats rows to the originating subagent Thread subagent identity through telemetry via an AsyncLocalStorage context so each API response knows which subagent (or main) emitted it. Aggregate a per-source breakdown alongside the existing per-model totals and render one row per (model, source) in /stats and /stats model. Main-only sessions collapse to the existing single-row display. Resolves #3215 * fix(cli): reserve `main` subagent name and stabilize /stats React keys Two latent correctness issues found during self-review of PR #3229: - A subagent named `main` would silently collide with the `MAIN_SOURCE` sentinel and be merged into the main bucket with no attribution. Add `main` to the reserved-names list so validation rejects it. - `flattenModelsBySource` used the normalized display label (with `-001` stripped) as the React key, which could collapse distinct models `foo` and `foo-001` into duplicate keys. Split `ModelSourceEntry` into `{ key, label, metrics }` with `key` built from the raw model name (plus `::source` in the split case), and update both `StatsDisplay` and `ModelStatsDisplay` to key rows/columns off it. Also surface invalid-subagent-file parse errors through the debug logger instead of swallowing them entirely, so users running with debug logging enabled can tell why a subagent failed to load. Add a dedicated unit test file for `flattenModelsBySource` covering the collapse rule, session-wide split, source order, the `foo`/`foo-001` key-collision regression, and the empty-bySource fallback. Extend the reserved-name test to include `main`.	2026-04-21 11:44:10 +08:00
Shaojin Wen	52c7a3d0ed	fix(cli): pin /recap above input and align defaults with fastModel (#3478 ) Some checks are pending Qwen Code CI / Lint (push) Waiting to run Details Qwen Code CI / Test (push) Blocked by required conditions Details Qwen Code CI / Test-1 (push) Blocked by required conditions Details Qwen Code CI / Test-2 (push) Blocked by required conditions Details Qwen Code CI / Test-3 (push) Blocked by required conditions Details Qwen Code CI / Test-4 (push) Blocked by required conditions Details Qwen Code CI / Test-5 (push) Blocked by required conditions Details Qwen Code CI / Test-6 (push) Blocked by required conditions Details Qwen Code CI / Test-7 (push) Blocked by required conditions Details Qwen Code CI / Test-8 (push) Blocked by required conditions Details Qwen Code CI / Post Coverage Comment (push) Blocked by required conditions Details Qwen Code CI / CodeQL (push) Waiting to run Details E2E Tests / E2E Test (Linux) - sandbox:docker (push) Waiting to run Details E2E Tests / E2E Test (Linux) - sandbox:none (push) Waiting to run Details E2E Tests / E2E Test - macOS (push) Waiting to run Details * fix(cli): pin /recap above input box and align defaults with fastModel The recap rendered as a regular history item, so as soon as the model streamed a new reply the "where you left off" reminder scrolled out of view. Move it to a sticky banner anchored just above the Composer (matching how btwItem is rendered) so it stays visible across turns. While reworking the surface, also: - Replace the chevron prefix with `※ recap:` so it reads as a labeled recap line instead of a generic dim message. - Mirror the placement in ScreenReaderAppLayout so screen-reader users see it in the same logical position. - Drop HistoryItemAwayRecap from the HistoryItemWithoutId union — it is no longer addItem-able, and leaving it in invited silent no-op bugs where addItem(awayRecap) would compile but render nothing. - Clear the banner on /clear, /reset, /new and on /resume into a different session, so a recap from a previous context doesn't bleed into a freshly started one. - Re-measure the controls box when the banner appears or disappears (its height changes by a couple of lines) so the main content area recomputes availableTerminalHeight and stays laid out correctly. Auto-trigger now defaults to "on iff fastModel is configured" rather than unconditionally on. Running an ambient background recap on the main coding model is too costly and slow to be a sane default; tying it to fastModel means the feature is silently opt-in for users who have set up a cheap fast model. An explicit `general.showSessionRecap` override still wins either way, and `/recap` itself is unaffected. Sharpen the slash-command description to match the new behavior. * fix(core): silence AbortSignal listener-leak warning in OpenAI pipeline Every chat.completions.create call wires up an abort listener on the incoming AbortSignal, and several layers — retryWithBackoff, the LoggingContentGenerator wrapper, the SDK's own internal stream/fetch plumbing — register their own listeners against the same signal. Five retry attempts plus those layers comfortably exceed Node's default 10-listener cap and produce a MaxListenersExceededWarning. With features that share or compose signals (e.g., recap + followup speculation firing on the same response cycle), even a higher cap gets blown past. The signals here are per-request and short-lived, so the accumulation is structural rather than a real memory leak — they get GC'd as soon as the request settles. setMaxListeners(0, signal) at the SDK boundary disables the warning for these specific signals only, without masking any genuine leak elsewhere in the process. Idempotent and confined to the one place where retry-bound API calls cross into the SDK. * fix(core): tighten recap to a single sentence within 80 chars The 1-3 sentence budget reliably wrapped onto two lines in the sticky banner above the input box, which made it visually heavy for what is supposed to be a glanceable reminder. Constrain the prompt to exactly one sentence with a hard 80-char cap, and merge the "high-level task + next step" rule into a single sentence instead of two adjacent ones. Also sweep the docs (settings, commands, design) so the user-facing copy and the internal design notes match the new format. * fix(cli): apply review feedback for recap PR Two issues from review: - The schema description for `general.showSessionRecap` still said "1-3 sentence summary" while the prompt, docs, and slash-command copy already say "one-line". Aligns the text in settingsSchema.ts and the regenerated VSCode JSON schema. - The /resume wrapper cleared the sticky recap synchronously, before the inner handler had a chance to discover that no session data was available. On a no-op resume the user would still lose the current recap. Make `useResumeCommand.handleResume` return Promise<boolean> reporting whether a session actually loaded, and only clear the recap on a confirmed switch. * fix(cli): default showSessionRecap to false and drop fastModel heuristic The earlier "enabled iff fastModel is configured" default made it hard for users to answer the simple question "is auto-recap on for me right now?" — the answer depended on a setting from a different category, and setting/unsetting fastModel silently changed recap behavior. Revert to a plain boolean with a conservative off-by-default: - Auto-trigger fires only when the user explicitly sets `general.showSessionRecap: true`. - Manual `/recap` keeps working regardless (that's a user-initiated call, not an ambient one). - Users never get ambient LLM calls billed to their main coding model without having opted in. Aligns settings.md, design doc, and the regenerated JSON schema.	2026-04-20 23:58:19 +08:00
jinye	bf561fa495	fix(core): prevent malformed permission rules from becoming tool-wide catch-alls (#3467 ) * fix(core): prevent malformed permission rules from becoming tool-wide catch-alls A permission rule with unbalanced parentheses (e.g. `Bash(rm -rf /)`) was silently parsed with `specifier: undefined`, causing `matchesRule` to treat it as a catch-all that matches every invocation of the tool. For deny rules this blocked all commands; for allow rules a typo could silently auto-approve everything. Add an `invalid` flag to `PermissionRule`. `parseRule` now marks rules with unbalanced parens as invalid, `matchesRule` short-circuits them to never match, and all entry points (`addSessionRule`, `addPersistentRule`, `parseRules`) warn on malformed input. `listRules` filters out invalid rules so they don't appear in the /permissions UI. * fix(cli): show error in /permissions dialog when adding malformed rule When a user enters a rule with unbalanced parentheses via the "Add Rule" input in the /permissions dialog, show an inline error message instead of silently accepting and then hiding the invalid rule. Closes #3459 Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>	2026-04-20 18:56:14 +08:00
Shaojin Wen	c74d7678cb	Revert "feat(core): add dynamic swarm worker tool (#3433 )" (#3468 ) This reverts commit `f7ebc372f1`.	2026-04-20 16:40:14 +08:00
Shaojin Wen	5fedf10419	feat(cli): add tool execution progress messages (#3155 ) * feat(cli): add tool execution progress messages with per-tool elapsed time, shell stats, and terminal progress bar - Show per-tool elapsed time (Ns) next to spinner after 3 seconds of execution, covering all tools (not just shell), by piping existing core startTime through to the UI layer via IndividualToolCallDisplay.executionStartTime - Add shell output statistics bar below ANSI output showing +N lines overflow count, byte size, and explicit timeout when set by user - Add terminal tab progress bar via OSC 9;4 sequences for iTerm2, Ghostty, and ConEmu, with tmux/screen DCS passthrough support - Extend AnsiOutputDisplay with optional totalLines/totalBytes/timeoutMs fields - Add ShellStatsBar component for rendering shell output statistics * fix(cli): address review feedback — use formatDuration for timeout, pass displayHeight to ShellStatsBar - Use existing formatDuration() from formatters.ts instead of inline timeout formatting for correct precision (e.g., "2m 3s" not "2m") - Add displayHeight prop to ShellStatsBar so +N lines overflow calculation respects actual terminal height, not hardcoded DEFAULT_HEIGHT * fix(cli): guard terminal progress bar against non-TTY stdout Check process.stdout.isTTY in isProgressBarSupported() so escape sequences are not emitted when stdout is piped, redirected to log files, or running in CI environments where TERM_PROGRAM may be set but stdout is not a TTY. Also add defensive isProgressBarSupported() guard in the effect cleanup. * fix(cli): format tool elapsed time with minutes/hours for long-running tools Previously showed raw seconds (e.g. "3600s") for long-running tools. Now formats as "3s" for under a minute, "1m 30s" for minutes, and "2h 15m" for hours, while keeping compact integer seconds for short durations. * fix(cli): audit fixes for terminal progress and shell output stats Three issues found by post-merge audit: - useTerminalProgress: WT_SESSION was wrongly used to exclude Windows Terminal. WT 1.6+ actually supports OSC 9;4 progress sequences (per Microsoft docs), so treat it as a positive indicator like iTerm2 and Ghostty. - useTerminalProgress: add process.on('exit'\|'SIGINT'\|'SIGTERM') handler that writes PROGRESS_CLEAR. Without it, killing the CLI mid-tool (Ctrl+C, SIGTERM) left the terminal tab stuck showing an indeterminate progress indicator because React cleanup never ran. Mirrors the useBracketedPaste cleanup pattern. - shell.ts: ANSI totalBytes used token.text.length (character count), inconsistent with the string path's Buffer.byteLength(..., 'utf-8'). Multi-byte chars (CJK, emoji) now count as their true UTF-8 byte length in both paths. * refactor(cli): right-align tool elapsed time, extract to its own component Move the executing-tool elapsed-seconds indicator out of ToolStatusIndicator (where it sat immediately after the spinner on the left edge) and into a new right-aligned ToolElapsedTime component. The left placement caused layout jitter: every second the elapsed text width would change (e.g. "9s" → "10s" → "1m" → "1m 15s"), shifting the tool name and description horizontally. Right-aligning the elapsed keeps the tool name anchored and only the far-right timer moves. - New packages/cli/src/ui/components/shared/ToolElapsedTime.tsx owns the setInterval + formatElapsed logic. - ToolStatusIndicator is now pure status again; the executionStartTime prop is gone from it. - ToolMessage and CompactToolGroupDisplay mount ToolElapsedTime as the last flex child of the status row, with marginLeft=1. - ToolInfo gains flexGrow=1 so the description fills the middle and the timer sits flush at the right edge of the row. * fix(core): measure tool elapsed from executing-transition, not validating-entry trackedCall.startTime is stamped when a tool is first registered with the scheduler (validating state), then preserved through awaiting_approval, scheduled, and executing transitions. Using it for the executing-row elapsed display meant any approval-wait time was counted as execution time — a tool that waited 30s for user approval would flash "30s" immediately when it actually began running. Add a separate executionStartTime on ExecutingToolCall, stamped at the moment of the transition into 'executing', and pipe that through useReactToolScheduler into IndividualToolCallDisplay.executionStartTime. startTime is kept as-is for durationMs bookkeeping. Also stops piping executionStartTime for validating/scheduled states, since those don't have a meaningful execution duration yet. * fix(cli): only hook 'exit' for terminal progress cleanup, not SIGINT/SIGTERM Registering SIGINT/SIGTERM handlers that neither re-raise nor exit inhibits Node's default termination behavior. If this hook were ever the only signal handler in play, Ctrl+C would leave the process hanging. Drop the signal handlers and rely on 'exit' alone. Other parts of the CLI already own the signal-to-shutdown path (gemini.tsx, telemetry shutdown, sharedTokenManager, etc.) and ultimately call process.exit(), which fires 'exit' and runs this cleanup. SIGKILL cannot be cleaned up either way. * fix(cli): thread executionStartTime through agent-view tool groups The main TUI renders per-tool elapsed time via IndividualToolCallDisplay. executionStartTime, but the agent-view adapter (agentHistoryAdapter.ts) constructed its display items without this field, so sub-agent tool groups never showed the elapsed indicator. Thread it through the sub-agent event pipeline: - AgentToolOutputUpdateEvent gains an optional executionStartTime, emitted once per callId by agent-core.onToolCallsUpdate the first time a call is seen in the scheduler's 'executing' state (carrying ExecutingToolCall.executionStartTime). This also fires for tools that produce no live output, so their elapsed indicator appears too. - AgentInteractive tracks executionStartTimes in a callId→timestamp map, analogous to liveOutputs/shellPids. First TOOL_OUTPUT_UPDATE with a value wins; later events that re-carry it are ignored. Cleared on TOOL_RESULT. - AgentChatView passes the map as the new fifth argument to agentMessagesToHistoryItems. - The adapter reads the map for Executing tools and sets IndividualToolCallDisplay.executionStartTime, matching the main-view plumbing. Agent-view tool_groups now render the same elapsed-time indicator the main view does. Adds three test cases covering set-when-executing, skip-when-completed, and skip-when-map-absent. * fix(core): skip stats accounting for string shell chunks totalLines/totalBytes are only emitted alongside AnsiOutputDisplay in the ANSI-array branch of updateOutput. Computing split('\n') and Buffer.byteLength for string chunks was wasted work — the values never left the function. Only compute stats when event.chunk is an AnsiLine[] now.	2026-04-20 16:04:58 +08:00
易良	33d0b4af00	fix(core): normalize Windows PATH for MCP stdio servers (#3451 ) Some checks are pending Qwen Code CI / Lint (push) Waiting to run Details Qwen Code CI / Test (push) Blocked by required conditions Details Qwen Code CI / Test-1 (push) Blocked by required conditions Details Qwen Code CI / Test-2 (push) Blocked by required conditions Details Qwen Code CI / Test-3 (push) Blocked by required conditions Details Qwen Code CI / Test-4 (push) Blocked by required conditions Details Qwen Code CI / Test-5 (push) Blocked by required conditions Details Qwen Code CI / Test-6 (push) Blocked by required conditions Details Qwen Code CI / Test-7 (push) Blocked by required conditions Details Qwen Code CI / Test-8 (push) Blocked by required conditions Details Qwen Code CI / Post Coverage Comment (push) Blocked by required conditions Details Qwen Code CI / CodeQL (push) Waiting to run Details E2E Tests / E2E Test (Linux) - sandbox:docker (push) Waiting to run Details E2E Tests / E2E Test (Linux) - sandbox:none (push) Waiting to run Details E2E Tests / E2E Test - macOS (push) Waiting to run Details * fix(core): normalize Windows PATH for MCP stdio servers * test(core): stabilize MCP stdio env assertion on Windows * refactor(core): extract shared windowsPath utility and fix PATH override semantics Address PR #3451 review comments: 1. [Critical] Fix PATH override semantics: normalize process.env before merging with server config so that a server-provided PATH fully replaces the parent value instead of being merged with a stale case-variant (Path). 2. [Suggestion] Extract mergeWindowsPathValues and normalizePathEnvForWindows into a shared utility module (src/utils/windowsPath.ts), eliminating near-identical copies in shellExecutionService.ts and mcp-client.ts. The shared version includes the fingerprint caching from shellExecutionService. 3. [Suggestion] Revert the "should connect via command" test to its original non-Windows form, restoring generic coverage. Add a dedicated test for server config PATH override behavior. * fix(core): mock process.platform in shellExecutionService PATH tests The shared windowsPath utility uses process.platform (not os.platform), so the two Windows PATH normalization tests need to mock both to ensure the utility detects the win32 platform correctly in test environments. * fix(core): use objectContaining for env assertion in command transport test On Windows, normalizePathEnvForWindows deduplicates PATH entries, so the env passed to StdioClientTransport won't exactly match a raw process.env spread. Use expect.objectContaining to verify server config env vars are passed through without requiring an exact match on all platform-dependent env keys.	2026-04-20 15:22:01 +08:00
Shaojin Wen	9d8201d206	feat(core): PDF text extraction fallback and Jupyter notebook parsing (#3160 ) * feat(core): add PDF text extraction fallback and Jupyter notebook parsing For text-only models (qwen3-coder, deepseek) that lack PDF modality support, read_file now falls back to pdftotext (poppler-utils) for text extraction instead of returning an unsupported error. A new `pages` parameter enables paginated PDF reading (e.g. "1-5", "3-"). Also adds structured .ipynb parsing — notebooks are displayed as labeled cells with code blocks and execution outputs rather than raw JSON. Key changes: - New utils/pdf.ts: pdftotext integration with availability caching, page range parsing, 5MB maxBuffer, and 100K char output truncation - New utils/notebook.ts: .ipynb JSON parser with per-cell output truncation (10K chars) and overall notebook truncation (100K chars) - Modified fileUtils.ts: new 'notebook' FileType, PDF fallback logic, pages parameter threading - Modified read-file.ts: pages parameter in schema/validation/execution * fix(core): avoid circular dependency via shell-utils in pdf.ts pdf.ts was importing execCommand from shell-utils.ts, which transitively pulled in tool-utils.ts → ../index.js (barrel), creating a circular dependency that caused AuthType to be undefined during vitest module initialization in 46 test files. Replace with a local execFile wrapper that has no transitive dependencies beyond node:child_process. * fix(core): use optional call on getContentGeneratorConfig Moving the modalities computation outside the if-block caused readManyFiles.test.ts to fail because its mock config doesn't implement getContentGeneratorConfig — previously the method was only called for media files (image/pdf/audio/video), never for text files. Use ?.() to gracefully fall back to an empty modalities object when the method is not defined. * fix(core): reject open-ended PDF page ranges to enforce 20-page limit Previously, parsePDFPageRange returned lastPage: Infinity for open-ended ranges like "3-", which bypassed the 20-page validation check and caused pdftotext to extract from the start page to EOF. This violated the documented "Max 20 pages per request" contract. Now validation explicitly rejects open-ended ranges with a helpful message telling users to specify an explicit end page within the limit. The pages parameter schema description and interface comment are also updated to reflect this constraint. * fix(core): tighten parsePDFPageRange to reject malformed tokens parseInt() silently truncates invalid input, so values like "1-2-3", "5abc", "1-2x", "1x-2", and "1.5" were accepted and then interpreted as the wrong range (e.g. "1-2-3" parsed as 1-2). Switch to regex-based whole-string validation so any non-matching input returns null at ReadFileTool.build() time instead of reaching pdftotext. * fix(core): surface processSingleFileContent errors in readManyFiles readManyFiles previously dropped any file whose processSingleFileContent result carried an error, so users only saw "No files matching the criteria were found or all were skipped." This hid actionable guidance such as the pdftotext-not-installed install hint, password-protected PDF notices, and the >10MB size-limit message. Now the per-file error message (already a human-readable string in llmContent) is included as a content part, so batch reads surface the same guidance as single-file reads. * fix(core): tolerate whitespace around hyphen in parsePDFPageRange The strict regex introduced in the previous commit stopped accepting inputs like "1 - 5" or "3 -", which the old parseInt-based parser handled (parseInt skips leading whitespace). Allow optional \s* on each side of the hyphen while still rejecting malformed trailing tokens such as "5abc" and "1-2-3". * fix(cli,core): render failed @file reads as Error in atCommandProcessor The previous commit surfaced per-file errors through readManyFiles, but FileReadInfo still lacked a status field and atCommandProcessor hardcoded ToolCallStatus.Success for every entry in result.files. So a failed read (missing pdftotext, password-protected PDF, >10MB file) rendered in the UI as if it had succeeded, just with the error text embedded in the LLM content. Add an optional `error` field on FileReadInfo, populate it in readFileContent, and use it in atCommandProcessor to pick ToolCallStatus.Error plus a resultDisplay string the user can see. * fix(core): treat pdftotext maxBuffer overrun as truncation When a text-dense PDF produced more than 5MB of stdout, Node killed the child and `execFile` delivered the error as `ERR_CHILD_PROCESS_STDIO_MAXBUFFER`, which fell into the generic `pdftotext failed:` branch — so a perfectly valid PDF failed instead of returning the usual truncated output. Detect the maxBuffer error code in the execFile wrapper, and in extractPDFText use the partial stdout with the existing truncation note. Also lower the maxBuffer to 2×MAX_PDF_TEXT_OUTPUT_CHARS (from 5MB) since anything past that is discarded anyway — this also caps RSS for pathological inputs. * fix(core): skip 10MB size gate for PDF text-extraction path The generic 9.9MB file-size check ran before the pdf branch knew whether we were taking the base64 inline path or the pdftotext text-extraction path. That meant `read_file("huge.pdf", pages="1-5")` was rejected up front even though pdftotext streams through the file and only emits a capped (100K char) text slice — never loading 15MB into Node memory. Move the size gate past the fileType/modalities decision point and skip it when the PDF will go through text extraction (pages parameter set, or model lacks pdf modality). The base64 inline path still carries its own encoded-size cap, so oversized PDFs continue to be rejected there. * fix(core): harden pdftotext wrapper against six audit findings An adversarial pass over the PDF utilities turned up several issues that warrant hardening before the PR lands: - Argument injection (C1): filenames starting with `-` (e.g. `-opw=foo.pdf`) are parsed as options by poppler's argv parser when passed positionally. Insert `--` before `filePath` in both `extractPDFText` and `getPDFPageCount` so the shell's option parser stops processing flags. Reproduced locally: `pdftotext -h -` prints help while `pdftotext -- -h -` treats `-h` as the input file. - Brittle availability signal (H1): `isPdftotextAvailable` used `stderr.length > 0` as the positive signal, so a sandbox that suppresses stderr would cache `false` for the whole process. Switch to the exit code. - Concurrent availability probes (H2): N parallel callers (e.g. an `@`-glob of PDFs) each spawned their own `pdftotext -v` before the first probe resolved. Cache the in-flight promise. - Precision-loss bypass of the 20-page cap (H3): `Number()` collapses any integer past 2^53 onto the same value, so the string `"999999999999999998-999999999999999999"` parsed as a 1-page range and slid past the validator. Cap accepted page numbers at 1,000,000. - Timeout error clarity (M2): 30s timeouts surfaced as the generic `pdftotext failed:` branch with empty stderr. Detect SIGTERM/killed and emit a dedicated "timed out after 30s" message. - Over-eager maxBuffer success (M1): the previous commit treated any maxBuffer overrun with non-empty stdout as a truncated success. If the overrun was driven by stderr spam (password warnings, corrupt- PDF diagnostics), that delivered garbage as success. Require at least MAX_PDF_TEXT_OUTPUT_CHARS of stdout before treating as truncated; otherwise re-run the password/corrupt detectors on the captured stderr. Added regression tests for each. * fix(core): gate non-regular files and oversized PDFs before extraction Two defense-in-depth guards suggested by the adversarial audit: - Non-regular files (FIFOs, sockets, /dev/zero, character devices) have meaningless `stats.size` (typically 0), so the 10MB size gate would happily wave them through. Handing `/dev/zero` to pdftotext then produced a 30s-timeout failure after the wrapper streamed megabytes into Node. Require `stats.isFile()` before routing into any extraction path. - The previous commit skipped the 10MB gate for the PDF text- extraction path so `read_file("huge.pdf", pages="1-5")` could work. Unbounded, though, a multi-GB PDF would make pdftotext run until the 30s timeout fires. Add a separate 100MB ceiling for the extraction path with a guidance error pointing the user at `pages` or document splitting. The base64 inline path keeps its own encoded- size cap. Added regression tests for both. * fix(core): strip ANSI escapes and surface non-text outputs in notebooks Two notebook-rendering issues surfaced by the audit: - ipykernel emits ANSI CSI/SGR escape sequences (`\x1B[0;31m...`) in error tracebacks by default. Those codes add noise and burn tokens without conveying anything useful once we're rendering to plain text. Strip them from stream, execute_result, display_data, and error outputs. - Cells whose only output was a non-text MIME type (image/png, text/html, application/vnd.jupyter.widget-view+json, ...) were silently dropped — the model saw the source code with no indication that a plot or HTML block existed. Emit a `[non-text output: <mime-types>]` placeholder so the model knows something was there without us inlining the payload. * fix(core): round-2 audit fixes (in-flight cleanup, Windows timeout, ANSI/MIME) Reverse audit on the previous three commits surfaced four medium- severity issues plus a polish item: - isPdftotextAvailable in-flight promise leak: the `.then(...)` cleared the cached promise on success but a synchronous throw inside the IIFE would have left a rejected promise stuck in the slot forever. Switch to `.finally` so the slot is always cleared. - Timeout detection on Windows: Node's `execFile` `timeout` terminates via TerminateProcess on Windows, where `signal` is typically `null` rather than `'SIGTERM'`. The previous SIGTERM-only check would let Windows timeouts fall through to the generic "pdftotext failed" branch. Accept null/undefined signal alongside SIGTERM. - ANSI regex was CSI-only: missed OSC hyperlinks (`ESC ]8;;url`), DCS, APC/PM/SOS, and lone two-byte escapes that ipykernel and related tools sometimes emit. Extend the pattern to cover all four families. - Non-text MIME placeholder was attacker-controlled: a malicious notebook could set `data: {"\nIGNORE PREVIOUS INSTRUCTIONS\n": ...}` and that key would flow unescaped into `[non-text output: ...]`, smuggling prompt-injection payload bytes into the LLM context. Filter keys against the IANA MIME-type grammar before joining. - Hoisted PDF_EXTRACTION_MAX_MB to module scope alongside the other size constants so it's discoverable in one place. * chore(core): correct ANSI comment example and rename cache-reset test Comment/test polish from the convergence audit: - The `[@-Z\-_]` C1-Fe branch of the ANSI regex does not actually match `ESC c` (RIS), `ESC 7`, or `ESC 8`, which sit at 0x63/0x37/0x38. It does match IND/NEL/HTS/RI (ESC D/E/H/M). Correct the jsdoc example. - The `should clear the in-flight promise after a probe to allow retries` test wasn't distinguishing the `.finally` behaviour from the `resetPdftotextCache()` call that immediately precedes the second probe. Rename it to reflect what it actually verifies; the `.finally` remains as defence-in-depth (a synchronous throw inside the IIFE's own handlers can't leave the in-flight slot stuck on a rejected promise).	2026-04-20 11:09:50 +08:00
ihubanov	0b8b3da836	feat(cli): add slashCommands.disabled setting to gate slash commands (#3445 ) * feat(cli): add slashCommands.disabled setting to gate slash commands Introduces a first-class way for operators to hide and refuse to execute specific slash commands. Useful for multi-tenant / enterprise / sandboxed deployments where different users should see different command subsets. The denylist is sourced from three unioned inputs: * `slashCommands.disabled` settings key (string[], UNION merge), so workspace scopes can only add to a denylist set at user or system scope, never shrink it — matching the shape already used by `permissions.deny`. * `--disabled-slash-commands` CLI flag (comma-separated or repeated). * `QWEN_DISABLED_SLASH_COMMANDS` environment variable. Matching is case-insensitive against the final (post-rename) command name, so extension commands are addressable by their disambiguated form (e.g. `firebase.deploy`). Disabled commands are removed from `CommandService`'s output, so they disappear from autocomplete and produce the standard unknown-command path in both interactive TUI and non-interactive (`--prompt`) modes. The scope of this change is slash commands only: it does not affect tool permissions (still `permissions.deny`) or keyboard shortcuts. * chore(cli): regenerate settings.schema.json for slashCommands.disabled Regenerates the companion JSON schema consumed by the VS Code extension after adding the `slashCommands.disabled` entry to the TS schema in the previous commit. Required by the "Check settings schema is up-to-date" CI lint step. * fix(cli): route disabled slash commands to unsupported, not no_command handleSlashCommand was passing the disabled denylist straight into CommandService.create, so disabled commands disappeared from `allCommands` too. The fallback existence check that distinguishes "known but not allowed in non-interactive mode" from "truly unknown" then failed, and disabled commands like `/help` fell through to `no_command` — causing the caller to forward them to the model as plain prompt text. Keep `allCommands` unfiltered and apply the denylist only when constructing the executable set and when producing the unsupported response. A disabled command now returns `unsupported` with a "disabled by the current configuration" reason and never reaches the model. Added three regression tests covering the primary case, case-insensitive match, and the preserved no_command path for genuinely unknown input.	2026-04-20 11:06:26 +08:00
易良	7cded6e0df	feat(vscode-ide-companion): support /insight command (#2593 ) * feat(vscode-ide-companion): support /insight command Add ACP support for /insight progress streaming and report opening in the VSCode companion. Resolves #2023 * fix(cli): defer insight command runtime deps * test(cli): cover acp slash command allowlist * Revert "test(cli): cover acp slash command allowlist" This reverts commit `3209274ab6`. * Revert "fix(cli): defer insight command runtime deps" This reverts commit `3b08491e46`. * Reapply "fix(cli): defer insight command runtime deps" This reverts commit `386c5c67d3`. * Reapply "test(cli): cover acp slash command allowlist" This reverts commit `e2716140dd`. * refactor(cli): simplify insight ACP integration - Replace `formatAcpInsightProgress` with `encodeAcpInsightProgress` using JSON payload - Move imports to top-level, no longer defer loading for non-ACP mode - Remove `INSIGHT_READY_MARKER` parsing from Session.ts as it's now handled by WebViewProvider * refactor: extract insight protocol markers to core package Move INSIGHT_PROGRESS_MARKER and INSIGHT_READY_MARKER from cli and vscode-ide-companion packages to @qwen-code/qwen-code-core for better shareability and to avoid duplication. Also extract ACP_ALLOWED_COMMANDS constant in Session.ts to improve readability and maintainability. * refactor(vscode-ide-companion): extract test helper to reduce webview mock duplication Introduce `setupAttachedProvider()` helper in WebViewProvider.test.ts to eliminate ~160 lines of repeated webview mock + provider setup code across 5 insight-related test cases. * feat(cli): 添加ACP执行模式到内置命令当ACP启用时，将executionMode参数传递给所有内置命令，使命令能够识别当前运行在ACP模式下并相应地调整行为。 test(cli): 为insight命令添加ACP进度消息流测试新增测试验证insight命令在ACP模式下能够正确流式传输进度消息，而不必等待生成完成。测试涵盖了从开始到完成的整个进度更新过程。 refactor(core): 重构insight协议消息格式将insight进度和就绪消息从基于标记字符串的格式改为结构化的JSON格式，提供更好的类型安全和解析可靠性。 feat(vscode-ide-companion): 支持新的insight消息协议更新WebViewProvider以支持新的结构化insight消息协议，能够正确解析和处理来自CLI的进度和就绪消息。 ``` * fix(vscode-ide-companion/insight): streamline insight progress handling Trim redundant CLI insight coverage around the ACP path. Keep the VS Code insight progress flow aligned with normalized slash commands and the updated progress layout. * fix(insight): restore slash commands after webview reload Cache available commands in the VS Code provider so webview restoration still exposes /insight without a manual login. Also remove the unused progress bar markup to keep the UI diff smaller. * Update packages/webui/src/index.ts Co-authored-by: Shaojin Wen <shaojin.wensj@alibaba-inc.com> * fix(webui): remove duplicate insight card export --------- Co-authored-by: Shaojin Wen <shaojin.wensj@alibaba-inc.com>	2026-04-20 10:02:18 +08:00
易良	41f71ab7e7	feat(cli): add bare startup mode (#3448 ) * feat(cli): add bare startup mode Skip implicit startup discovery in bare mode while keeping explicit inputs such as include directories and extension overrides. Add a repository plan document and targeted tests for config, startup, skills, extensions, and memory discovery. * fix(bare): enforce explicit-only startup behavior * fix(cli): preserve bare tools in non-interactive mode * chore(docs): remove bare mode planning note	2026-04-20 10:01:59 +08:00
Shaojin Wen	60a6dfc14c	feat(cli): add session recap with /recap and auto-show on return (#3434 ) Some checks are pending Qwen Code CI / Lint (push) Waiting to run Details Qwen Code CI / Test (push) Blocked by required conditions Details Qwen Code CI / Test-1 (push) Blocked by required conditions Details Qwen Code CI / Test-2 (push) Blocked by required conditions Details Qwen Code CI / Test-3 (push) Blocked by required conditions Details Qwen Code CI / Test-4 (push) Blocked by required conditions Details Qwen Code CI / Test-5 (push) Blocked by required conditions Details Qwen Code CI / Test-6 (push) Blocked by required conditions Details Qwen Code CI / Test-7 (push) Blocked by required conditions Details Qwen Code CI / Test-8 (push) Blocked by required conditions Details Qwen Code CI / Post Coverage Comment (push) Blocked by required conditions Details Qwen Code CI / CodeQL (push) Waiting to run Details E2E Tests / E2E Test (Linux) - sandbox:docker (push) Waiting to run Details E2E Tests / E2E Test (Linux) - sandbox:none (push) Waiting to run Details E2E Tests / E2E Test - macOS (push) Waiting to run Details * feat(cli): add session recap with /recap and auto-show on return Users often open an old session days later and need to scroll through pages to remember where they left off. This change adds a short "where did I leave off" recap — a 1-3 sentence summary generated by the fast model — so they can resume without re-reading the history. Two triggers: - /recap: manual slash command. - Auto: when the terminal has been blurred for 5+ minutes and gets focused again (uses the existing DECSET 1004 focus protocol via useFocus). Gated on streamingState === Idle so it never interrupts an active turn. Only fires once per blur cycle. The recap is rendered in dim color with a chevron prefix, visually distinct from assistant replies. A new `general.showSessionRecap` setting controls the auto-trigger (default on). /recap works independent of the setting. Implementation notes: - generateSessionRecap uses fastModel (falls back to main model), tools: [], maxOutputTokens: 300, and a tight system prompt. It strips tool calls / responses from history before sending — tool responses can hold 10K+ tokens of file content that drown the recap in irrelevant detail. The 30-message window respects turn boundaries (slice never starts on a dangling model/tool response). - Output is wrapped in <recap>...</recap> tags; the extractor returns empty (skips render) if the tag is missing, preventing model reasoning from leaking into the UI. - All failures are silent (return null) and logged via a scoped debugLogger; recap is best-effort and must never break main flow. - /recap refuses to run while a turn is pending. * fix(cli): abort in-flight recap when showSessionRecap is disabled If the user disables showSessionRecap while an auto-recap LLM call is already in flight, the previous code returned early without aborting. The pending .then would still pass its idle/abort guards and append the recap, producing an unwanted message after the user has opted out. Abort the controller and clear it eagerly so the resolved promise no longer adds to history. * fix(cli): gate /recap and auto-recap on streaming idle state Two related issues from review: 1. /recap was only refusing when ui.pendingItem was set, but a normal model reply runs with streamingState === Responding and a null pendingItem. Invoking /recap mid-stream would generate a recap from a partial conversation and insert it between the user prompt and the assistant reply. 2. useAwaySummary cleared blurredAtRef before checking isIdle, so if focus returned during a still-streaming turn (after a >5min blur) the recap was permanently dropped — there was no later retry when the turn became idle, because isIdle was not in the effect deps. Fixes: - Expose isIdleRef on CommandContext.ui (mirrors btwAbortControllerRef pattern). Plumb it from AppContainer through useSlashCommandProcessor. - recapCommand now refuses when isIdleRef.current is false OR pendingItem is non-null. - useAwaySummary preserves blurredAtRef on the !isIdle bail and adds isIdle to the effect deps, so the trigger re-evaluates when the current turn finishes. - Brief blurs (< AWAY_THRESHOLD_MS) still reset blurredAtRef. Also seeds isIdleRef in nonInteractiveUi and mockCommandContext so the new field has a sensible default outside the interactive UI. * docs: document /recap command, showSessionRecap setting, and design - User docs: add /recap to the Session and Project Management table in features/commands.md and a dedicated subsection covering manual use, the auto-trigger, the dim-color rendering, and the fast-model tip. - User docs: add general.showSessionRecap row to the configuration settings reference. - Design doc: docs/design/session-recap/session-recap-design.md covers motivation, the two trigger paths, the per-file architecture, prompt design with the <recap> tag and three-tier extractor, history filtering rationale (functionResponse can be 10K+ tokens), the useAwaySummary state machine, the isIdleRef gating for /recap, model selection, observability, and out-of-scope items. * fix(core): exclude thought parts from session recap context filterToDialog kept any non-empty text part, but @google/genai's Part type also marks model reasoning with part.thought / part.thoughtSignature. That hidden chain-of-thought was being fed to the recap LLM and could get summarized as if it were user-visible dialogue. Drop parts where either flag is set. Update the design doc's History 过滤 section to call this out alongside the existing tool-call/response rationale. * docs(session-recap): correct debug-logging guidance, fill in state machine, sharpen UX wording Audit of the session recap docs against the implementation found three issues worth fixing: - Design doc claimed debug logs were enabled via a QWEN_CODE_DEBUG_LOGGING env var. That var does not exist; debug logs are written to ~/.qwen/debug/<sessionId>.txt by default, gated by QWEN_DEBUG_LOG_FILE. Replace with the accurate path + opt-out behavior, and tell the reader to grep for the [SESSION_RECAP] tag. - Design doc's useAwaySummary state machine table was missing the isFocused && blurredAtRef === null path (taken on first render and right after a brief-blur reset). Add the row. - User doc's "Refuses to run ... failures are silent" line conflated the inline-error refusal with silent generation failures, and "(when the conversation is idle)" used internal jargon. Split the two cases and spell out what "idle" means, including the wait-then-fire behavior when focus returns mid-turn. * docs(session-recap): correctly describe /recap vs auto-trigger failure modes The previous wording said "Generation/network failures are silent — the recap simply does not appear", but recapCommand returns a user-facing info message ("Not enough conversation context for a recap yet.") in exactly that path, and also returns inline messages for the config-not-loaded and busy-turn guards. Only the auto-trigger path is truly silent (it just skips addItem when generateSessionRecap returns null). Split the two paths in the doc so the manual command's "always responds with something" behavior is distinguished from the auto-trigger's no-op-on-failure behavior. * docs(session-recap): align prompt-rules section with the actual prompt Two doc-vs-code mismatches in the design doc's "System Prompt" section, caught with the same lens as yiliang114's failure-mode review: - The bullet list claimed RECAP_SYSTEM_PROMPT forbids "推测用户意图" and "用 'you' 称呼用户". Those rules existed in an early draft but were dropped when the <recap> tag rules were added; the current prompt has no such restrictions. Replace with the actual rules and add a "与 RECAP_SYSTEM_PROMPT 一一对应" marker so future edits stay in sync. - The doc said systemInstruction "覆盖" the main agent prompt. True for the agent prompt portion, but GeminiClient.generateContent internally calls getCustomSystemPrompt which appends user memory (QWEN.md / 自动 memory) as a suffix. Spell that out — the final system prompt is recap prompt + user memory, which is actually useful project context for the recap. * docs(session-recap): translate design doc to English The repo convention for docs/design is English (7 of 8 existing files; auto-memory/memory-system.md is the only Chinese one). The first version of this design doc followed the auto-memory example, which turned out to be the wrong sample. Translate to English while preserving the existing structure, the state-machine table, the prompt-vs-doc 1:1 alignment, the QWEN_DEBUG_LOG_FILE description, and the failure-mode notes added in prior commits. * fix(cli): drop empty info return from /recap interactive success path The interactive success path inserts the away_recap history item directly via ui.addItem and then returned `{type: 'message', messageType: 'info', content: ''}`. The slash-command processor's 'message' case unconditionally calls addMessage, which adds another HistoryItemInfo with empty text. The empty info renders as nothing (StatusMessage early-returns null), but it still bloats the in-memory history list and shows up in /export and saved sessions. Return void on the interactive success path and on the abort path so the processor's `if (result)` check skips the message-handler branch entirely. Widen the action's return type to `void \| SlashCommandActionReturn` to match (same shape as btwCommand).	2026-04-19 21:38:48 +08:00
euxaristia	c175fd3d4a	feat(core): enhanced loop detection with stagnation + validation-retry checks (#3236 )	2026-04-19 18:06:43 +08:00
Reid	28d5722955	fix(core): remove abort listener during cleanup (#3438 )	2026-04-19 17:28:38 +08:00
gin-lsl	a02c115445	feat(tools): add Markdown for Agents support to WebFetch tool (#2734 ) Closes #2025	2026-04-19 17:23:09 +08:00
Reid	f7ebc372f1	feat(core): add dynamic swarm worker tool (#3433 ) * feat(core): add dynamic swarm worker tool Add a swarm tool for ad-hoc parallel worker execution with bounded concurrency, wait-all and first-success modes, per-worker failure isolation, and aggregated results. Register the tool in core, prevent nested worker recursion, and document the new workflow. * fix(core): harden swarm worker execution Prevent swarm calls from bypassing the outer scheduler concurrency budget. Disallow interactive question prompts in swarm workers by default, and avoid incomplete Markdown table escaping by using an HTML entity for pipe characters. Add focused tests for the scheduler behavior, worker tool restrictions, and result formatting.	2026-04-19 14:46:59 +08:00
Reid	cd8d9dce6a	fix(core): support older Git during repository initialization Replace git init --initial-branch with git init followed by symbolic-ref HEAD refs/heads/main. This keeps new repositories on main without requiring Git 2.28 or newer. Also ensure checkpoint shadow repository setup uses its dedicated git config during the initial commit.	2026-04-19 14:24:01 +08:00
Shaojin Wen	4bf5bf22de	feat(cli): support refreshInterval in statusLine for periodic refresh (#3383 ) * feat(cli): support refreshInterval in statusLine for periodic refresh The statusLine (#3311) re-runs only when Agent state changes (token count, model, git branch, etc.). Commands that display external data — a clock, rate-limit counters, CI build status — have no Agent event to hook into and go stale between messages. Add an optional `ui.statusLine.refreshInterval` field (seconds, minimum 1) that schedules a setInterval alongside the existing event-driven updates. Overlap with state-change debounce is safe: `doUpdate` kills any in-flight child and bumps the generation counter, so only the most recent output reaches the footer. Validation lives in `getStatusLineConfig`: - Must be `number`, `Number.isFinite(...)`, `>= 1` - Anything else is silently dropped (no interval scheduled) No changes to the default behavior — configs without `refreshInterval` behave exactly as before. * fix(cli): yield periodic statusLine tick when previous exec is in flight Review feedback on #3383: with `refreshInterval: 1` and a command whose real exec time exceeds 1s, each tick was unconditionally calling `doUpdate()` — which kills the in-flight child and bumps the generation counter — so the prior exec's callback was always discarded as stale. `setOutput` was never reached and the statusline stayed empty until `refreshInterval` was removed or the command became faster. Guard the interval callback with an `activeChildRef` check so a pending exec is allowed to finish. State-change triggers (model switch, token count, branch, etc.) still go through `scheduleUpdate` → `doUpdate` directly and legitimately preempt stale children; only the periodic tick yields. The existing 5s exec timeout is still the hard ceiling. Also drop the redundant `'refreshInterval' in raw` check — the `typeof raw.refreshInterval === 'number'` guard already excludes missing / undefined values. Tests: - Add regression test `'skips periodic ticks while a previous exec is still running'` — three ticks during one unfinished exec trigger zero new spawns; the next tick after callback completion does spawn. - Update two existing tests to resolve the mount exec before expecting subsequent ticks (the old tests implicitly relied on the starvation behavior being tolerated). * test(cli): assert user-visible lines state in starvation regression Self-review insight: the existing `skips periodic ticks while a previous exec is still running` test only counted `exec` calls — it confirmed the guard prevents redundant spawns, but would have silently passed even if the eventual callback was still being discarded as stale (which is the actual user-visible symptom of the starvation bug). Add `expect(result.current.lines).toEqual(['done'])` after resolving the mount's pending callback. Without the guard, generationRef would have bumped 3 times during the yielded ticks, the callback's captured gen would fail the stale check, `setOutput` would never fire, and `lines` would stay empty — now caught explicitly. * perf(cli): dedupe statusLine output to skip unchanged Footer re-renders Review feedback on #3383 (narrow terminal stacking): when `refreshInterval` fires at 1s and the command output is unchanged, the mount-and-setOutput cycle still allocates a new array and triggers a Footer re-render. Under certain narrow-terminal conditions, Ink's erase-line accounting mis-counts wrapped rows and stale content accumulates on screen. The Footer-layout root cause is in #3311's narrow-mode flex setup and Ink's truncate semantics, which is out of scope for this PR. But we can cut the re-render surface here by preserving the `lines` array reference when the command produces identical output — a strict Pareto improvement for any caller (clock-style statuslines with second-precision still re-render; rate-limit / branch / CI-status style statuslines that change infrequently stop triggering work every tick). Tests: - `preserves the same lines array reference when output is unchanged` asserts referential equality after a re-exec with identical stdout. - `produces a new reference when output changes` guards against over-eager dedup that would miss legitimate updates. * fix(cli): stabilize Footer rendering in narrow terminals Narrow-terminal E2E feedback on #3383: with `refreshInterval` at 1s, empty lines were accumulating above the input prompt each tick. Root cause is in the Footer flex layout — originally from #3311 — where Ink miscounts logical rows vs the physical rows the terminal actually uses. Two adjustments, both idiomatic (used elsewhere in the repo already): 1. Left column — `minWidth={0}`. Without this, Yoga's `min-width: auto` default keeps the Box at its natural content width, so a statusline wider than the terminal doesn't engage `<Text wrap="truncate">`; the text renders at content-width and the terminal wraps it physically. `minWidth={0}` lets the column shrink so the text child can truncate at container width. 2. Right section — `flexWrap="wrap"`. With multiple indicators (sandbox label, debug badge, dream, context-usage) the row can exceed a narrow terminal's width. Without `flexWrap` Ink lays them out in a single logical row, but the terminal physically wraps to two — Ink's erase sequence (`\e[2K\e[1A…` per logical row) then clears one row while two exist, and the extra row ghosts every re-render. With `wrap` Ink tracks the second row explicitly and erases correctly. Together these make the Footer's row count match between Ink's logical view and the terminal's physical view, so frequent re-renders (as `refreshInterval` enables) stop accumulating ghost rows. Needs verification in a real narrow TTY — from this environment I can reason about the flex semantics and confirm both props are supported by Ink's Box, but actually observing ghost-row elimination requires process.stdout.columns on a real terminal. * Revert "fix(cli): stabilize Footer rendering in narrow terminals" This reverts commit `9758cda85f`. Reason: I could not reproduce BZ-D's reported ghost-row stacking in tmux (40x25, 2-line statusline + real exec + Static history + refreshInterval: 1) over 14+ ticks. Both `minWidth={0}` and `flexWrap="wrap"` are legitimate defensive idioms, but without a failing repro I can't verify they address the reported bug, and I shouldn't ship a speculative layout change as "the fix". Keeping the output-dedup commit (`e1d321186`) — that one is a strict improvement regardless of the underlying Ink behavior. Will request BZ-D's specific terminal setup and reopen with a verified fix (or confirm the issue is specific to a particular emulator, not flex/Ink).	2026-04-19 11:12:16 +08:00
Edenman	4ee9ca912c	feat(mcp): add OSC 52 copy hotkey for OAuth authorization URL (#3337 ) (#3393 ) Some checks are pending Qwen Code CI / Lint (push) Waiting to run Details Qwen Code CI / Test (push) Blocked by required conditions Details Qwen Code CI / Test-1 (push) Blocked by required conditions Details Qwen Code CI / Test-2 (push) Blocked by required conditions Details Qwen Code CI / Test-3 (push) Blocked by required conditions Details Qwen Code CI / Test-4 (push) Blocked by required conditions Details Qwen Code CI / Test-5 (push) Blocked by required conditions Details Qwen Code CI / Test-6 (push) Blocked by required conditions Details Qwen Code CI / Test-7 (push) Blocked by required conditions Details Qwen Code CI / Test-8 (push) Blocked by required conditions Details Qwen Code CI / Post Coverage Comment (push) Blocked by required conditions Details Qwen Code CI / CodeQL (push) Waiting to run Details E2E Tests / E2E Test (Linux) - sandbox:docker (push) Waiting to run Details E2E Tests / E2E Test (Linux) - sandbox:none (push) Waiting to run Details E2E Tests / E2E Test - macOS (push) Waiting to run Details When MCP OAuth authentication falls back to the "copy this URL into your browser" path (e.g. remote/web terminal where the browser can't auto-open), long URLs wrap across lines inside the bordered dialog and the trailing │ border characters get selected alongside the URL, forcing the user to manually strip them out before pasting. Surface the URL on a dedicated event and let the user press 'c' to push it to the local clipboard via an OSC 52 escape sequence. Works through SSH and modern web terminals (iTerm2, Windows Terminal, xterm.js-based emulators, tmux with set-clipboard, etc.) without a subprocess, and falls back to a visible "copy the URL above manually" hint when the terminal is not a TTY or OSC 52 is blocked. Key points: - OAuth provider emits OAUTH_AUTH_URL_EVENT carrying the full URL. - AuthenticateStep listens, tracks it in state, and binds 'c' while authenticating (modifier/paste keys are filtered out). - copyToClipboardViaOsc52 writes to stderr when it's a TTY, falls back to stdout, and wraps the sequence for tmux/GNU screen via DCS passthrough so multiplexed sessions still work. - Honest feedback: distinct "copy request sent" / "cannot write to terminal" states with a short auto-revert so repeated presses reset the timer. Fixes #3337	2026-04-18 20:22:06 +08:00
Reid	7eba1c4635	test(core): update scheduler registry mock (#3415 ) Some checks are pending Qwen Code CI / Lint (push) Waiting to run Details Qwen Code CI / Test (push) Blocked by required conditions Details Qwen Code CI / Test-1 (push) Blocked by required conditions Details Qwen Code CI / Test-2 (push) Blocked by required conditions Details Qwen Code CI / Test-3 (push) Blocked by required conditions Details Qwen Code CI / Test-4 (push) Blocked by required conditions Details Qwen Code CI / Test-5 (push) Blocked by required conditions Details Qwen Code CI / Test-6 (push) Blocked by required conditions Details Qwen Code CI / Test-7 (push) Blocked by required conditions Details Qwen Code CI / Test-8 (push) Blocked by required conditions Details Qwen Code CI / Post Coverage Comment (push) Blocked by required conditions Details Qwen Code CI / CodeQL (push) Waiting to run Details E2E Tests / E2E Test (Linux) - sandbox:docker (push) Waiting to run Details E2E Tests / E2E Test (Linux) - sandbox:none (push) Waiting to run Details E2E Tests / E2E Test - macOS (push) Waiting to run Details Update the CoreToolScheduler retry-loop test registry mock to match the current ToolRegistry interface. Add ensureTool and getAllToolNames so the tests exercise the scheduler path used in production.	2026-04-18 13:46:46 +08:00
jinye	9f4734e84d	fix(tool-registry): add lazy factory registration with inflight concurrency dedup (#3297 ) Closes #3221. Introduces a lazy factory API on ToolRegistry (registerFactory, ensureTool, warmAll, getAllToolNames) as infrastructure for future esbuild code-splitting (#3226). With the current single-bundle build, the lazy API does not change startup time on its own — the primary immediate value is fixing three pre-existing bugs uncovered while designing it. Bug fixes: - Concurrent instantiation (P0): the original ensureTool had no concurrency protection around `await factory()` — two concurrent calls for the same tool both passed the cache check and each ran the factory, producing two instances. AgentTool and SkillTool register SubagentManager listeners in their constructors, so the extra instance leaked listeners. Fix: a per-name `inflight: Map<string, Promise<Tool>>` so concurrent ensureTool() calls share a single promise. On factory rejection the inflight entry is cleared so a subsequent call can retry. - stop() resource leak: stop() only disposed tools already in `this.tools`; tools still loading in `inflight` when stop() ran finished afterward and were never disposed. Fix: await Promise.allSettled(inflight.values()) before the dispose loop. - Cache hit left stale factory: ensureTool's cache-hit branch did not delete the factory entry, so warmAll() would re-invoke the factory for an already-loaded tool. Fix: delete the factory on cache hit. Additional hardening in response to review feedback: - warmAll({ strict?: boolean }): strict mode re-throws the first factory failure rather than swallowing it. Config.initialize() uses strict: true so a broken built-in tool fails startup fast instead of silently leaving a partially initialized registry; runtime-path callers (GeminiChat, agent runtime, etc.) continue to use the non-strict default and log failures via debugLogger. - getAllTools() and getFunctionDeclarationsFiltered() emit a debug warning when called while unloaded factories remain, nudging callers toward warmAll() without hard-breaking existing code paths. - copyDiscoveredToolsFrom() now iterates source.tools.values() directly instead of source.getAllTools() — the copy path deals only with already-discovered MCP/command tools and should not trigger the unloaded-factory warning. - MemoryTool and SkillTool config parsing was extracted into memory-config.ts and skill-utils.ts so a factory can resolve tool metadata without importing the tool module. Tests: - tool-registry.test.ts adds 128 lines covering: concurrent ensureTool runs the factory exactly once, warmAll and ensureTool overlap, retries succeed after a prior factory failure, stop() disposes tools that finish loading after stop was called, and warmAll strict vs default behavior. - 33 existing call sites across cli, core, agents, and subagents were updated to await warmAll() before bulk tool access.	2026-04-18 10:31:50 +08:00
euxaristia	5facd8738b	feat(core): detect tool validation retry loops and inject stop directive (#3178 ) Primary change: prevent the model from burning tokens in an infinite retry loop when a tool call repeatedly fails schema validation with the same error (observed with ask_user_question and a malformed `questions` parameter retrying 10+ times with the same validation error). - Track consecutive validation failures per (tool name, error message) pair in CoreToolScheduler via a `validationRetryCounts` Map. - After 3 consecutive failures for the same (tool, error) pair, append a RETRY LOOP DETECTED directive to the error response instructing the model to stop, re-examine the schema, try a fundamentally different approach, or surface the issue to the user. - Reset per-tool counters when the tool invocation succeeds; reset globally when an incoming batch shares no tool name with any previously failing tool; reset the per-tool counter when the tool returns a different validation error so unrelated mistakes do not accumulate toward the threshold. - Distinct from LoopDetectionService, which tracks model-behavior loops (repeated thoughts, stagnant actions); this change catches tool-API misuse loops at the scheduler layer. Piggyback fixes bundled in the same PR: - packages/cli/index.ts, packages/core/src/services/shellExecutionService.ts: treat PTY `EAGAIN` on the read path as an expected read error alongside `EIO`, avoiding noisy surface-level failures from transient non-blocking reads. - scripts/build.js: switch the settings-schema generation step from `npx tsx` to `node --import tsx/esm` for Bun compatibility. Tests: - Unit tests in coreToolScheduler.test.ts cover: directive injection on the 3rd consecutive failure, counter reset when a different tool is called, and counter reset after a successful invocation of the same tool (fail → fail → succeed → fail → fail must not trip the directive).	2026-04-18 10:24:46 +08:00
ChiGao	9e26424aa7	feat(cli): add dual-output sidecar mode for TUI (#3352 ) Some checks are pending Qwen Code CI / Lint (push) Waiting to run Details Qwen Code CI / Test (push) Blocked by required conditions Details Qwen Code CI / Test-1 (push) Blocked by required conditions Details Qwen Code CI / Test-2 (push) Blocked by required conditions Details Qwen Code CI / Test-3 (push) Blocked by required conditions Details Qwen Code CI / Test-4 (push) Blocked by required conditions Details Qwen Code CI / Test-5 (push) Blocked by required conditions Details Qwen Code CI / Test-6 (push) Blocked by required conditions Details Qwen Code CI / Test-7 (push) Blocked by required conditions Details Qwen Code CI / Test-8 (push) Blocked by required conditions Details Qwen Code CI / Post Coverage Comment (push) Blocked by required conditions Details Qwen Code CI / CodeQL (push) Waiting to run Details E2E Tests / E2E Test (Linux) - sandbox:docker (push) Waiting to run Details E2E Tests / E2E Test (Linux) - sandbox:none (push) Waiting to run Details E2E Tests / E2E Test - macOS (push) Waiting to run Details * feat(cli): add dual-output sidecar mode for TUI Adds an optional dual-output mode for the interactive TUI: while Qwen Code keeps rendering normally on stdout, it concurrently emits a structured JSON event stream on a second channel (--json-fd / --json-file) and optionally watches a JSONL command file (--input-file) for prompts and tool-permission responses written by an external program. This unlocks programmatic embedding of the TUI from IDE extensions, web frontends, CI agents, or automation scripts without forcing them to give up the rich interactive UI in favor of --output-format=stream-json. ## Design The TUI already has a battle-tested JSON event emitter (`StreamJsonOutputAdapter`). This change makes that adapter pluggable on its output stream and wires a small `DualOutputBridge` that forwards TUI events to a second instance of the adapter writing to fd / file. For tool approvals, when a tool enters awaiting_approval the bridge emits `control_request` (subtype `can_use_tool`); whichever side resolves first (TUI's native UI or `confirmation_response` via --input-file) wins, and a `control_response` is mirrored back so all observers stay in sync. `session_start` is announced once when the bridge is constructed so consumers can correlate the channel with a session before any other event arrives. ## CLI surface - `--json-fd <n>` — write JSON events to fd n (n >= 3; provided via spawn stdio). - `--json-file <path>` — write JSON events to a file / FIFO / /dev/fd/N. - `--input-file <path>` — watch this file for JSONL commands. `--json-fd` and `--json-file` are mutually exclusive. fds 0/1/2 are rejected to prevent corrupting the TUI. ## Wire protocol Output: existing stream-json schema with `includePartialMessages` always enabled, plus: - `system` / `subtype: session_start` — emitted once on bridge construction. - `control_request` / `subtype: can_use_tool` — pending tool approval. - `control_response` — final approval outcome (mirrors TUI-native or external resolution). Input (--input-file): {"type":"submit","text":"What does this function do?"} {"type":"confirmation_response","request_id":"...","allowed":true} `submit` is queued and retried when the TUI returns to idle. `confirmation_response` is dispatched immediately — a pending tool call is blocking and the response cannot wait behind earlier submits. See `docs/users/features/dual-output.md` for the full schema, latency notes, failure modes, and a spawn example. ## What changes when the flags are absent Nothing. The bridge and watcher are constructed only when the relevant flags are set; otherwise the React Context providers carry `null` and every callsite short-circuits. No overhead, no behavioral change for existing users. ## Failure handling - Bad fd / unopenable path → warning on stderr, dual output stays disabled, TUI launches normally. - Consumer disconnect (EPIPE) → bridge silently disables itself, TUI keeps running. - Any exception inside the adapter → caught, logged, bridge disabled. The TUI is never crashed by a dual-output failure. ## Files New: - packages/cli/src/dualOutput/{DualOutputBridge,DualOutputContext,index}.{ts,tsx} - packages/cli/src/remoteInput/{RemoteInputWatcher,RemoteInputContext,index}.{ts,tsx} - packages/cli/src/nonInteractive/io/index.ts - docs/users/features/dual-output.md Modified: - packages/core/src/config/config.ts — 3 new ConfigParameters fields + getters - packages/cli/src/config/config.ts — yargs options + mutex validation - packages/cli/src/gemini.tsx — instantiate bridge / watcher in startInteractiveUI, wrap with Context Providers, register cleanup - packages/cli/src/ui/AppContainer.tsx — connect RemoteInput to submitQuery, bridge tool confirmations - packages/cli/src/ui/hooks/useGeminiStream.ts — call dualOutput?.processEvent(...) at five existing event points - packages/cli/src/nonInteractive/io/{Base,Stream}JsonOutputAdapter.ts — StreamJsonOutputAdapter accepts an injected output stream; base adapter exposes emitPermissionRequest / emitControlResponse through a new emitControlMessageImpl hook (default no-op in batch mode). ## Tests - packages/cli/src/dualOutput/DualOutputBridge.test.ts — fd validation, auto session_start, control-event routing, post-shutdown safety. - packages/cli/src/remoteInput/RemoteInputWatcher.test.ts — submit forwarding, immediate confirmation dispatch, busy/idle retry, malformed-line tolerance, shutdown. - packages/cli/src/nonInteractive/io/StreamJsonOutputAdapter.dualOutput.test.ts — custom outputStream injection and new emitPermissionRequest / emitControlResponse paths. tsc --noEmit -p packages/cli/tsconfig.json is clean. vitest run src/nonInteractive src/dualOutput src/remoteInput → 297 passed, 1 skipped, 11 files. * feat(cli): dual-output capability handshake, session_end, control_error, settings.json Incremental improvements on top of the initial dual-output PR based on reviewer feedback. All extensions are additive; older consumers that ignore unknown fields keep working. ## Capability handshake in session_start `session_start.data` now carries three new fields so consumers can feature-detect without sniffing the stream: - `protocol_version` (integer, currently 1) — bumped on any protocol change consumers might care about. - `version` (string) — the Qwen Code CLI version, threaded in from `gemini.tsx`. - `supported_events` (string[]) — the event kinds this bridge version is known to emit, exported as `SUPPORTED_EVENTS` from the module. ## session_end on bridge shutdown DualOutputBridge.shutdown() now emits a final `system` / `session_end` event carrying `session_id` before closing the stream. Gives consumers a definitive termination signal rather than requiring them to infer it from EPIPE. Idempotent — calling shutdown twice emits exactly one session_end. ## control_error emission path `ControlErrorResponse` (already defined in types.ts) now has a first- class emission path: `BaseJsonOutputAdapter.emitControlError(requestId, message)` → `control_response` with `subtype: 'error'`. Wired into AppContainer's remote-input confirmation handler so that a `confirmation_response` referencing an unknown / already-resolved request_id produces a structured error reply instead of silently dropping, letting consumers retry or surface the error. ## settings.json support New `dualOutput` top-level settings block with `jsonFile` and `inputFile` properties. `--json-fd` has no settings equivalent (fd passing is a spawn-time concern). CLI flag wins over settings when both are present, so scripted one-off runs still work unchanged. `requiresRestart: true` since the bridge is constructed once at startup. ## Documentation `docs/users/features/dual-output.md` gains three major sections: - Use cases — concrete integration scenarios (terminal+chat dual sync, IDE extensions, web frontends, CI observers, multi-agent orchestration, session replay, observability, QA). - Why two output flags? — detailed rationale for coexisting `--json-fd` and `--json-file`, including the PTY constraint (`node-pty` / `bun-pty` expose no stdio array, and `forkpty(3)` / `login_tty` actively close fds >= 3 before exec). - Comparison with Claude Code's stream-json — schema-parity matrix, transport-topology differences, permission-control-plane behavioral notes, and a "room to improve" section as a design horizon. - Runnable demos — seven copy-paste POCs: event observer, remote submit, permission bridge, Node embedder with capability feature-detection, session_end handling, failure drills. - Settings-based configuration — example settings.json snippet and precedence rules. ## Tests - DualOutputBridge.test.ts: new cases for capability handshake shape, session_end on shutdown, shutdown idempotency, and emitControlError. - StreamJsonOutputAdapter.dualOutput.test.ts: new case for emitControlError at the adapter level. 302 passed, 1 skipped, 11 files. tsc --noEmit -p packages/cli is clean. * docs(dual-output): shrink Claude Code comparison to one honest sentence After actually reading the Claude Code source (src/cli/structuredIO.ts, src/bridge/, src/utils/messages/systemInit.ts), the previous "Comparison with Claude Code's stream-json" section was overstated: - Claude Code has no equivalent of TUI + sidecar running simultaneously. Its stream-json only works with --print (non-interactive); the bridge in src/bridge/ is Anthropic's own remote worker protocol, not a local embedding surface. - CC uses `system/init` (not `session_start`) and has no session_end in the wire protocol, so the schema-parity table contained false ticks. - Framing this PR as "parity with Claude Code" is therefore inaccurate; it's filling a gap Claude Code does not address. Replace the whole multi-section comparison (schema matrix, transport table, permission notes, borrow list, roadmap) with a single sentence stating the accurate relation: same event format in spirit, different topology — CC's is non-interactive only. * fix(cli): address review feedback on dual-output sidecar mode - Fix control_response mirror: external-initiated confirmations now emit control_response via the same mirror useEffect as TUI-native resolutions, making the emission path symmetric for all observers. - Fix ENOENT: --json-file with a non-existent path now falls back to createWriteStream (auto-creates the file) instead of throwing. - Fix race: add reading guard to RemoteInputWatcher.readNewLines() preventing duplicate command processing on rapid appends. - Refactor confirmationHandler to use refs (pendingToolCallsRef, dualOutputRef) and register once (deps: [remoteInput]) to eliminate teardown/re-registration churn. - Add debug logging to shutdown bare catch for ops correlation. - Add ENOENT fallback test case for DualOutputBridge. - Regenerate settings.schema.json for dualOutput section. Generated with AI Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com> * fix(cli): make RemoteInputWatcher poll interval configurable for CI reliability RemoteInputWatcher.test.ts was timing out in CI (5s default) because fs.watchFile's 500ms poll interval is unreliable under load. Fix: - Accept optional `pollIntervalMs` in constructor (default 500ms). - Tests use 100ms poll interval for faster feedback. - Increase per-test timeout to 15s and waitFor timeout to 10s. - Increase "TUI busy" wait from 800ms to 1500ms for CI headroom. Generated with AI Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com> * fix(cli): eliminate fs.watchFile timing dependency in RemoteInputWatcher tests Tests were flaky across all CI platforms (macOS/ubuntu/windows) because fs.watchFile polling (even at 100ms) is unreliable under CI load. Fix: expose checkForNewInput() as a public method that directly triggers file reading and returns a Promise. Tests now call it synchronously after writing to the input file — no polling, no timeouts, deterministic. Also fixes: - Windows ENOTEMPTY: add delay in afterEach before rmSync - Add active check in readNewLines to respect shutdown state - readNewLines now returns Promise<void> for awaitable reads Generated with AI Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com> --------- Co-authored-by: 秦奇 <gary.gq@alibaba-inc.com> Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>	2026-04-18 02:14:53 +08:00
Shaojin Wen	355ac5d54a	feat(core): add path-based context rule injection from .qwen/rules/ (#3339 ) Some checks are pending Qwen Code CI / Lint (push) Waiting to run Details Qwen Code CI / Test (push) Blocked by required conditions Details Qwen Code CI / Test-1 (push) Blocked by required conditions Details Qwen Code CI / Test-2 (push) Blocked by required conditions Details Qwen Code CI / Test-3 (push) Blocked by required conditions Details Qwen Code CI / Test-4 (push) Blocked by required conditions Details Qwen Code CI / Test-5 (push) Blocked by required conditions Details Qwen Code CI / Test-6 (push) Blocked by required conditions Details Qwen Code CI / Test-7 (push) Blocked by required conditions Details Qwen Code CI / Test-8 (push) Blocked by required conditions Details Qwen Code CI / Post Coverage Comment (push) Blocked by required conditions Details Qwen Code CI / CodeQL (push) Waiting to run Details E2E Tests / E2E Test (Linux) - sandbox:docker (push) Waiting to run Details E2E Tests / E2E Test (Linux) - sandbox:none (push) Waiting to run Details E2E Tests / E2E Test - macOS (push) Waiting to run Details * feat(core): add path-based context rule injection from .qwen/rules/ Support multiple rule files in `.qwen/rules/` directories with optional YAML frontmatter for conditional loading based on glob patterns. Rules with a `paths:` field only load when matching files exist in the project. Rules without `paths:` always load as baseline rules. Key behaviors: - Global rules from ~/.qwen/rules/ always load - Project rules from <root>/.qwen/rules/ require folder trust - HTML comments stripped to save tokens - Files sorted alphabetically for deterministic ordering - Deduplication when project root equals home directory - Uses globIterate for early termination on first match * feat(core): align rules loading with Claude Code reference implementation Closes three gaps with Claude Code's .claude/rules/ feature: 1. Recursive directory scanning — .qwen/rules/ now supports subdirectories like frontend/, backend/ for organized rule hierarchies. 2. Exclusion patterns — new `contextRuleExcludes` config parameter accepts glob patterns to skip specific rule files (useful in monorepos with other teams' rules). 3. Turn-level lazy loading — conditional rules (with `paths:` frontmatter) are no longer injected eagerly at session start. Instead, they are stored in a per-session ConditionalRulesRegistry and injected on-demand via <system-reminder> when the model reads/edits a matching file (read_file, edit, write_file). Each rule is injected at most once per session. Internals: - loadRules() now returns { content, ruleCount, conditionalRules } — only baseline rules flow into the system prompt; conditional rules are deferred. - ConditionalRulesRegistry pre-compiles picomatch matchers for efficiency and tracks injected rules to avoid duplicate injection. - coreToolScheduler.ts injects matched rules after PostToolUse hooks but before the tool response is sent to the model. - Path matching defensively rejects files outside the project root. - /memory refresh and /directory add keep the registry in sync via setConditionalRulesRegistry(). * fix(core): correct field placement in config.test.ts mocks after merge Earlier replace_all inserted ruleCount/conditionalRules/projectRoot into the wrong mock call (readAutoMemoryIndex instead of loadServerHierarchicalMemory), breaking the build with syntax errors. Move the fields back to the correct mocked return value. * fix(core): normalize rule display paths to forward slashes for Windows On Windows, path.relative() returns backslash-separated paths, causing the "Rule from:" marker to differ from Linux/macOS and breaking the formats-rules-with-source-markers test on Windows CI. Normalize to forward slashes for cross-platform consistency, matching the convention used in glob patterns (paths: field) so that the model sees the same format regardless of the host OS. * fix(core): harden rulesDiscovery path checks and sort determinism Two small defensive improvements surfaced by the audit: 1. matchAndConsume now rejects the exact '..' relative path in addition to '../'-prefixed paths. path.relative returns '..' (no trailing slash) when the target equals the parent of projectRoot — rare in practice but worth guarding against. 2. loadRulesFromDir now uses Array.sort() default (UTF-16 code point comparison) instead of localeCompare. The previous sort was locale-dependent and could produce different rule loading order on machines with non-English locales (e.g. zh-CN). Rule filenames are typically ASCII so behaviour is unchanged in common cases, but deterministic ordering is preferable across environments. Adds one test case for the '..' rejection path. * fix(core): address CodeQL incomplete HTML comment sanitization stripHtmlComments only matched complete <!-- ... --> pairs in a single pass, so input like 'A<!-- one --><!-- two -->B<!--unclosed' would leave a residual '<!--' marker — flagged by CodeQL as incomplete-multi-character-sanitization. Not a security issue in our context (the output goes to an LLM system prompt, not an HTML renderer), but worth fixing to: - clear the CodeQL alert in CI - avoid token waste from dangling markers - produce deterministic output Strategy: iteratively strip <!-- ... --> pairs until stable, then remove any residual <!-- markers (leaving the following content visible since the author probably intended it to appear in the rule).	2026-04-17 22:05:50 +08:00
tanzhenxin	7e83c08062	feat: background subagents with headless and SDK support (#3076 ) * feat(core): add run_in_background support for Agent tool Enable sub-agents to run asynchronously via `run_in_background: true` parameter. Background agents execute independently from the parent, which receives an immediate launch confirmation and continues working. A notification is injected into the parent conversation when the background agent completes. Key changes: - BackgroundTaskRegistry tracks lifecycle of background agents - Agent tool gains async execution path with fire-and-forget semantics - Background agents use YOLO approval mode to prevent deadlock - Independent AbortControllers survive parent ESC cancellation - CLI bridges notifications via useMessageQueue for between-turn delivery - State race guards prevent complete/fail after cancellation - Session cleanup aborts all running background agents * feat(background): improve notification formatting and UI handling - Add prefix/separator protocol to distinguish background notifications from user input - Show concise summary in UI while sending full details to LLM - Add 'notification' history item type with specialized display - Add 'background' agent status for background-running agents - Prevent notifications from polluting prompt history (up-arrow) - Truncate long descriptions in display text This improves the UX for background agents by showing cleaner, more concise notifications while preserving full context for the LLM. Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com> * fix(background): reject run_in_background in non-interactive mode Headless mode skips AppContainer, so the notification callback is never registered and background agent results would be silently dropped. Return an error prompting the model to retry without run_in_background. * refactor(background): replace prefix/separator protocol with typed notification queue Replace the stringly-typed \x00__BG_NOTIFY__\x00 prefix/separator encoding with a typed notification path using SendMessageType.Notification. - Add SendMessageType.Notification to the enum - Change BackgroundNotificationCallback to emit (displayText, modelText) - Move notification queue from AppContainer into useGeminiStream (mirrors the cron queue pattern): register on registry, queue structured items, drain on idle via submitQuery - prepareQueryForGemini short-circuits for Notification type (skips slash commands, shell mode, @-commands, prompt history logging) - Remove BACKGROUND_NOTIFICATION_PREFIX/SEPARATOR constants * refactor(background): move abortAll to Config.shutdown Background agent cleanup belongs in Config.shutdown() alongside other resource teardown (skillManager, toolRegistry, arenaRuntime), not in AppContainer's registerCleanup. This also ensures headless mode gets cleanup for free. * fix(background): persist notification items for session resume Background agent notifications were missing after session resume because they were never recorded in the chat history. The model text was absent from the API history and the display item was lost. - Add recordNotification() to ChatRecordingService — stores as user-role message with subtype 'notification' and displayText payload - Thread notificationDisplayText through submitQuery → sendMessageStream - Restore as HistoryItemNotification in resumeHistoryUtils * fix(background): replace YOLO with deny-by-default for background agents Background agents were using YOLO approval mode which auto-approves all tool calls — too permissive. Replace with shouldAvoidPermissionPrompts which auto-denies tool calls that need interactive approval, matching claw-code's approach. The permission flow for background agents is now: 1. L3/L4 permission rules (allow/deny) — same as foreground 2. Approval mode overrides (AUTO_EDIT for edits) — same as foreground 3. PermissionRequest hooks — can override the denial 4. Auto-deny — if no hook decided, deny because prompts are unavailable * fix(background): add missing getBackgroundTaskRegistry mock in useGeminiStream tests * refactor(core): move fork subagent params from execute() to construction time Identity-shaping fork inputs (parent history, generationConfig, tool decls, env-skip flag) were threaded through `AgentHeadless.execute()`'s options bag and re-passed by the SubagentStop hook retry loop. They belong on the agent's construction-time configs, not its per-invocation options. - PromptConfig gains `renderedSystemPrompt` (verbatim, bypasses templating and userMemory injection) and drops the `systemPrompt`/`initialMessages` XOR so fork can carry both. createChat skips env bootstrap when `initialMessages` is non-empty. - AgentHeadless.execute() shrinks to (context, signal?). Fork dispatch in agent.ts builds synthetic PromptConfig/ModelConfig/ToolConfig from the parent's cache-safe params and calls AgentHeadless.create directly (bypassing SubagentManager). Parent's tool decls flow through verbatim including the `agent` tool itself for cache parity. - Recursive-fork prevention switches from fork-side tool stripping to a runtime guard. The previous `isInForkChild(history)` helper was dead code (it scanned the main GeminiClient's history, not the fork child's chat). Replaced with `isInForkExecution()` backed by AsyncLocalStorage: the fork's background execution runs inside `runInForkContext`, and the ALS frame propagates through the standard async chain into nested AgentTool.execute() calls where the guard fires. * refactor(core): move agent tool files into dedicated tools/agent/ directory Move agent.ts, agent.test.ts, and fork-subagent.ts under tools/agent/ and update all import paths accordingly. * refactor(core): remove dead temp and top_p fields from ModelConfig These fields were never populated from subagent frontmatter and served no purpose in the fork path either. The ModelConfig interface retains only the actively-used model field. * refactor(core): read parent generation config directly instead of getCacheSafeParams Fork subagent now reads system instruction and tool declarations from the live GeminiChat via getGenerationConfig() instead of the global getCacheSafeParams() snapshot. This removes the cross-module coupling between the agent tool and the followup infrastructure. * fix(core): prevent duplicate tool declarations when toolConfig has only inline decls prepareTools() treated asStrings.length === 0 as "add all registry tools", which is correct when no tools are specified at all, but wrong when the caller provides only inline FunctionDeclaration[] (no string names). The fork path passes parent tool declarations as inline decls for cache parity, so prepareTools was adding the full registry set on top — duplicating every non-excluded tool. Add onlyInlineDecls.length === 0 to the condition so that pure-inline toolConfigs bypass the registry entirely. * feat(core): support agent-level `background: true` in frontmatter Subagent definitions can now declare `background: true` in their YAML frontmatter to always run as background tasks. This is OR'd with the `run_in_background` tool parameter — useful for monitors, watchers, and proactive agents so the LLM doesn't need to remember to set the flag. * fix(core): address background subagent lifecycle gaps - Inherit bgConfig from agentConfig so the resolved approval mode is preserved for background agents (foreground would run AUTO_EDIT but background fell back to DEFAULT, which combined with shouldAvoid- PermissionPrompts would auto-deny every permission request). - Honor SubagentStop blocking decisions in background runs by looping on hook output up to 5 iterations, matching runSubagentWithHooks. - Check terminate mode before reporting completion; non-GOAL modes (ERROR, MAX_TURNS, TIMEOUT) are now reported as failures instead of emitting a success notification for an incomplete run. - Exclude SendMessageType.Notification from the UserPromptSubmit hook guard so background completion messages are not rewritten or blocked as if they were user input. * feat(cli): headless support and SDK task events for background agents (#3379) * feat(cli): unify notification queue for cron and background agents Migrate cron from its own queue (cronQueueRef / cronQueue) to the shared notification queue used by background agents. Both producers now push the same item shape { displayText, modelText, sendMessageType } and a single drain effect / helper processes them in FIFO order. Cron fires render as HistoryItemNotification (● prefix) instead of HistoryItemUser (> prefix), with a "Cron: <prompt>" display label. Records use subtype 'cron' for clean resume and analytics separation. Lift the non-interactive rejection for background agents. Register a notification callback in nonInteractiveCli.ts with a terminal hold-back phase (100ms poll) that keeps the process alive until all background agents complete and their notifications are processed. * feat(cli): emit SDK task events for background subagents Emit `task_started` when a background agent registers and `task_notification` when it completes, fails, or is cancelled, so headless/SDK consumers can track lifecycle without parsing display text. Model-facing text is now structured XML with status, summary, truncated result, and usage stats. Completion stats (tokens, tool uses, duration) are captured from the subagent and included in both the SDK payload and the model XML. * fix: address codex review issues for background subagents - Background subagents now inherit the resolved approval mode from agentConfig instead of the raw session config, so a subagent with `approvalMode: auto-edit` (or execution in a trusted folder) keeps that override when it runs asynchronously. - Non-interactive cron drains are single-flight: concurrent cron fires now await the same in-flight drain, and the cron-done check gates on it, preventing the final result from being emitted while a cron turn is still streaming. - Background forks go through createForkSubagent so they retain the parent's rendered system prompt and inherited history instead of degrading to a plain FORK_AGENT. * fix(cli): restore cancellation, approval, and error paths in queued drain - Hold-back loop now reacts to SIGINT/SIGTERM: when the main abort signal fires it calls registry.abortAll() so background agents with their own AbortControllers stop promptly instead of pinning the process open. - Queued-turn tool execution forwards the stream-json approval update callback (onToolCallsUpdate) so permission-gated tools inside a background-notification follow-up emit can_use_tool requests. - Queued-turn stream loop mirrors the main loop's text-mode handling of GeminiEventType.Error, writing to stderr and throwing so provider errors produce a non-zero exit code instead of silently succeeding. - Interactive cron prompts go through the normal slash/@-command/shell preprocessing again; only Notification messages skip that path. * fix(cli): skip duplicate user-message item for cron prompts Cron prompts already render as a `● Cron: …` notification via the queue drain, so adding them again as a `USER` history item produced a duplicate `> …` line. * fix(cli): honor SIGINT/SIGTERM during cron scheduler wait The non-interactive cron phase awaits a Promise that resolves only when scheduler.size reaches 0 and no drain is in flight. Recurring cron jobs never drop the scheduler size to 0 on their own, so the previous abort handling (added to the hold-back loop) was unreachable — the process hung indefinitely after SIGINT/SIGTERM. Attach an abort listener inside the promise so abort stops the scheduler and resolves immediately, allowing the hold-back loop to run and the process to exit cleanly. * feat(core): propagate tool-use id through background agent notifications Plumb the scheduler's callId into AgentToolInvocation via an optional setCallId hook on the invocation, detected structurally in buildInvocation. The agent tool forwards it as toolUseId on the BackgroundTaskRegistry entry so completion notifications can carry a <tool-use-id> tag and SDK task_started / task_notification events can emit tool_use_id — letting consumers correlate background completions back to the original Agent tool-use that spawned them. * fix(cli): drain single-flight race kept task_notification from emitting drainLocalQueue wrapped its body in an async IIFE and cleared the promise reference via finally. When the queue is empty the IIFE has no awaits, so its finally runs synchronously as part of the RHS of the assignment `drainPromise = (async () => {...})()` — clearing drainPromise BEFORE the outer assignment overwrites it with the resolved promise. The reference then stayed stuck on that fulfilled promise forever, so later calls short-circuited through `if (drainPromise) return drainPromise` and never processed queued notifications. Symptom: in headless `--output-format json` (and `stream-json`), task_started emitted but task_notification never did, even after the background agent completed. The process sat in the hold-back loop until SIGTERM. Fix: move the null-clearing out of the async body into an outer `.finally()` on the returned promise. `.finally()` runs as a microtask after the current synchronous block, so it clears the latest drainPromise reference instead of the pre-assignment null. * fix(cli): append newline to text-mode emitResult so zsh PROMPT_SP doesn't erase the line Headless text mode wrote `resultMessage.result` without a trailing newline. In a TTY, zsh themes that use PROMPT_SP (powerlevel10k, agnoster, …) detect the missing `\n` and emit `\r\033[K` before drawing the next prompt, which wipes the final line off the screen. Pipe-captured output was unaffected, so the bug only surfaced for interactive shell users — most visibly in the background-agent flow where the drain-loop's final assistant message is the only stdout write in text mode. Append `\n` to both the success (stdout) and error (stderr) writes. * docs(skill): tighten worked-example blurb in structured-debugging Mirror the simplified blurb from .claude/skills/structured-debugging/SKILL.md (knowledge repo). Drops the round-by-round narrative; keeps the contradiction + two lessons. * docs(skill): mirror SKILL.md improvements (reframing failure mode, generalized path, value-logging guidance) Mirror of knowledge repo commit 38eb28d into the qwen-code .qwen/skills copy. * docs(skill): mirror worked example into .qwen/skills/structured-debugging/ Mirrors knowledge/.claude/skills/structured-debugging/examples/ headless-bg-agent-empty-stdout.md so the .qwen copy of the skill links resolve. * docs(skill): mirror generalized side-note path guidance * fix(cli): harden headless cron and background-agent failure paths Three regressions surfaced by Codex review of feat/background-subagent: - Cron drain rejections were dropped by a bare `void`, so a failing queued turn left the outer Promise unresolved and hung the run. Route drain failures through the Promise's reject so they propagate to the outer catch. - The background-agent registry entry was inserted before `createForkSubagent()` / `createAgentHeadless()` was awaited. Failed init returned an error from the tool call but left a phantom `running` entry, and the headless hold-back loop (`registry.getRunning()`) waited forever. Register only after init succeeds. - SIGINT/SIGTERM during the hold-back phase aborted background tasks, then fell through to `emitResult({ isError: false })`, so a cancelled `qwen -p ...` exited 0 with the prior assistant text. Route through `handleCancellationError()` so cancellation exits non-zero, matching the main turn loop. * test(cli): update stdout/stderr assertions for trailing newline ``feadf052f`` appended `\n` to text-mode `emitResult` output, but the nonInteractiveCli tests still asserted the pre-change strings. Update the 11 affected assertions to expect the trailing newline. * fix: address review comments on background-agent notifications Four additional issues from the PR review that the prior regression-fix commit didn't cover: - Escape XML metacharacters when interpolating `description`, `result`, `error`, `agentId`, `toolUseId`, and `status` into the task-notification envelope. Subagent output (which itself may carry untrusted tool output, fetched HTML, or another agent's notification) could contain `</result>` or `</task-notification>` and forge sibling tags the parent model would treat as trusted metadata. Truncate result text before escaping so the truncation never slices through an entity like `&`. - Emit the terminal notification from `cancel()` and `abortAll()`. The fire-and-forget `complete()`/`fail()` from the subagent task is guarded by `status !== 'running'` and was no-op'd after cancellation, so SDK consumers saw `task_started` with no matching `task_notification`, breaking the contract this PR establishes. Updated two race-guard tests that asserted the old behavior. - Call `adapter.finalizeAssistantMessage()` before the abort-triggered early return inside `drainOneItem`'s stream loop. Without it, `startAssistantMessage()` had already been called, so stream-json mode left `message_start` unpaired. - Enforce `config.getMaxSessionTurns()` in `drainOneItem` for symmetry with the main turn loop. Cron fires and notification replies otherwise bypass the budget cap in headless runs. * fix: address codex review comments for background subagents - Wrap background fork execute() in runInForkContext so the recursive-fork guard (AsyncLocalStorage-based) fires when a background fork's child model calls `agent` again. Previously only the foreground fork path was wrapped, so background forks could spawn nested implicit forks. - Emit queued terminal task_notifications on SIGINT/SIGTERM before handleCancellationError exits. abortAll() enqueues cancellation notifications via the registry callback, but the process was exiting before the drain loop had a chance to flush them — leaving stream-json consumers that already saw task_started without a matching terminal task_notification. Extracted the SDK-emit block into a shared emitNotificationToSdk helper reused by the normal drain and the cancellation flush. - Skip notification/cron subtypes in ACP HistoryReplayer. These records are persisted as type: 'user' so the model's chat history keeps them for continuity, but they were never user input — replaying them leaked raw <task-notification> XML (and cron prompts) back into the ACP session as if the user typed them. * test(cli): sync JsonOutputAdapter text-mode assertions with trailing newline Commit `0da1182b7` appended a newline to text-mode emitResult output (zsh PROMPT_SP fix) and updated the nonInteractiveCli tests, but four assertions in JsonOutputAdapter.test.ts were missed. Update them to expect the trailing newline so CI passes. * refactor: simplify background subagent plumbing - Extract the SubagentStop hook blocking-decision loop into a runSubagentStopHookLoop helper so the foreground and background paths no longer duplicate the iteration/abort/log scaffolding. - Unify BackgroundTaskRegistry.abortAll to delegate to cancel, removing copy-pasted abort/notification bookkeeping. - Drop the unused findByName and BackgroundAgentEntry.name field. - In nonInteractiveCli drain, hoist inputFormat and toolCallUpdateCallback out of the inner tool loop, and drop the unreachable try/catch around the readonly registry. - Trim boilerplate doc/narration comments while keeping load-bearing WHY comments. * fix: address codex review comments for background subagents - Use tool callId (or short random suffix) instead of Date.now() for background agentIds; avoids registry collisions when parallel same-type agents launch in the same millisecond. - Reset loopDetector and lastPromptId for Notification turns so a prior turn's loop count doesn't trip LoopDetected on the notification response. - Replay notification/cron displayText in ACP HistoryReplayer so the assistant reply has an antecedent in resumed transcripts. --------- Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>	2026-04-17 18:23:06 +08:00
jinye	2080686a62	feat(skills): add /batch skill for parallel batch operations (#3079 ) * feat(skills): add /batch skill for parallel batch operations Add a new built-in skill `/batch` that orchestrates large-scale parallel changes across multiple files. The skill automatically: - Discovers target files using glob patterns - Splits files into chunks for parallel processing - Launches multiple worker agents concurrently - Aggregates results with success/failure statistics - Supports --dry-run mode for preview Resolves #3043 * fix(skills): address PR review feedback for /batch skill - Extend chunking table to cover 51-100 files (now uses 5 chunks) - Clarify that `task` tool is the Agent tool for spawning workers - Add SKIPPED status to Agent Prompt Template output format - Shorten description field, move examples to body - Expand test file exclusion patterns (.test.js, __tests__/, etc.) - Expand Dry-Run Mode section with detailed example output - Fix example math: "24 files → 3 chunks of ~8 each" fix(skills): address wenshao's review feedback for /batch skill - Add copyright header - Add ask_user_question to allowedTools - Fix edit_file -> edit (canonical tool name) - Fix chunking table math (3-8 each, not ~7-8) - Add zero file match handling - Reframe dry-run from --flag to natural language pattern * fix(skills): remove unnecessary copyright header from SKILL.md --------- Co-authored-by: jinye.djy <jinye.djy@alibaba-inc.com>	2026-04-17 15:52:21 +08:00
tanzhenxin	503c9c638d	refactor(core): move fork subagent params from execute() to construction time (#3255 ) * refactor(core): move fork subagent params from execute() to construction time Identity-shaping fork inputs (parent history, generationConfig, tool decls, env-skip flag) were threaded through `AgentHeadless.execute()`'s options bag and re-passed by the SubagentStop hook retry loop. They belong on the agent's construction-time configs, not its per-invocation options. - PromptConfig gains `renderedSystemPrompt` (verbatim, bypasses templating and userMemory injection) and drops the `systemPrompt`/`initialMessages` XOR so fork can carry both. createChat skips env bootstrap when `initialMessages` is non-empty. - AgentHeadless.execute() shrinks to (context, signal?). Fork dispatch in agent.ts builds synthetic PromptConfig/ModelConfig/ToolConfig from the parent's cache-safe params and calls AgentHeadless.create directly (bypassing SubagentManager). Parent's tool decls flow through verbatim including the `agent` tool itself for cache parity. - Recursive-fork prevention switches from fork-side tool stripping to a runtime guard. The previous `isInForkChild(history)` helper was dead code (it scanned the main GeminiClient's history, not the fork child's chat). Replaced with `isInForkExecution()` backed by AsyncLocalStorage: the fork's background execution runs inside `runInForkContext`, and the ALS frame propagates through the standard async chain into nested AgentTool.execute() calls where the guard fires. * refactor(core): move agent tool files into dedicated tools/agent/ directory Move agent.ts, agent.test.ts, and fork-subagent.ts under tools/agent/ and update all import paths accordingly. * refactor(core): remove dead temp and top_p fields from ModelConfig These fields were never populated from subagent frontmatter and served no purpose in the fork path either. The ModelConfig interface retains only the actively-used model field. * refactor(core): read parent generation config directly instead of getCacheSafeParams Fork subagent now reads system instruction and tool declarations from the live GeminiChat via getGenerationConfig() instead of the global getCacheSafeParams() snapshot. This removes the cross-module coupling between the agent tool and the followup infrastructure. * fix(core): prevent duplicate tool declarations when toolConfig has only inline decls prepareTools() treated asStrings.length === 0 as "add all registry tools", which is correct when no tools are specified at all, but wrong when the caller provides only inline FunctionDeclaration[] (no string names). The fork path passes parent tool declarations as inline decls for cache parity, so prepareTools was adding the full registry set on top — duplicating every non-excluded tool. Add onlyInlineDecls.length === 0 to the condition so that pure-inline toolConfigs bypass the registry entirely. * refactor(core): remove dead temp and skipEnvHistory fields from AgentPathParams These fields were carried over from earlier designs but have no remaining effect after the fork subagent refactor: - `temp` was never forwarded into ModelConfig, which this PR already stripped of the temperature field. - `skipEnvHistory` is redundant with the auto-skip in `AgentCore.createChat`, which already bypasses env bootstrap whenever `initialMessages` is non-empty — the condition under which any caller would set this flag. Also drops the corresponding `skipEnvHistory: true` at the one caller in the memory extraction planner.	2026-04-17 15:45:57 +08:00
顾盼	875f3ffeb4	fix(core): add shell argument quoting guidance to prevent special char errors (#3327 ) * fix(core): add shell argument quoting guidance to prevent special char errors When models pass arguments containing special characters (parentheses, backticks, single quotes, dollar signs, etc.) to shell commands like `gh pr create --body '...'`, bash can misinterpret them as shell syntax, causing the command to fail with cryptic errors. Add an explicit quoting guide to `getShellToolDescription()` covering: - Single quotes: pass everything literally but cannot contain `'` - ANSI-C quoting (`$'...'`): supports escape sequences including `\'` - Heredoc: the most robust approach for multi-line or mixed-quote text, with a concrete `gh pr create` example Fixes #3300 * test: update shell tool description snapshots	2026-04-17 15:37:22 +08:00
tanzhenxin	f7733cfc7e	fix(core): strip thinking blocks from history on model switch (#3304 ) (#3315 ) When switching models mid-session, reasoning_content fields from thinking-capable models leaked into API requests sent to the new provider, causing 422 errors on strict OpenAI-compatible endpoints. Call stripThoughtsFromHistory() in handleModelChange() so thought parts are removed before the next request is built for the new model.	2026-04-17 15:28:58 +08:00
tanzhenxin	0f8e8db9ae	fix(core): limit skill watcher depth to prevent FD exhaustion (#3320 ) * fix(core): limit skill watcher depth to prevent FD exhaustion (#3289) The chokidar file watcher in SkillManager.updateWatchersFromCache() had no depth limit or ignored paths. When skill directories contained heavy subtrees like node_modules, chokidar recursively watched every file, exhausting file descriptors and breaking child-process I/O (node-pty onData/onExit callbacks silently stop firing). Fix: set depth to 2 (skills use a fixed <skill-name>/SKILL.md layout) and add an ignored function that filters out special file types (sockets, FIFOs, devices) and .git directories. Made-with: Cursor * fix(core): use path.join in watcher test for Windows compat The watcherIgnored test used hardcoded forward-slash paths which don't split correctly on Windows where path.sep is backslash. Made-with: Cursor	2026-04-17 15:28:46 +08:00
Shaojin Wen	b004450d7f	feat(cli): support multi-line status line output (#3311 ) Some checks are pending Qwen Code CI / Lint (push) Waiting to run Details Qwen Code CI / Test (push) Blocked by required conditions Details Qwen Code CI / Test-1 (push) Blocked by required conditions Details Qwen Code CI / Test-2 (push) Blocked by required conditions Details Qwen Code CI / Test-3 (push) Blocked by required conditions Details Qwen Code CI / Test-4 (push) Blocked by required conditions Details Qwen Code CI / Test-5 (push) Blocked by required conditions Details Qwen Code CI / Test-6 (push) Blocked by required conditions Details Qwen Code CI / Test-7 (push) Blocked by required conditions Details Qwen Code CI / Test-8 (push) Blocked by required conditions Details Qwen Code CI / Post Coverage Comment (push) Blocked by required conditions Details Qwen Code CI / CodeQL (push) Waiting to run Details E2E Tests / E2E Test (Linux) - sandbox:docker (push) Waiting to run Details E2E Tests / E2E Test (Linux) - sandbox:none (push) Waiting to run Details E2E Tests / E2E Test - macOS (push) Waiting to run Details * feat(cli): support multi-line status line output (#3211) Remove the single-line hard limit (.split('\n')[0]) from the status line hook so user scripts can output multiple rows. Footer renders each line as a separate <Text wrap="truncate"> element, preserving per-line horizontal truncation. Ink's virtual DOM handles re-rendering without manual ANSI cursor management. * feat(cli): cap status line output at 3 lines Prevent runaway scripts from flooding the footer — lines beyond the third are silently discarded. * docs: mention 3-line cap in status line docs and agent prompt * fix(cli): cap status line at 2 lines to keep footer within 3 rows Footer has a fixed bottom row (hint/mode indicator), so status line gets at most 2 lines to keep the total footer height at 3 rows max. * test(cli): improve useStatusLine coverage to 100% lines Add tests for: per-model metrics payload, contextWindowSize/version/ model fallbacks, config removal with pending debounce, command change cancelling pending debounce. * docs: update status line ASCII diagram for multi-line layouts Also fix TS error in test (null → null as never for mock return). * refactor(cli): return string[] from useStatusLine, filter empty lines Address review feedback: - Hook returns `lines: string[]` instead of `text: string \| null`, eliminating the join/split round-trip with Footer. - Filter empty lines before slicing so leading blanks don't eat real content (e.g. "\n\nreal content" no longer yields ["", ""]). - Export MAX_STATUS_LINES with comment explaining the 3-row constraint. - Use `status-line-${i}` as React key for clarity. * test(cli): add Footer multi-line rendering, \r\n, and pure-newline tests Address remaining review feedback: - Footer test: mock useStatusLine, verify multi-line rendering and hint suppression. - useStatusLine test: add \r\n line ending and pure-newline edge case. * fix(cli): align right footer indicators to top When the status line has multiple rows, the left column becomes taller than the right section. The outer Box defaults to `alignItems: stretch` which caused the indicators to visually center; add `alignItems="flex-start"` on the right Box so they stay anchored to the top row. Reported via e2e test in #3311.	2026-04-17 12:44:30 +08:00
Reid	12b24e2d28	test(core): stabilize glob truncation tests (#3322 ) Some checks are pending Qwen Code CI / Lint (push) Waiting to run Details Qwen Code CI / Test (push) Blocked by required conditions Details Qwen Code CI / Test-1 (push) Blocked by required conditions Details Qwen Code CI / Test-2 (push) Blocked by required conditions Details Qwen Code CI / Test-3 (push) Blocked by required conditions Details Qwen Code CI / Test-4 (push) Blocked by required conditions Details Qwen Code CI / Test-5 (push) Blocked by required conditions Details Qwen Code CI / Test-6 (push) Blocked by required conditions Details Qwen Code CI / Test-7 (push) Blocked by required conditions Details Qwen Code CI / Test-8 (push) Blocked by required conditions Details Qwen Code CI / Post Coverage Comment (push) Blocked by required conditions Details Qwen Code CI / CodeQL (push) Waiting to run Details E2E Tests / E2E Test (Linux) - sandbox:docker (push) Waiting to run Details E2E Tests / E2E Test (Linux) - sandbox:none (push) Waiting to run Details E2E Tests / E2E Test - macOS (push) Waiting to run Details * test(core): stabilize glob truncation tests Mock glob results in truncation-specific tests instead of creating large numbers of real files. This keeps the tests focused on GlobTool boundary logic and avoids filesystem timing issues on Windows CI. * test(cli): stabilize selection list scroll test Wait for the newly active item to render after rerendering the list so the scroll assertions do not read a stale frame on slower Windows CI runs.	2026-04-16 20:51:24 +08:00
顾盼	9e2f63a1ca	feat(memory): managed auto-memory and auto-dream system (#3087 ) * docs: add auto-memory implementation log * feat(core): add managed auto-memory storage scaffold * feat(core): load managed auto-memory index * feat(core): add managed auto-memory recall * feat(core): add managed auto-memory extraction * feat(cli): add managed auto-memory dream commands * feat(core): add auxiliary side-query foundation * feat(memory): add model-driven recall selection * feat(memory): add model-driven extraction planner * feat(core): add background task runtime foundation * feat(memory): schedule auto dream in background * feat(core): add background agent runner foundation * feat(memory): add extraction agent planner * feat(core): add dream agent planner * feat(core): rebuild managed memory index * feat(memory): add governance status commands * feat(memory): add managed forget flow * feat(core): harden background agent planning * feat(memory): complete managed parity closure * test(memory): add managed lifecycle integration coverage * feat: same to cc * feat(memory-ui): add memory saved notification and memory count badge Feature 3 - Memory Saved Notification: - Add HistoryItemMemorySaved type to types.ts - Create MemorySavedMessage component for rendering '● Saved/Updated N memories' - In useGeminiStream: detect in-turn memory writes via mapToDisplay's memoryWriteCount field and emit 'memory_saved' history item after turn - In client.ts: capture background dream/extract promises and expose via consumePendingMemoryTaskPromises(); useGeminiStream listens post-turn and emits 'Updated N memories' notification for background tasks Feature 4 - Memory Count Badge: - Add isMemoryOp field to IndividualToolCallDisplay - Add memoryWriteCount/memoryReadCount to HistoryItemToolGroup - Add detectMemoryOp() in useReactToolScheduler using isAutoMemPath - ToolGroupMessage renders '● Recalled N memories, Wrote N memories' badge at the top of tool groups that touch memory files Fix: process.env bracket-access in paths.ts (noPropertyAccessFromIndexSignature) Fix: MemoryDialog.test.tsx mock useSettings to satisfy SettingsProvider requirement * fix(memory-ui): auto-approve memory writes, collapse memory tool groups, fix MEMORY.md path Problem 1 - Auto-approve memory file operations: - write-file.ts: getDefaultPermission() checks isAutoMemPath; returns 'allow' for managed auto-memory files, 'ask' for all other files - edit.ts: same pattern Problem 2 - Feature 4 UX: collapse memory-only tool groups: - ToolGroupMessage: detect when all tool calls have isMemoryOp set (pure memory group) and all are complete; render compact '● Recalled/Wrote N memories (ctrl+o to expand)' instead of individual tool call rows - ctrl+o toggles expand/collapse when isFocused and group is memory-only - Mixed groups (memory + other tools) keep badge-at-top behaviour - Expanded state shows individual tool calls with '● Memory operations (ctrl+o to collapse)' header Problem 3 - MEMORY.md path mismatch: - prompt.ts: Step 2 now references full absolute path ${memoryDir}/MEMORY.md so the model writes to the correct location inside the memory directory, not to the parent project directory Fix tests: - write-file.test.ts: add getProjectRoot to mockConfigInternal - prompt.test.ts: update assertion to match full-path section header * fix(memory-ui): fix duplicate notification, broken ctrl+o, and Edit tool detection - Remove duplicate 'Saved N memories' notification: the tool group badge already shows 'Wrote N memories'; the separate HistoryItemMemorySaved addItem after onComplete was double-counting. Keep only the background-task path (consumePendingMemoryTaskPromises). - Remove ctrl+o expand: Ink's Static area freezes items on first render and cannot respond to user input. useInput/useState(isExpanded) in a Static item is a no-op. Removed the dead code; memory-only groups now always render as the compact summary (no fake interactive hint). - Fix Edit tool detection: detectMemoryOp was checking for 'edit_file' but the real tool name constant is 'edit'. Also removed non-existent 'create_file' (write_file covers all writes). Now editing MEMORY.md is correctly identified as a memory write op, collapses to 'Wrote N memories', and is auto-approved. * fix(dream): run /dream as a visible submit_prompt turn, not a silent background agent The previous implementation ran an AgentHeadless background agent that could take 5+ minutes with zero UI feedback — user saw a blank screen for the entire duration and then at most one line of text. Fix: /dream now returns submit_prompt with the consolidation task prompt so it runs as a regular AI conversation turn. Tool calls (read_file, write_file, edit, grep_search, list_directory, glob) are immediately visible as collapsed tool groups as the model works through the memory files — identical UX to Claude Code. Also export buildConsolidationTaskPrompt from dreamAgentPlanner so dreamCommand can reuse the same detailed consolidation prompt that was already written. * fix(memory): auto-allow ls/glob/grep on memory base directory Add getMemoryBaseDir() to getDefaultPermission() allow list in ls.ts, glob.ts, and grep.ts — mirrors the existing pattern in read-file.ts. Without this, ListFiles/Glob/Grep on ~/.qwen/* would trigger an approval dialog, blocking /dream at its very first step. * fix(background): prevent permission prompt hangs in background agents Match Claude Code's headless-agent intent: background memory agents must never block on interactive permission prompts. Wrap background runtime config so getApprovalMode() returns YOLO, ensuring any ask decision is auto-approved instead of hanging forever. Add regression test covering the wrapped approval mode. * fix(memory): run auto extract through forked agent Make managed auto-memory extraction follow the Claude Code architecture: background extraction now uses a forked agent to read/write memory files directly, instead of planning patches and applying them with a separate filesystem pipeline. Keep the old patch/model path only as fallback if the forked agent fails. Add regression tests covering the new execution path and tool whitelist. * refactor(memory): remove legacy extract fallback pipeline Delete the old patch/model/heuristic extraction path entirely. Managed auto-memory extract now runs only through the forked-agent execution flow, with no planner/apply fallback stages remaining. Also remove obsolete exports/tests and update scheduler/integration coverage to use the forked-agent-only architecture. * refactor(memory): move auxiliary files out of memory/ directory meta.json, extract-cursor.json, and consolidation.lock are internal bookkeeping files, not user-visible memories. Move them one level up to the project state dir (parent of memory/) so that the memory/ directory contains only MEMORY.md and topic files, matching the clean layout of the upstream reference implementation. Add getAutoMemoryProjectStateDir() helper in paths.ts and update the three path accessors + store.test.ts path assertions accordingly. * fix(memory): record lastDreamAt after manual /dream run The /dream command submits a prompt to the main agent (submit_prompt), which writes memory files directly. Because it bypasses dreamScheduler, meta.json was never updated and /memory always showed 'never'. Fix by: - Exporting writeDreamManualRunToMetadata() from dream.ts - Adding optional onComplete callback to SubmitPromptActionReturn and SubmitPromptResult (types.ts / commands/types.ts) - Propagating onComplete through slashCommandProcessor.ts - Firing onComplete after turn completion in useGeminiStream.ts - Providing the callback in dreamCommand.ts to write lastDreamAt * fix(memory): remove scope params from /remember in managed auto-memory mode --global/--project are legacy save_memory tool concepts. In managed auto-memory mode the forked agent decides the appropriate type (user/feedback/project/reference) based on the content of the fact. Also improve the prompt wording to explicitly ask the agent to choose the correct type, reducing the tendency to default to 'project'. * feat(ui): show '✦ dreaming' indicator in footer during background dream Subscribe to getManagedAutoMemoryDreamTaskRegistry() in Footer via a useDreamRunning() hook. While any dream task for the current project is pending or running, display '✦ dreaming' in the right section of the footer bar, between Debug Mode and context usage. * refactor(memory): align dream/extract infrastructure with Claude Code patterns Five improvements based on Claude Code parity audit: 1. Memoize getAutoMemoryRoot (paths.ts) - Add _autoMemoryRootCache Map, keyed by projectRoot - findCanonicalGitRoot() walks the filesystem per call; memoize avoids repeated git-tree traversal on hot-path schedulers/scanners - Expose clearAutoMemoryRootCache() for test teardown 2. Lock file stores PID + isProcessRunning reclaim (dreamScheduler.ts) - acquireDreamLock() writes process.pid to the lock file body - lockExists() reads PID and calls process.kill(pid, 0); dead/missing PID reclaims the lock immediately instead of waiting 2h - Stale threshold reduced to 1h (PID-reuse guard, same as CC) 3. Session scan throttle (dreamScheduler.ts) - Add SESSION_SCAN_INTERVAL_MS = 10min (same as CC) - Add lastSessionScanAt Map<projectRoot, number> to ManagedAutoMemoryDreamRuntime - When time-gate passes but session-gate doesn't, throttle prevents re-scanning the filesystem on every user turn 4. mtime-based session counting (dreamScheduler.ts) - Replace fragile recentSessionIdsSinceDream Set in meta.json with filesystem mtime scan (listSessionsTouchedSince) - Mirrors Claude Code's listSessionsTouchedSince: reads session JSONL files from Storage.getProjectDir()/chats/, filters by mtime > lastDreamAt - Immune to meta.json corruption/loss; no per-turn metadata write - ManagedAutoMemoryDreamRuntime accepts injectable SessionScannerFn for clean unit testing without real session files 5. Extraction mutual exclusion extended to write_file/edit (extractScheduler.ts) - historySliceUsesMemoryTool() now checks write_file/edit/replace/create_file tool calls whose file_path is within isAutoMemPath() - Previously only detected save_memory; missed direct file writes by the main agent, causing redundant background extraction * docs(memory): add user-facing memory docs, i18n for all locales, simplify /forget - Add docs/users/features/memory.md: comprehensive user-facing guide covering QWEN.md instructions, auto-memory behaviour, all memory commands, and troubleshooting; replaces the placeholder auto-memory.md - Update docs/users/features/_meta.ts: rename entry auto-memory → memory - Update docs/users/features/commands.md: add /init, /remember, /forget, /dream rows; fix /memory description; remove /init duplicate - Update docs/users/configuration/settings.md: add memory.* settings section (enableManagedAutoMemory, enableManagedAutoDream) between tools and permissions - Remove /forget --apply flag: preview-then-apply flow replaced with direct deletion; update forgetCommand.ts, en.js, zh.js accordingly - Add all auto-memory i18n keys to de, ja, pt, ru locales (18 keys each): Open auto-memory folder, Auto-memory/Auto-dream status lines, never/on/off, ✦ dreaming, /forget and /remember usage strings, all managed-memory messages - Remove dead save_memory branch from extractScheduler.partWritesToMemory() - Add ✦ dreaming indicator to Footer.tsx with i18n; fix Footer.test.tsx mocks - Refactor MemoryDialog.tsx auto-dream status line to use i18n - Remove save_memory tool (memoryTool.ts/test); clean up webui references - Add extractionPlanner.ts, const.ts and associated tests - Delete stale docs/users/configuration/memory.md and docs/developers/tools/memory.md (content superseded) * refactor(memory): remove all Claude Code references from comments and test names * test(memory): remove empty placeholder test files that cause vitest to fail * fix eslint * fix test in windows * fix test * fix(memory): address critical review findings from PR #3087 - fix(read-file): narrow auto-allow from getMemoryBaseDir() (~/.qwen) to isAutoMemPath(projectRoot) to prevent exposing settings.json / OAuth credentials without user approval (wenshao review) - fix(forget): per-entry deletion instead of whole-file unlink - assign stable per-entry IDs (relativePath:index for multi-entry files) so the model can target individual entries without removing siblings - rewrite file keeping unmatched entries; only unlink when file becomes empty (wenshao review) - fix(entries): round-trip correctness for multi-entry new-format bodies - parseAutoMemoryEntries: plain-text line closes current entry and opens a new one (was silently ignored when current was already set) - renderAutoMemoryBody: emit blank line between adjacent entries so the parser can detect entry boundaries on re-read (wenshao review) - fix(entries): resolve two CodeQL polynomial-regex alerts - indentedMatch: \s{2,}(?:[-]\s+)? → [\t ]{2,}(?:[-][\t ]+)? - topLevelMatch: :\s(.+)$ → :[ \t](\S.)$ (github-advanced-security review) - fix(scan.test): use forward-slash literal for relativePath expectation since listMarkdownFiles() normalises all separators to '/' on all platforms including Windows fix(memory): replace isAutoMemPath startsWith with path.relative() Using path.relative() instead of string startsWith() is more robust across platforms — it correctly handles Windows path-separator differences and avoids potential edge cases where a path prefix match could succeed on non-separator boundaries. Addresses github-actions review item 3 (PR #3087). * feat(telemetry): add auto-memory telemetry instrumentation Add OpenTelemetry logs + metrics for the five auto-memory lifecycle events: extract, dream, recall, forget, and remember. Telemetry layer (packages/core/src/telemetry/): - constants.ts: 5 new event-name constants (qwen-code.memory.{extract,dream,recall,forget,remember}) - types.ts: 5 new event classes with typed constructor params (MemoryExtractEvent, MemoryDreamEvent, MemoryRecallEvent, MemoryForgetEvent, MemoryRememberEvent) - metrics.ts: 8 new OTel instruments (5 Counters + 3 Histograms) with recordMemoryXxx() helpers; registered inside initializeMetrics() - loggers.ts: logMemoryExtract/Dream/Recall/Forget/Remember() — each emits a structured log record and calls its recordXxx() counterpart - index.ts: re-exports all new symbols Instrumentation call-sites: - extractScheduler.ts ManagedAutoMemoryExtractRuntime.runTask(): emits extract event with trigger=auto, completed/failed status, patches_count, touched_topics, and wall-clock duration - dream.ts runManagedAutoMemoryDream(): emits dream event with trigger=auto, updated/noop status, deduped_entries, touched_topics, and duration; covers both agent-planner and mechanical fallback paths - recall.ts resolveRelevantAutoMemoryPromptForQuery(): emits recall event with strategy, docs_scanned/selected, and duration; covers model, heuristic, and none paths - forget.ts forgetManagedAutoMemoryEntries(): emits forget event with removed_entries_count, touched_topics, and selection_strategy (model/heuristic/none) - rememberCommand.ts action(): emits remember event with topic=managed\|legacy at command invocation time (before agent decides the actual memory type) * refactor(telemetry): remove memory forget/remember telemetry events Remove EVENT_MEMORY_FORGET and EVENT_MEMORY_REMEMBER along with all associated infrastructure that is no longer needed: - constants.ts: remove EVENT_MEMORY_FORGET, EVENT_MEMORY_REMEMBER - types.ts: remove MemoryForgetEvent, MemoryRememberEvent classes - metrics.ts: remove MEMORY_FORGET_COUNT, MEMORY_REMEMBER_COUNT constants, memoryForgetCounter, memoryRememberCounter module vars, their initialization in initializeMetrics(), and recordMemoryForgetMetrics(), recordMemoryRememberMetrics() functions - loggers.ts: remove logMemoryForget(), logMemoryRemember() functions and their imports - index.ts: remove all re-exports for the above symbols - memory/forget.ts: remove logMemoryForget call-site and import - cli/rememberCommand.ts: remove logMemoryRemember call-sites and import * change default value * fix forked agent * refactor(background): unify fork primitives into runForkedAgent + cleanup - Merge runForkedQuery into runForkedAgent via TypeScript overloads: with cacheSafeParams → GeminiChat single-turn path (ForkedQueryResult) without cacheSafeParams → AgentHeadless multi-turn path (ForkedAgentResult) - Delete forkedQuery.ts; move its test to background/forkedAgent.cache.test.ts - Remove forkedQuery export from followup/index.ts - Migrate all callers (suggestionGenerator, speculation, btwCommand, client) to import from background/forkedAgent - Add getFastModel() / setFastModel() to Config; expose in CLI config init and ModelDialog / modelCommand - Remove resolveFastModel() from AppContainer — now delegated to config.getFastModel() - Strip Claude Code references from code comments * fix(memory): address wenshao's critical review findings - dream.ts: writeDreamManualRunToMetadata now persists lastDreamSessionId and resets recentSessionIdsSinceDream, preventing auto-dream from firing again in the same session after a manual /dream - config.ts: gate managed auto-memory injection on getManagedAutoMemoryEnabled(); when disabled, previously saved memories are no longer injected into new sessions - rememberCommand.ts: remove legacy save_memory branch (tool was removed); fall back to submit_prompt directing agent to write to QWEN.md instead - BuiltinCommandLoader.ts: only register /dream and /forget when managed auto-memory is enabled, matching the feature's runtime availability - forget.ts: return early in forgetManagedAutoMemoryMatches when matches is empty, avoiding unnecessary directory scaffolding as a side effect * fix test * fix ci test * feat(memory): align extract/dream agents to Claude Code patterns - fix(client): move saveCacheSafeParams before early-return paths so extract agents always have cache params available (fixes extract never triggering in skipNextSpeakerCheck mode) - feat(extract): add read-only shell tool + memory-scoped write permissions; create inline createMemoryScopedAgentConfig() with PermissionManager wrapper (isToolEnabled + evaluate) that allows only read-only shell commands and write/edit within the auto-memory dir - feat(extract): align prompt to Claude Code patterns — manifest block listing existing files, parallel read-then-write strategy, two-step save (memory file then index) - feat(dream): remove mechanical fallback; runManagedAutoMemoryDream is now agent-only and throws without config - feat(dream): align prompt to Claude Code 4-phase structure (Orient/Gather/Consolidate/Prune+Index); add narrow transcript grep, relative→absolute date conversion, stale index pruning, index size cap - fix(permissions): add isToolEnabled() to MemoryScopedPermissionManager to prevent TypeError crash in CoreToolScheduler._schedule - test: update dreamScheduler tests to mock dream.js; replace removed mechanical-dedup test with scheduler infrastructure verification * move doc to design * refactor(memory): unify extract+dream background task management into MemoryBackgroundTaskHub - Add memoryTaskHub.ts: single BackgroundTaskRegistry + BackgroundTaskDrainer shared by all memory background tasks; exposes listExtractTasks() / listDreamTasks() typed query helpers and a unified drain() method - extractScheduler: ManagedAutoMemoryExtractRuntime accepts hub via constructor (defaults to defaultMemoryTaskHub); test factory gets isolated fresh hub - dreamScheduler: same pattern — sessionScanner + hub injection; BackgroundTask- Scheduler initialized from injected hub; test factory gets isolated hub - status.ts: replace two separate getRegistry() calls with defaultMemoryTaskHub typed query methods - Footer.tsx (useDreamRunning): subscribe to shared registry, filter by DREAM_TASK_TYPE so extract tasks do not trigger the dream spinner - index.ts: re-export memoryTaskHub.ts so defaultMemoryTaskHub/DREAM_TASK_TYPE/ EXTRACT_TASK_TYPE are available as top-level package exports * refactor(background): introduce general-purpose BackgroundTaskHub Replace memory-specific MemoryBackgroundTaskHub with a domain-agnostic BackgroundTaskHub in the background/ layer. Any future background task runtime (3rd, 4th, …) plugs in by accepting a hub via constructor injection — no new infrastructure required. Changes: - Add background/taskHub.ts: BackgroundTaskHub (registry + drainer + createScheduler() + listByType(taskType, projectRoot?)) and the globalBackgroundTaskHub singleton. Zero knowledge of any task type. - Delete memory/memoryTaskHub.ts: its narrow listExtractTasks / listDreamTasks helpers are replaced by the generic listByType() call. - Move EXTRACT_TASK_TYPE to extractScheduler.ts (owned by the runtime that defines it); replace 3 hardcoded string literals with the const. - Move DREAM_TASK_TYPE to dreamScheduler.ts; use hub.createScheduler() instead of manually wiring new BackgroundTaskScheduler(reg, drain). - status.ts: globalBackgroundTaskHub.listByType(EXTRACT_TASK_TYPE, ...) - Footer.tsx: globalBackgroundTaskHub.registry (shared, filtered by type) - index.ts: export background/taskHub.js; drop memory/memoryTaskHub.js * test(background): add BackgroundTaskHub unit tests and hub isolation checks - background/taskHub.test.ts (11 tests): - createScheduler(): tasks registered via scheduler appear in hub registry; multiple calls return distinct scheduler instances - listByType(): filters by taskType, filters by projectRoot, returns [] for unknown types, two types co-exist in registry but stay separated - drain(): resolves false on timeout, resolves true when tasks complete, resolves true immediately when no tasks in flight - isolation: tasks in hubA do not appear in hubB - globalBackgroundTaskHub: is a BackgroundTaskHub instance with registry/drainer - extractScheduler.test.ts (+1 test): - factory-created runtimes have isolated registries; tasks in runtimeA are invisible to runtimeB; all tasks carry EXTRACT_TASK_TYPE - dreamScheduler.test.ts (+1 test): - factory-created runtimes have isolated registries; tasks in runtimeA are invisible to runtimeB; all tasks carry DREAM_TASK_TYPE * refactor(memory): consolidate all memory state into MemoryManager Replace BackgroundTaskRegistry/Drainer/Scheduler/Hub helper classes and module-level globals with a single MemoryManager class owned by Config. ## Changes ### New - packages/core/src/memory/manager.ts — MemoryManager with: - scheduleExtract / scheduleDream (inline queuing + deduplication logic) - recall / forget / selectForgetCandidates / forgetMatches - getStatus / drain / appendToUserMemory - subscribe(listener) compatible with useSyncExternalStore - storeWith() atomic record registration (no double-notify) - Distinct skippedReason 'scan_throttled' vs 'min_sessions' for dream - packages/core/src/utils/forkedAgent.ts — pure cache util (moved from background/) - packages/core/src/utils/sideQuery.ts — pure util (moved from auxiliary/) ### Deleted - background/taskRegistry, taskDrainer, taskScheduler, taskHub and all tests - background/forkedAgent (moved to utils/) - auxiliary/sideQuery (moved to utils/) - memory/extractScheduler, dreamScheduler, state and all tests ### Modified - config/config.ts — Config owns MemoryManager instance; getMemoryManager() - core/client.ts — all memory ops via config.getMemoryManager() - core/client.test.ts — mock MemoryManager instead of individual modules - memory/status.ts — accepts MemoryManager param, drops globalBackgroundTaskHub - index.ts — memory exports reduced from 14 modules to 5 (manager/types/paths/store/const) - cli/commands/dreamCommand.ts — via config.getMemoryManager() - cli/commands/forgetCommand.ts — via config.getMemoryManager() - cli/components/Footer.tsx — useSyncExternalStore replacing setInterval polling - cli/components/Footer.test.tsx — add getMemoryManager mock	2026-04-16 20:05:45 +08:00
Reid	07475026f6	fix(cli): remember "Start new chat session" until summary changes (#3308 ) * fix(cli): remember "Start new chat session" until summary changes Persist a project-scoped Welcome Back restart choice keyed to the current PROJECT_SUMMARY fingerprint. This suppresses the Welcome Back dialog after choosing "Start new chat session", while still showing it again after the project summary is updated. * fix conflict	2026-04-16 13:54:14 +08:00
DennisYu07	b5115e731e	feat(hooks): Add HTTP Hook, Function Hook and Async Hook support (#2827 ) * add http/async/function type * fix url error * resolve comment * align cc non blocking error * fix hookRunner for async * fix(hooks): update hook type validation to support http and function types - Change validated hook types from ['command', 'plugin'] to ['command', 'http', 'function'] - Add validation for HTTP hooks requiring url field - Add validation for function hooks requiring callback field - Add comprehensive test coverage for all hook type validations Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com> * fix(hooks): align SSRF protection with Claude Code behavior - Allow 127.0.0.0/8 (loopback) for local dev hooks - Allow localhost hostname for local dev hooks - Allow ::1 (IPv6 loopback) for local dev hooks - Add 100.64.0.0/10 (CGNAT) to blocked ranges (RFC 6598) - Update tests to match Claude Code's ssrfGuard.ts behavior This fixes HTTP hooks failing to connect to local dev servers. Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com> * refactor(hooks): align HTTP hook security with Claude Code behavior - Add CRLF/NUL sanitization for env var interpolation (header injection) - Implement combined abort signal (external signal + timeout) - Upgrade SSRF protection to DNS-level with ssrfGuard - Allow loopback (127.0.0.0/8, ::1) for local dev hooks - Block CGNAT (100.64.0.0/10) and IPv6 private ranges - Increase default HTTP hook timeout to 10 minutes - Fix VS Code hooks schema to support http type - Add url, headers, allowedEnvVars, async, once, statusMessage, shell fields - Note: "function" type is SDK-only (callback cannot be serialized to JSON) * feat(hooks): enhance Function Hook with messages, skillRoot, shell, and matcher support - Add MessagesProvider for automatic conversation history passing to function hooks - Add FunctionHookContext with messages, toolUseID, and signal - Add skillRoot support for skill-scoped session hooks - Add shell parameter support for command hooks (bash/powershell) - Add regex matcher support for hook pattern matching - Add statusMessage to CommandHookConfig - Change default function hook timeout from 60s to 5s - Add comprehensive unit tests for all new features Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com> * add session hook for skill * fix function hook parsing * refactor ui for http hook/async hook/function hook * update doc and add integration test * change telemetryn type and refactor SSRF * fix project level bug --------- Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>	2026-04-16 10:10:33 +08:00
DennisYu07	08d3d6eb6f	feat(acp): add complete hooks support for ACP integration (#3248 ) * complete hooks for acp * resolve comment * reslove test * resolve comment for SessionEnd/SessionStart/PostToolUseFailure/PostToolUse	2026-04-16 09:28:26 +08:00
tanzhenxin	17269fa0e6	chore(release): bump version to 0.14.5 (#3298 ) Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>	2026-04-15 22:43:29 +08:00
tanzhenxin	f6271c61b6	feat(auth): discontinue Qwen OAuth free tier (2026-04-15 cutoff) (#3291 ) * feat(auth): discontinue Qwen OAuth free tier (2026-04-15 cutoff) The Qwen OAuth free tier has reached its end-of-life date. This updates all client-side messaging, blocks new OAuth signups, and guides existing users to alternative providers. * fix(test): add getModelsConfig mock and update QWEN_OAUTH test expectations - Add getModelsConfig() to Config mocks in gemini.test.tsx (3 failures) - Update validateNonInterActiveAuth test to expect exit for QWEN_OAUTH since validateAuthMethod now returns an error for discontinued free tier	2026-04-15 22:30:20 +08:00
Shaojin Wen	e4f7a7f380	fix(core): allow thought-only responses in GeminiChat stream validation (#3251 ) Models using thinking/reasoning modes may emit only thought content without explicit text output. The stream validation previously rejected these as 'empty' responses. Now accepts responses that contain either text content or thought content when a finish reason is present. (cherry picked from commit `a0b13911f4`) Co-authored-by: mingholy.lmh <mingholy.lmh@alibaba-inc.com>	2026-04-15 10:29:57 +08:00
jinye	17d9d8c706	fix(core): respect custom Gemini baseUrl from modelProviders (#3212 ) * fix(core): pass gemini baseUrl via httpOptions --------- Co-authored-by: QwenCode <qwen-coder@alibabacloud.com>	2026-04-15 08:58:37 +08:00
Shaojin Wen	6254b85cb6	fix(core): detect rate-limit errors from streamed SSE frames (#3246 ) DashScope throttling (`Throttling.AllocationQuota`) surfaces as an SSE `event:error` frame mid-stream with `:HTTP_STATUS/429` as a comment and a non-numeric `code` in the payload. The existing detection paths missed it, so subagents failed immediately with `Failed to run subagent: id:1 event:error ...` instead of retrying. - `getErrorStatus`: add a final fallback that parses `HTTP_STATUS/NNN` out of `error.message`, bounded by `\b` and the 100-599 range, so streamed errors where the SDK never sees a real HTTP status can still be classified. - `getErrorCode`: fix three `\|\| null` early-returns that swallowed later fall-through paths when the provider code was non-numeric (`isApiError` top-level, `isApiError` JSON-in-message, and `isStructuredError`). The branches now fall through on non-numeric values so `.status` or the new `HTTP_STATUS/NNN` fallback can recover the real code.	2026-04-14 19:58:26 +08:00
Shaojin Wen	83b394e423	feat(core): implement fork subagent for context sharing (#2936 ) * feat(core): implement fork subagent for context sharing - Make subagent_type optional in AgentTool - Add forkSubagent.ts to build identical tool result prefixes - Run fork processes in the background to preserve UX * fix(core): fix test failures related to root execution and optional subagent_type - Skip pathReader and edit tool permission tests when running as root - Fix agent.test.ts to correctly mock execute call with extraHistory - Remove unused imports in forkSubagent.ts * fix(core): fix fork subagent bugs and add CacheSafeParams integration Bug fixes: - Fix AgentParams.subagent_type type: string -> string? (match schema) - Fix undefined agentType passed to hook system (fallback to subagentConfig.name) - Fix hook continuation missing extraHistory parameter - Fix functionResponse missing id field (match coreToolScheduler pattern) - Fix consecutive user messages in Gemini API (ensure history ends with model) - Fix duplicate task_prompt when directive already in extraHistory - Fix FORK_AGENT.systemPrompt empty string causing createChat to throw - Fix redundant dynamic import of forkSubagent.js (merge into single import) - Fix non-fork agent returning empty string on execution failure - Fix misleading fork child rule referencing non-existent system prompt config - Fix functionResponse.response key from {result:} to {output:} for consistency CacheSafeParams integration: - Retrieve parent's generationConfig via getCacheSafeParams() for cache sharing - Add generationConfigOverride to CreateChatOptions and AgentHeadless.execute() - Add toolsOverride to AgentHeadless.execute() for parent tool declarations - Fork API requests now share byte-identical prefix with parent (DashScope cache hits) - Graceful degradation when CacheSafeParams unavailable (first turn) Docs: - Add Fork Subagent section to sub-agents.md user manual - Add fork-subagent-design.md design document * fix(core): apply subagent tool exclusion to forked agents Fork children were inheriting parent's cached tool declarations directly, bypassing prepareTools() filtering and gaining access to AgentTool and cron tools. Extract EXCLUDED_TOOLS_FOR_SUBAGENTS as a shared constant and apply it to forkToolsOverride. * fix(core): skip env history whenever extraHistory is provided Previously gated on generationConfigOverride, which meant the no-cache fallback path (CacheSafeParams unavailable) still ran getInitialChatHistory and duplicated env bootstrap messages already present in the parent's history. Gate on extraHistory instead so both fork paths skip env init. * fix(core): use explicit skipEnvHistory flag for fork env handling The previous fix gated env-init skipping on the presence of extraHistory, but agent-interactive (arena) also passes extraHistory — its chatHistory is env-stripped by stripStartupContext() and DOES need fresh env init for the child's working directory. Skipping env there broke the interactive path. Replace the implicit gate with an explicit skipEnvHistory option that only fork sets (when extraHistory is present, since fork's history comes from getHistory(true) and already contains env). * fix(core): defend skipEnvHistory gate against empty extraHistory Edge case: when the parent's rawHistory ends with a user message and has length 1, extraHistory becomes []. The previous gate (extraHistory !== undefined) would set skipEnvHistory: true, leaving the fork with neither env bootstrap nor parent history. Check length > 0 so empty arrays fall through to the normal env-init path. * fix(core): apply skipEnvHistory to stop-hook retry execute The second subagent.execute() call in the SubagentStop retry loop was missing skipEnvHistory, so on retry the fork's env context would be duplicated — same bug as the initial tanzhenxin report, just on a less common code path.	2026-04-14 14:27:38 +08:00
pomelo	e90abf4c35	docs: update quota exceeded alternatives to OpenRouter and Fireworks (#3217 ) * docs: update quota exceeded alternatives to OpenRouter and Fireworks - Update README.md news section to recommend OpenRouter and Fireworks as primary alternatives, with ModelStudio as third option - Update retry.ts quota error message to include OpenRouter and Fireworks URLs for users whose OAuth quota has been exhausted Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com> * fix(test): update retry test assertions to match new quota error message * docs: update free tier quota to 100 req/day with sunset notice and alternatives Update all references to reflect the Qwen OAuth free tier policy change: - 1,000 → 100 requests/day across code, i18n, and docs - Add 2026-04-15 sunset date everywhere - Guide users to OpenRouter, Fireworks AI, or ModelStudio in docs - Remove CHANGELOG.md --------- Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com> Co-authored-by: tanzhenxin <tanzhenxing1987@gmail.com>	2026-04-13 21:45:38 +08:00

1 2 3 4 5 ...

2163 commits