qwen-code

mirror of https://github.com/QwenLM/qwen-code.git synced 2026-05-05 23:42:03 +00:00

Author	SHA1	Message	Date
Shaojin Wen	35fe97e0f6	feat(review): expand review pipeline + qwen review CLI subcommands (#3754 ) Some checks are pending Qwen Code CI / Lint (push) Waiting to run Details Qwen Code CI / Test (push) Blocked by required conditions Details Qwen Code CI / Test-1 (push) Blocked by required conditions Details Qwen Code CI / Test-2 (push) Blocked by required conditions Details Qwen Code CI / Test-3 (push) Blocked by required conditions Details Qwen Code CI / Test-4 (push) Blocked by required conditions Details Qwen Code CI / Test-5 (push) Blocked by required conditions Details Qwen Code CI / Test-6 (push) Blocked by required conditions Details Qwen Code CI / Test-7 (push) Blocked by required conditions Details Qwen Code CI / Test-8 (push) Blocked by required conditions Details Qwen Code CI / Post Coverage Comment (push) Blocked by required conditions Details Qwen Code CI / CodeQL (push) Waiting to run Details E2E Tests / E2E Test (Linux) - sandbox:docker (push) Waiting to run Details E2E Tests / E2E Test (Linux) - sandbox:none (push) Waiting to run Details E2E Tests / E2E Test - macOS (push) Waiting to run Details * feat(review): expand review pipeline + add `qwen review` CLI subcommands Review skill (SKILL.md) changes: - Step 4: 5 → 9 parallel agents (split Correctness/Security, add Test Coverage, 3 undirected personas: attacker / 3am-oncall / maintainer) - Step 5: verification "uncertain → reject" → "uncertain → low-confidence" (terminal-only "Needs Human Review" bucket; never posted as PR comments) - Step 6: single reverse audit → iterative (terminate on no-new-findings, hard cap 3 rounds) - Step 9: self-PR detection (downgrade APPROVE/REQUEST_CHANGES → COMMENT when GitHub forbids self-review with HTTP 422); CI status check (downgrade APPROVE → COMMENT on red/pending CI); existing-Qwen-comment classification with priority order Stale > Resolved > Overlap > NoConflict (only Overlap blocks for confirmation) `qwen review` CLI subcommands (packages/cli/src/commands/review/): - fetch-pr — clean stale + fetch PR ref + create worktree + metadata - pr-context — emit Markdown context file with security preamble + already-discussed dedup section - load-rules — read review rules from base branch (4 source files) - deterministic— run tsc, eslint, ruff, cargo-clippy, go-vet, golangci-lint on changed files; filtered + structured findings JSON (TypeScript/JavaScript, Python, Rust, Go) - presubmit — self-PR + CI status + existing-comment classification in a single JSON report - cleanup — worktree + branch ref + per-target temp files (idempotent) Cross-platform: execFileSync (no shell), path.join, CRLF normalization, which/where for tool detection. Replaces bash-style inline commands in SKILL.md; works identically on macOS/Linux/Windows. Path consistency: SKILL.md temp files moved from /tmp/qwen-review-* to .qwen/tmp/qwen-review-* — matches what os.tmpdir() resolves to across platforms (macOS returns /var/folders/... not /tmp). DESIGN.md gains five "Why ..." sections explaining each design decision; docs/users/features/code-review.md synced for user-visible changes. * feat(review): expose full reply chains in pr-context output `qwen review pr-context` now renders each replied-to inline-comment thread as the original reviewer comment + chronological reply chain, instead of only listing the root-comment snippet. This lets review agents see at a glance whether a topic has been addressed (e.g. a "Fixed in <commit>" reply closes the thread) and avoids re-reporting already-resolved concerns without forcing the LLM driver to manually summarise each reply chain in agent prompts. - Walk `in_reply_to_id` chain to group replies under their root comment - Sort replies chronologically (by id, monotonic on GitHub) - Render thread block: root snippet as a quote + bulleted reply list - Sort threads by `(path, line)` for deterministic output - SKILL.md note updated to point agents at the new chain format * feat(review): include review-level summaries in pr-context output `qwen review pr-context` now also fetches `gh api repos/{owner}/{repo}/pulls/{n}/reviews` and renders a "Review summaries" section listing each reviewer's overall body (the comment they typed alongside an APPROVED / CHANGES_REQUESTED / COMMENTED submission). Closes a real gap found during the PR #3684 review: > "@wenshao [CHANGES_REQUESTED]: The previously identified exported > type rename issue no longer maps to the current PR diff, so this > review only includes the remaining high-confidence blocker." Without this section, the LLM driver's review agents would have missed that integration note from the prior reviewer. - New `RawReview` type + extra `ghApi` call - Filter: skip empty bodies + the canonical "No issues found. LGTM!" template the qwen-review pipeline auto-emits — those carry no agent-actionable content beyond the review state itself - Sort meaningful reviews by `submitted_at` for chronological output - Stdout summary now reports `M/N review summaries` (M = kept after filter) Smoke-tested on PR #3684: 30 inline, 3 issue, 1/30 review summaries correctly surfaces the @wenshao CHANGES_REQUESTED body and filters the 29 LGTM templates. * fix(review): paginate gh API calls to capture comments past page 1 `gh api <path>` defaults to per_page=30. Busy PRs cross that limit on inline comments, issue comments, and reviews — the latest entries (the ones most likely to contain new reviewer feedback or in-flight reply chains) end up on page 2+ and were silently truncated. Concrete bug found while re-reviewing PR #3684: Before: `30 inline, 3 issue comments, 1/30 review summaries` After: `97 inline, 3 issue comments, 6/67 review summaries` 5 additional reviewer-level summaries surfaced — including the @wenshao 2026-04-30 "Multi-agent re-review (Phase C)" body with the explicit verification notes that this PR's pipeline is supposed to chain forward into the next review. Changes: - `lib/gh.ts`: new `ghApiAll(path)` helper using `gh api --paginate`, which walks every `next` link and concatenates each page's array. - `pr-context.ts`: 3 fetches (inline / issue / reviews) → `ghApiAll`. - `presubmit.ts`: PR comments fetch → `ghApiAll` too (existing-comment classification was equally susceptible to dropping page 2+ overlap candidates). `check-runs` and `commits/<sha>/status` calls retain `ghApi` — those return objects (with embedded arrays) and rarely cross 30 entries. --------- Co-authored-by: wenshao <wenshao@U-K7F6PQY3-2157.local>	2026-05-01 18:30:35 +08:00
wenshao	820191d99d	docs: add zero-findings PR tip to follow-up table User doc and PR description now include the "PR review, zero findings → post comments → approve PR" row in the follow-up actions table. Also fixed PR description: "Step 4" → "Step 9" for post comments. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>	2026-04-07 15:33:03 +08:00
wenshao	a9038a1769	fix(review): 5 issues — CI security, incremental+comment, doc accuracy 1. CI config auto-discovery: read from base branch for PR reviews (PR branch is untrusted, malicious PR could inject commands) 2. Incremental early-exit: don't block --comment on unchanged PR — allow posting comments from previous review findings 3. Doc: review summary not always posted (Comment verdict skips it) 4. Doc: cross-repo reviews skip report persistence 5. Doc: clarify "Agents 1-4 findings verified" (not all — reverse audit findings skip verification) Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>	2026-04-07 14:40:16 +08:00
wenshao	c826c24f6e	fix(review): skip Agent 5 in cross-repo mode, update token counts Cross-repo lightweight mode has no local codebase — Agent 5 (build/test) is pointless. Now launches 4 agents instead of 5 in cross-repo mode. Updated token count tables in SKILL.md, user doc, and DESIGN.md: same-repo = 7 LLM calls, cross-repo = 6. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>	2026-04-07 09:24:59 +08:00
wenshao	8f723575bd	docs: add CI config auto-discovery to user doc and design doc - User doc: added "Other" row to language table + explanation that CI config is read for unrecognized projects - DESIGN.md: added "Why auto-discover from CI config" decision section + added .qwen/review-tools.md to rejected alternatives Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>	2026-04-07 07:47:07 +08:00
wenshao	cbf2aa06ac	feat(review): add Java and C/C++ support for deterministic analysis Step 3 now supports: - Java: mvn compile, checkstyle, spotbugs, pmd (Maven); gradle compileJava, checkstyleMain (Gradle) - C/C++: clang-tidy (when compile_commands.json available) Agent 5 build/test precedence now includes Maven and Gradle before Makefile to avoid duplicate builds. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>	2026-04-07 07:19:56 +08:00
wenshao	720796e699	fix(review): safe branch cleanup + cross-repo agent count - git branch -D: add 2>/dev/null \|\| true to both cleanup sites (Step 1 stale cleanup + Step 11) to prevent abort if ref missing - Cross-repo doc: clarify Agents 1-4 only (Agent 5 build/test requires local codebase, not available in cross-repo mode) Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>	2026-04-07 05:02:13 +08:00
wenshao	559f2efd42	fix(review): fix cross-repo mode and add documentation SKILL.md: - Step 9 must use owner/repo from URL (not gh repo view) for cross-repo - Step 2 (project rules) skipped in cross-repo mode (no local files) User doc: add Cross-repo PR Review section with same-repo vs cross-repo capability comparison table. DESIGN.md: add "Why cross-repo uses lightweight mode" section explaining CLI tools are inherently repo-local and our approach is best available. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>	2026-04-07 04:18:55 +08:00
wenshao	01c105b5e7	refactor(review): renumber steps from sub-steps to sequential 1-11 Replace confusing sub-step numbering (1, 1.1, 1.5, 2, 2.5, 2.6, 3, 3.5, 4, 4.5, 5) with clean sequential numbering (1-11). Mapping: 1→1, 1.1→2, 1.5→3, 2→4, 2.5→5, 2.6→6, 3→7, 3.5→8, 4→9, 4.5→10, 5→11 Updated all cross-references in SKILL.md, user docs, and PR description. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>	2026-04-07 04:12:45 +08:00
wenshao	09c95c44c8	fix(review): address 5 Copilot comments - Add error handling for git fetch and gh pr view failures in Step 1 - Skip worktree cleanup on autofix commit/push failure (preserve uncommitted fixes for manual recovery) - Fix Agent 5 counting: it's 1 of the 5 LLM agents (not a separate zero-cost stage). Remove misleading "zero LLM" annotation and duplicate row from token efficiency table. - Reverse audit skip-verification already implemented (comment #53 was stale) Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>	2026-04-07 03:42:20 +08:00
wenshao	6d7afafe3c	perf(review): skip verification for reverse audit findings Reverse audit agent already has full context (all confirmed findings + entire diff), so its findings don't need a second opinion. This brings the actual LLM call count to 7 (5 review + 1 verify + 1 reverse), matching the documented claim. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>	2026-04-07 03:38:49 +08:00
wenshao	d6b9b35350	fix(review): handle interrupted review cleanup If a previous review was interrupted (Ctrl+C, crash), stale worktree and local ref would block the next review. Now Step 1 checks for and cleans up stale .qwen/tmp/review-pr-<N> worktree and qwen-review/pr-<N> ref before creating new ones. Step 5 also cleans up the local ref alongside the worktree. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>	2026-04-07 03:35:00 +08:00
wenshao	6b28920a07	docs: fix autofix section redundancy and add pre-fix verdict note - Remove duplicate worktree commit+push bullet (lines 107 vs 109) - Add note that PR submission uses pre-fix verdict since remote isn't updated until autofix push completes Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>	2026-04-07 03:31:48 +08:00
wenshao	ce47f64ae6	docs: replace Mermaid with plain-text pipeline diagram Mermaid only renders on GitHub; shows as raw code on Nextra, Docusaurus, VS Code preview, and offline viewing. Plain-text ASCII diagram is universally compatible and includes LLM call cost annotations on each stage. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>	2026-04-07 03:30:34 +08:00
wenshao	1db1dec517	docs: replace step list with Mermaid flowchart in How It Works Visual pipeline diagram showing: - Sequential flow from scope detection to cleanup - 5 parallel agents subgraph - Decision branches for autofix and PR comments - Zero-LLM-cost stages marked - GitHub renders Mermaid natively Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>	2026-04-07 03:29:20 +08:00
wenshao	295e907d25	fix(review): address 4 Copilot comments on worktree and verification - Step 4.5: use absolute paths for reports/cache in worktree mode (relative paths would land in worktree and be deleted) - Step 1: fetch into qwen-review/pr-<N> ref to avoid clobbering existing local branches - Step 2.6: reverse audit findings use batch verification (not one-per-finding), consistent with Step 2.5 - Doc: clarify reverse audit findings are also batch-verified Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>	2026-04-07 03:25:24 +08:00
wenshao	9b9bccd27d	docs: update user doc with token efficiency, fix follow-up table - Add Token Efficiency section showing fixed 7 LLM calls breakdown - Fix follow-up table: "fix these issues" is local-only (worktree cleaned up after PR review) - Update PR description with worktree, batch verification, cross-model review, PR comment dedup, and expanded test plan Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>	2026-04-07 03:22:54 +08:00
wenshao	e65e5bd353	perf(review): replace N verification agents with single batch verification Previously, each finding got its own independent verification agent (N findings = N LLM calls). Now a single verification agent receives all findings at once and verifies them in one pass. Token cost: 6+N variable calls → 7 fixed calls (5 review + 1 verify + 1 reverse audit) Quality: minimal impact — batch verification has fuller cross-finding context Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>	2026-04-07 03:19:23 +08:00
wenshao	08a797cf76	fix(review): address 4 Copilot comments - Add model attribution to no-findings LGTM path - Handle empty string from getModel() with .trim() \|\| 'unknown' - Add tests for {{model}} with args and empty model ID - Fix doc contradiction: PR autofix pushes automatically from worktree Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>	2026-04-07 03:13:23 +08:00
wenshao	a5cc2c38cb	fix(review): fix 5 worktree issues found in audit 1. Remove gh pr checkout --detach (modifies working tree, defeats worktree purpose). Use git fetch only. 2. Add dependency installation step (npm ci etc.) in worktree — without it, all TS/JS linting/building fails. 3. Cache and reports written to main project dir, not worktree (would be deleted in Step 5). 4. "fix these issues" tip only for local reviews — worktree is cleaned up after PR review, so interactive fixing not possible. 5. Autofix push uses explicit remote branch name from Step 1. 6. Move incremental check before dependency install to avoid wasting time when no new changes. 7. Fix Step 3 reference: "from Steps 2.5 and 2.6" (includes reverse audit findings). Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>	2026-04-07 03:11:16 +08:00
wenshao	dd2de17de5	feat(review): use ephemeral worktree for PR reviews Replace the stash + checkout + restore flow with an isolated git worktree for PR reviews. This eliminates: - Stash orphan risks (multiple early exit paths) - Wrong-branch risks (Step 5 restore failures) - Build cache pollution (worktree has its own state) - All stash-related error handling complexity New flow: - Step 1: git worktree add .qwen/tmp/review-pr-<number> - All agents operate in the worktree directory - Autofix commits and pushes from the worktree - Step 5: git worktree remove (--force for dirty worktrees) User's working tree is never modified during PR reviews. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>	2026-04-07 03:07:14 +08:00
wenshao	5effbb696f	feat(review): read existing PR comments to avoid duplicate feedback For PR reviews, fetch existing inline and general comments via gh api before launching agents. A summary of already-discussed issues is passed to agents so they don't re-report problems that humans or other tools have already flagged. Added to Exclusion Criteria: "Issues already discussed in existing PR comments." Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>	2026-04-07 01:07:35 +08:00
wenshao	757bd9865a	feat(review): model-aware incremental cache for cross-model review The incremental review cache now stores modelId alongside commitSha. When the same PR is re-reviewed with a different model: - Cache detects model change → runs full review (not skipped) - Informs user: "Previous review used X. Running full review with Y for a second opinion." Same SHA + same model still skips as before. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>	2026-04-07 01:01:52 +08:00
wenshao	50d25733d7	feat(review): add reverse audit step to find coverage gaps Add Step 2.6: after all findings are verified and aggregated, a single reverse audit agent reviews the diff with full knowledge of what was already found, specifically looking for important issues that all previous agents missed. - Only reports Critical/Suggestion level gaps (not Nice to have) - Findings go through the same verification as other agents - Single agent call — minimal cost overhead - If nothing is found, initial review had strong coverage This formalizes the "multi-round undirected audit" pattern that proved effective during the development of this PR (14 rounds, 40+ issues). Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>	2026-04-07 00:56:08 +08:00
wenshao	95a62da039	docs: fix review doc accuracy and remove non-existent /simplify code-review.md: - Add PR URL support to Quick Start - Add "no changes" behavior note - Fix copilot-instructions.md precedence (prefer .github/, not both) - Fix "automatically gitignored" → user must ensure .gitignore coverage - Clarify reports directory is project-relative - Add "What's NOT Flagged" section (exclusion criteria) commands.md: - Replace non-existent /simplify with actual bundled skills (/loop, /qc-helper) Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>	2026-04-07 00:50:26 +08:00
wenshao	ac179b0e02	docs: add /review user documentation Add comprehensive user documentation for the /review command covering: - Quick start examples for all modes (local, PR, file, --comment) - Pipeline overview with all steps explained - Review agents table (5 agents + their focus areas) - Deterministic analysis (supported languages and tools) - Severity levels and PR comment filtering rules - Autofix workflow - PR inline comments (what gets posted vs terminal-only) - Follow-up actions (fix/post comments/commit) - Project review rules (.qwen/review-rules.md etc.) - Incremental review and caching - Review report persistence - Cross-file impact analysis - Design philosophy Also add /review and /simplify to the commands reference page under a new "Built-in Skills" section with link to full docs. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>	2026-04-07 00:47:31 +08:00

26 commits