mirror of
https://github.com/QwenLM/qwen-code.git
synced 2026-08-08 08:15:38 +00:00
* fix(autofix): keep a still-red check visible until its head is judged A red check is a persistent STATE, but the scan only counted checks that failed AFTER the watermark. The moment the watermark passed the failure the PR went quiet while still red. Measured on the live fleet: #6451 watermark 10:55 3 reds completed 09:30, 09:30, 09:51 #7357 watermark 09:18 1 red completed 07:59 #7390 watermark 11:27:37 red completed 11:27:37 — a strict `>` hid it the instant it appeared All three sat red for hours while every scan logged "nothing new", and #6451 wrote two consecutive no-ops whose reasoning never mentions the three failures, because they were not in its feedback at all. A currently-red check now counts as feedback until the head it ran against has been evaluated. The address job records that head in its own `autofix-redcheck` marker — carried inside the eval comment, so no ts/acted/round parser changes and the agent still never sees it as feedback — and the scan skips a PR whose recorded head still matches. That bounds this to ONE look per head rather than every scan, which is what keeps a permanently-red PR from being re-selected forever. The head comes from the REMOTE, not local HEAD: after a rejected push the two differ, and recording a sha that never landed would suppress the reds on the head that actually exists. Empty on failure — matches no marker, so the reds stay visible. Two existing count assertions are replaced by the property they stood for: every check selector in the scan carries the address carve-out. * fix(autofix): pair REPORT_HEAD with the steps that emit its marker Review found the assignment had landed in issue-autofix's "Report dry-run / failure" step, which emits no redcheck marker and has no ${PR} in scope — dead code plus a malformed, swallowed API call. Verifying it surfaced a second half the review did not state: review-address's OWN handoff step emits the marker at line 3362 with REPORT_HEAD never assigned in that step, since shell variables do not cross step boundaries. Neither step sets `set -u`, so it expanded empty and the marker recorded no head — fail-open, but the handoff path never recorded one. Deletes the dead assignment, adds the missing one, and rewords the scan log so the two overlapping counts no longer read as a sum. The test now asserts the PAIRING per step block — emits iff defines — rather than counting each kind. Counting was what let this through: both counts were "right". The first fix for it keyed the sets by step NAME, which merged the two identically-named "Report dry-run / failure" steps and still passed with the bug reintroduced; keying by step block catches it. * fix(autofix): note fail-closed asymmetry on empty LIVE_HEAD (#7438) * fix(autofix): close three state-transition gaps in persistent red-check tracking (#7438) - Forward persistent red checks into agent feedback: the scan selects via N_RED_NOW but the prepare renderer only showed checks that failed AFTER the watermark, leaving the agent with an empty Failed checks section. Add a Still-red checks section with the complement filter. - Omit the redcheck marker on sentinel/retry handoffs: a sentinel ts means the agent evaluated nothing, so recording a judged head would suppress the retry the handoff promises. - Record the checked-out head, not the report-time remote head: capture the SHA in prepare before agent mutations and forward it as a step output, so a mid-run branch move cannot stamp an unevaluated head as judged. * fix(autofix): test empty-LIVE_HEAD fail-closed path (#7438) * fix(autofix): discard no-op same-head duplicates in the queued-job stale gate (#7438) Two near-simultaneous scans can both enqueue the same PR with the same watermark. When the first serialized job ends in a no-op, it records a redcheck marker for the head it judged but leaves both the eval timestamp and the round UNCHANGED — so the live-watermark/round revalidation never fires, and the second job re-runs the agent and posts a duplicate report for the same head. Parse the latest live redcheck marker during prepare (mirroring the scan's RED_HEAD parse) and add its head match against CHECKED_OUT_HEAD as a third stale-duplicate signature, reusing the existing "nothing newer" revalidation so newer feedback or a live conflict still keeps the target actionable. --------- Co-authored-by: wenshao <wenshao@example.com> Co-authored-by: Qwen Code <qwen-code@users.noreply.github.com> Co-authored-by: qwen-code-dev-bot <qwen-code-dev-bot@users.noreply.github.com> |
||
|---|---|---|
| .. | ||
| ai-release-notes-workflow.test.js | ||
| build-and-publish-image-workflow.test.js | ||
| check-build-status.test.js | ||
| check-i18n.test.ts | ||
| chrome-extension-package.test.js | ||
| ci-flaky-rerun-workflow.test.js | ||
| ci-flaky-rerun.test.js | ||
| clean-package-build-artifacts.test.js | ||
| cli-entry.test.js | ||
| comment-attachment-guard-workflow.test.js | ||
| dev.test.js | ||
| generate-changelog.test.js | ||
| generate-release-notes.test.js | ||
| get-release-version-python-sdk.test.js | ||
| get-release-version.test.js | ||
| install-script.test.js | ||
| issue-triage-ownership-workflow.test.js | ||
| lint.test.js | ||
| main-ci-failure-issue-workflow.test.js | ||
| no-ak-integration-ci.test.js | ||
| package-assets.test.js | ||
| package-scripts.test.js | ||
| pr-force-push-reminder-workflow.test.js | ||
| qwen-autofix-workflow.test.js | ||
| qwen-fleet-shepherd-workflow.test.js | ||
| qwen-pr-review-workflow.test.js | ||
| qwen-resolve-workflow.test.js | ||
| qwen-triage-workflow.test.js | ||
| release-helpers.test.js | ||
| release-sdk-workflow.test.js | ||
| sandbox-command.test.js | ||
| sdk-node-exporter-stub.test.js | ||
| serve-fast-path-bundle-check.test.js | ||
| start.test.js | ||
| test-setup.ts | ||
| upload-aliyun-oss-assets.test.js | ||
| vitest.config.ts | ||
| workspaces.test.js | ||