qwen-code/scripts/tests
Shaojin Wen 86b4b281ee
fix(autofix): keep a still-red check visible until its head is judged (#7438)
* fix(autofix): keep a still-red check visible until its head is judged

A red check is a persistent STATE, but the scan only counted checks that
failed AFTER the watermark. The moment the watermark passed the failure
the PR went quiet while still red. Measured on the live fleet:

  #6451  watermark 10:55  3 reds completed 09:30, 09:30, 09:51
  #7357  watermark 09:18  1 red  completed 07:59
  #7390  watermark 11:27:37  red completed 11:27:37 — a strict `>` hid it
                             the instant it appeared

All three sat red for hours while every scan logged "nothing new", and
#6451 wrote two consecutive no-ops whose reasoning never mentions the
three failures, because they were not in its feedback at all.

A currently-red check now counts as feedback until the head it ran
against has been evaluated. The address job records that head in its own
`autofix-redcheck` marker — carried inside the eval comment, so no
ts/acted/round parser changes and the agent still never sees it as
feedback — and the scan skips a PR whose recorded head still matches.
That bounds this to ONE look per head rather than every scan, which is
what keeps a permanently-red PR from being re-selected forever.

The head comes from the REMOTE, not local HEAD: after a rejected push
the two differ, and recording a sha that never landed would suppress the
reds on the head that actually exists. Empty on failure — matches no
marker, so the reds stay visible.

Two existing count assertions are replaced by the property they stood
for: every check selector in the scan carries the address carve-out.

* fix(autofix): pair REPORT_HEAD with the steps that emit its marker

Review found the assignment had landed in issue-autofix's "Report dry-run
/ failure" step, which emits no redcheck marker and has no ${PR} in scope
— dead code plus a malformed, swallowed API call. Verifying it surfaced a
second half the review did not state: review-address's OWN handoff step
emits the marker at line 3362 with REPORT_HEAD never assigned in that
step, since shell variables do not cross step boundaries. Neither step
sets `set -u`, so it expanded empty and the marker recorded no head —
fail-open, but the handoff path never recorded one.

Deletes the dead assignment, adds the missing one, and rewords the scan
log so the two overlapping counts no longer read as a sum.

The test now asserts the PAIRING per step block — emits iff defines —
rather than counting each kind. Counting was what let this through: both
counts were "right". The first fix for it keyed the sets by step NAME,
which merged the two identically-named "Report dry-run / failure" steps
and still passed with the bug reintroduced; keying by step block catches
it.

* fix(autofix): note fail-closed asymmetry on empty LIVE_HEAD (#7438)

* fix(autofix): close three state-transition gaps in persistent red-check tracking (#7438)

- Forward persistent red checks into agent feedback: the scan selects
  via N_RED_NOW but the prepare renderer only showed checks that failed
  AFTER the watermark, leaving the agent with an empty Failed checks
  section. Add a Still-red checks section with the complement filter.

- Omit the redcheck marker on sentinel/retry handoffs: a sentinel ts
  means the agent evaluated nothing, so recording a judged head would
  suppress the retry the handoff promises.

- Record the checked-out head, not the report-time remote head: capture
  the SHA in prepare before agent mutations and forward it as a step
  output, so a mid-run branch move cannot stamp an unevaluated head as
  judged.

* fix(autofix): test empty-LIVE_HEAD fail-closed path (#7438)

* fix(autofix): discard no-op same-head duplicates in the queued-job stale gate (#7438)

Two near-simultaneous scans can both enqueue the same PR with the same
watermark. When the first serialized job ends in a no-op, it records a
redcheck marker for the head it judged but leaves both the eval timestamp
and the round UNCHANGED — so the live-watermark/round revalidation never
fires, and the second job re-runs the agent and posts a duplicate report
for the same head.

Parse the latest live redcheck marker during prepare (mirroring the scan's
RED_HEAD parse) and add its head match against CHECKED_OUT_HEAD as a third
stale-duplicate signature, reusing the existing "nothing newer" revalidation
so newer feedback or a live conflict still keeps the target actionable.

---------

Co-authored-by: wenshao <wenshao@example.com>
Co-authored-by: Qwen Code <qwen-code@users.noreply.github.com>
Co-authored-by: qwen-code-dev-bot <qwen-code-dev-bot@users.noreply.github.com>
2026-07-22 01:08:23 +00:00
..
ai-release-notes-workflow.test.js ci(release): finalize stable releases asynchronously (#6868) 2026-07-15 00:30:51 +00:00
build-and-publish-image-workflow.test.js ci(autofix): restore sandbox image flow (#6261) 2026-07-03 15:30:58 +00:00
check-build-status.test.js fix(review): report what the transcripts prove; build the roster in one call (#7033) 2026-07-18 00:43:57 +00:00
check-i18n.test.ts fix(cli): localize approval mode UI labels (#6592) 2026-07-11 00:07:03 +00:00
chrome-extension-package.test.js fix(ci): cover release integration regressions (#5994) 2026-06-29 11:54:11 +00:00
ci-flaky-rerun-workflow.test.js fix(ci): stop a slow patrol classifier from killing every flaky rerun (#7358) 2026-07-21 02:33:51 +00:00
ci-flaky-rerun.test.js feat(ci): auto-open a deflake fix issue for confirmed flaky tests (#7231) 2026-07-19 16:49:29 +00:00
clean-package-build-artifacts.test.js test(core): stabilize file history eviction test (#6637) 2026-07-10 06:39:52 +00:00
cli-entry.test.js fix(cli): update npm installs safely in background (#7322) 2026-07-21 06:32:24 +00:00
comment-attachment-guard-workflow.test.js ci: add suspicious comment attachment guard (#6599) 2026-07-10 09:40:54 +00:00
dev.test.js fix(review): report what the transcripts prove; build the roster in one call (#7033) 2026-07-18 00:43:57 +00:00
generate-changelog.test.js feat(release): generate AI-assisted release notes (#6756) 2026-07-12 13:00:22 +00:00
generate-release-notes.test.js feat(release): generate AI-assisted release notes (#6756) 2026-07-12 13:00:22 +00:00
get-release-version-python-sdk.test.js feat(sdk-python): add network timeouts to release version helper (#3833) 2026-05-05 19:25:00 +08:00
get-release-version.test.js fix(scripts): handle missing NPM dist-tags gracefully in release versioning (#6476) (#6481) 2026-07-08 12:21:14 +00:00
install-script.test.js fix(cli): avoid updating active CLI processes (#6874) 2026-07-15 00:33:17 +00:00
issue-triage-ownership-workflow.test.js fix(ci): consolidate issue triage ownership (#7180) 2026-07-19 10:58:40 +00:00
lint.test.js revert: remove local PR verification gate (#7031) 2026-07-16 11:24:38 +00:00
main-ci-failure-issue-workflow.test.js feat(release): generate AI-assisted release notes (#6756) 2026-07-12 13:00:22 +00:00
no-ak-integration-ci.test.js fix(ci): avoid apt on self-hosted Playwright smoke (#6865) 2026-07-14 13:19:40 +00:00
package-assets.test.js fix(release): raise prepared package size limit to 96 MB (#6687) (#6691) 2026-07-11 01:01:59 +00:00
package-scripts.test.js fix(autofix): resolve owning package for nested paths; report verify-failed handoffs as not pushed (#7330) 2026-07-20 14:39:56 +00:00
pr-force-push-reminder-workflow.test.js ci(autofix): run agents on dedicated ECS runners (#6207) 2026-07-03 07:40:07 +00:00
qwen-autofix-workflow.test.js fix(autofix): keep a still-red check visible until its head is judged (#7438) 2026-07-22 01:08:23 +00:00
qwen-fleet-shepherd-workflow.test.js feat(autofix): label-driven takeover and release; fix forced-dispatch green no-op (#7165) 2026-07-19 07:35:13 +00:00
qwen-pr-review-workflow.test.js fix(web-shell): restore scheduled task reference interactions (#7313) 2026-07-21 07:10:26 +00:00
qwen-resolve-workflow.test.js fix(ci): serialise the two workflows that push to a PR head branch (#7392) 2026-07-21 07:40:56 +00:00
qwen-triage-workflow.test.js fix(ci): tell a triage action crash apart from a silent agent (#7418) 2026-07-21 11:07:11 +00:00
release-helpers.test.js refactor: extract shared release helper utilities (#3834) 2026-05-05 10:15:17 +08:00
release-sdk-workflow.test.js fix(ci): skip empty SDK release PR (#6861) 2026-07-14 13:19:42 +00:00
sandbox-command.test.js fix(scripts): avoid shell injection in sandbox command detection (#6108) 2026-07-01 16:20:40 +08:00
sdk-node-exporter-stub.test.js perf(telemetry): lazy-load the SDK and split OTLP exporter chains by protocol (#7276) 2026-07-21 07:35:30 +00:00
serve-fast-path-bundle-check.test.js perf(telemetry): lazy-load the SDK and split OTLP exporter chains by protocol (#7276) 2026-07-21 07:35:30 +00:00
start.test.js fix(review): report what the transcripts prove; build the roster in one call (#7033) 2026-07-18 00:43:57 +00:00
test-setup.ts feat(installer): add standalone archive installation (#3776) 2026-05-11 13:25:48 +08:00
upload-aliyun-oss-assets.test.js feat(installer): add standalone hosted install and uninstall flow (#3828) 2026-05-21 11:57:10 +08:00
vitest.config.ts refactor(auth): unify provider config in core, simplify /auth as "Connect a Provider" (#4287) 2026-05-20 23:48:52 +08:00
workspaces.test.js feat(desktop): Add desktop app package with Qwen ACP SDK integration (#3778) 2026-06-11 21:57:20 +08:00