qwen-code/scripts/tests
Shaojin Wen ee0cc79739
ci(autofix): fan out review targets and stop route-scan starvation (#7127)
* ci(autofix): fan out review targets and stop route-scan starvation

Two throughput fixes for the review loop, both observed live:

- review-scan emitted ONE newest-first target per scan ("single-target
  worker"). With sparse cron ticks this starves older armed PRs for hours —
  an armed PR sat unprocessed for 16h while newer PRs took every tick. Emit
  EVERY eligible target instead: the address matrix's max-parallel (3) bounds
  simultaneity and the per-PR concurrency groups already prevent duplicate
  same-PR runs, so one surviving scan drains the whole backlog.

- route used a single shared concurrency group with cancel-in-progress. Under
  runner backlog a route job sits QUEUED for minutes, and any newer event
  (review submissions arrive constantly) cancelled it — five consecutive
  dispatched scans died this way; during event storms no full scan survived
  at all. Cron ticks keep deduping through a shared 'route-cron' group, but
  dispatches and review/issue events now get unique per-run groups: route is
  a seconds-long job, so never cancelling it costs nothing and every trigger
  is guaranteed to route.

Contract test updated: fan-out asserted (no single-target break, matrix
max-parallel), new route concurrency expression pinned. 50/50.

* ci(autofix): cap targets emitted per scan (review defense-in-depth)

Review note on the fan-out: bound the scan's output for a pathological
backlog. Clarifications recorded in-thread — the loop lives in review-scan
(timeout 15m), not route (5m), and the pre-change worst case already walked
the full candidate list (break fired on the first ELIGIBLE PR, not the first
candidate) — but an explicit bound is good hygiene: emit at most
MAX_TARGETS_PER_SCAN (10) targets, LOG the deferral (never a silent cap),
and let the next scan pick up the remainder since their signals persist.
Contract test pins the cap, the deferral log, and the slice.

* ci(autofix): review round 2 — per-target route coalescing, busy-PR skip, in-loop budget

Both criticals and the suggestion from review, each verified against live
campaign observations:

- Route concurrency is now keyed by TARGET: cron ticks still coalesce with
  each other; review events coalesce PER PR (near-simultaneous reviews on one
  PR route once — the one useful side effect of the old shared group,
  restored — without events on other PRs cancelling this one); issue events
  coalesce per issue; dispatches stay unique and are never cancelled. This
  keeps the starvation fix while closing the duplicate-forced-scan window the
  per-run_id grouping had opened.
- The scan now skips any PR whose review-address job is RUNNING OR QUEUED in
  a live autofix run (one runs-list plus a jobs-view per live run). A
  fanned-out matrix holds queued jobs past a 10-minute tick and
  schedule/dispatch runs never surface in the PR's checks, so without this
  the next scan re-emitted the same PRs and per-PR groups accumulated
  duplicates that later replayed stale watermarks — the exact duplicate-round
  behavior observed live on the fleet.
- The per-scan target budget now BREAKS the candidate loop instead of slicing
  after it, so it genuinely bounds scan runtime and API usage (each candidate
  costs several serial reads); the deferral is logged and the remainder keeps
  its signals for the next scan.

Contract test updated for all three (route expression per target, busy-skip
message + capture regex, in-loop budget break). 50/50.

* ci(autofix): discard stale duplicate targets via live-watermark revalidation

Review: the busy-set closes the queued-matrix window but not the pre-matrix
one — two near-simultaneous same-PR triggers can both scan before either has
emitted a matrix job, so both emit the PR with the same stale watermark, and
the per-PR address group QUEUES (not discards) the duplicate.

That queueing is exactly what makes revalidation sound: address jobs for one
PR run strictly one at a time, so when the duplicate reaches prepare, the
first job's eval marker is already posted. Prepare now recomputes the
watermark from LIVE markers; if it advanced past the matrix watermark and
nothing (reviews / inline / issue comments / failed checks) is newer — and
there is no conflict — the run marks itself stale and the address + verify
steps are skipped entirely: no agent run, no marker, no comment, no push.
Contract test pins the revalidation, both step gates, and the now-three
shared address-carve-out sites. 50/50.

* ci(autofix): filter live runs server-side in the busy-set listing

A client-side status filter over the 15 newest runs loses a long-lived
fanned-out run once cron traffic (~6 runs/hr) pushes it past the
window — its queued review-address PRs silently stop looking busy and
the next scan re-emits them with a stale watermark. Query in_progress
and queued server-side instead, so the limit applies to LIVE runs only
(at most a handful) and the window cannot be starved by completed runs.
One status query failing does not hide the other (|| true per query);
an empty set stays fail-open by design — the address-side live-marker
revalidation is the second line of defense.

* ci(autofix): document route-group cases and busy-set fail-open contract

Review notes (non-blocking, adopted): a four-case summary above the
chained route-concurrency ternary for the next reader hitting it in a
blame, and an explicit contract at the busy-set listing — a double
status-query failure is deliberately fail-open because the skip is an
optimization and the address-side live-marker revalidation is the
correctness gate; if that revalidation is ever removed, this read must
become fail-closed.

* ci(autofix): discard conflict-only duplicates; fix dead recount in the stale gate

Review round: the ts-only revalidation missed a conflict-only
duplicate. Two overlapping scans emit the same conflicted PR with
watermark W; the first serialized job resolves the conflict and — with
no newer feedback — its marker keeps ts=W while its round advances.
The second job then sees CONFLICT=false live but LIVE_EVAL_WM == W, so
the strict > gate never fired and the agent re-ran against resolved
work. The gate now also extracts LIVE_MAX_ROUND and treats
same-ts-with-newer-round (conflict cleared) as a duplicate signature,
still subject to the nothing-newer recount.

The behavioral replay the reviewer asked for immediately caught a
latent bug in the previous fix: the recount jq opened with '((' and
never closed it, so it failed to compile, LIVE_NEW stayed empty, and
the whole stale gate was dead code in production. Fixed to a single
paren; the replay now proves five transitions (conflict-only duplicate
discards, first conflict job proceeds, live conflict always proceeds,
ts-advanced duplicate discards, round-advanced-with-new-feedback
proceeds).

---------

Co-authored-by: wenshao <wenshao@example.com>
2026-07-18 12:06:26 +00:00
..
ai-release-notes-workflow.test.js ci(release): finalize stable releases asynchronously (#6868) 2026-07-15 00:30:51 +00:00
build-and-publish-image-workflow.test.js ci(autofix): restore sandbox image flow (#6261) 2026-07-03 15:30:58 +00:00
check-build-status.test.js fix(review): report what the transcripts prove; build the roster in one call (#7033) 2026-07-18 00:43:57 +00:00
check-i18n.test.ts fix(cli): localize approval mode UI labels (#6592) 2026-07-11 00:07:03 +00:00
chrome-extension-package.test.js fix(ci): cover release integration regressions (#5994) 2026-06-29 11:54:11 +00:00
ci-flaky-rerun-workflow.test.js feat(ci): add automated PR failure patrol (#6766) 2026-07-15 00:51:11 +00:00
ci-flaky-rerun.test.js feat(ci): add automated PR failure patrol (#6766) 2026-07-15 00:51:11 +00:00
clean-package-build-artifacts.test.js test(core): stabilize file history eviction test (#6637) 2026-07-10 06:39:52 +00:00
cli-entry.test.js fix(review): report what the transcripts prove; build the roster in one call (#7033) 2026-07-18 00:43:57 +00:00
comment-attachment-guard-workflow.test.js ci: add suspicious comment attachment guard (#6599) 2026-07-10 09:40:54 +00:00
dev.test.js fix(review): report what the transcripts prove; build the roster in one call (#7033) 2026-07-18 00:43:57 +00:00
generate-changelog.test.js feat(release): generate AI-assisted release notes (#6756) 2026-07-12 13:00:22 +00:00
generate-release-notes.test.js feat(release): generate AI-assisted release notes (#6756) 2026-07-12 13:00:22 +00:00
get-release-version-python-sdk.test.js feat(sdk-python): add network timeouts to release version helper (#3833) 2026-05-05 19:25:00 +08:00
get-release-version.test.js fix(scripts): handle missing NPM dist-tags gracefully in release versioning (#6476) (#6481) 2026-07-08 12:21:14 +00:00
install-script.test.js fix(cli): avoid updating active CLI processes (#6874) 2026-07-15 00:33:17 +00:00
lint.test.js revert: remove local PR verification gate (#7031) 2026-07-16 11:24:38 +00:00
main-ci-failure-issue-workflow.test.js feat(release): generate AI-assisted release notes (#6756) 2026-07-12 13:00:22 +00:00
no-ak-integration-ci.test.js fix(ci): avoid apt on self-hosted Playwright smoke (#6865) 2026-07-14 13:19:40 +00:00
package-assets.test.js fix(release): raise prepared package size limit to 96 MB (#6687) (#6691) 2026-07-11 01:01:59 +00:00
package-scripts.test.js feat(web-shell): git status chip, visual working-tree diff, and sidebar git status (#7054) 2026-07-18 10:06:07 +00:00
pr-force-push-reminder-workflow.test.js ci(autofix): run agents on dedicated ECS runners (#6207) 2026-07-03 07:40:07 +00:00
qwen-autofix-workflow.test.js ci(autofix): fan out review targets and stop route-scan starvation (#7127) 2026-07-18 12:06:26 +00:00
qwen-resolve-workflow.test.js fix(tests): sync qwen-resolve-workflow test expectations with PR #6706 timeout changes (#6720) 2026-07-11 09:27:24 +00:00
qwen-triage-workflow.test.js fix(ci): notify silent triage re-runs (#7079) 2026-07-17 16:07:10 +00:00
release-helpers.test.js refactor: extract shared release helper utilities (#3834) 2026-05-05 10:15:17 +08:00
release-sdk-workflow.test.js fix(ci): skip empty SDK release PR (#6861) 2026-07-14 13:19:42 +00:00
sandbox-command.test.js fix(scripts): avoid shell injection in sandbox command detection (#6108) 2026-07-01 16:20:40 +08:00
serve-fast-path-bundle-check.test.js feat(daemon): Profile ACP channel initialization (#7145) 2026-07-18 05:27:59 +00:00
start.test.js fix(review): report what the transcripts prove; build the roster in one call (#7033) 2026-07-18 00:43:57 +00:00
test-setup.ts feat(installer): add standalone archive installation (#3776) 2026-05-11 13:25:48 +08:00
upload-aliyun-oss-assets.test.js feat(installer): add standalone hosted install and uninstall flow (#3828) 2026-05-21 11:57:10 +08:00
vitest.config.ts refactor(auth): unify provider config in core, simplify /auth as "Connect a Provider" (#4287) 2026-05-20 23:48:52 +08:00
workspaces.test.js feat(desktop): Add desktop app package with Qwen ACP SDK integration (#3778) 2026-06-11 21:57:20 +08:00