mirror of
https://github.com/QwenLM/qwen-code.git
synced 2026-08-09 16:56:27 +00:00
* ci(autofix): fan out review targets and stop route-scan starvation
Two throughput fixes for the review loop, both observed live:
- review-scan emitted ONE newest-first target per scan ("single-target
worker"). With sparse cron ticks this starves older armed PRs for hours —
an armed PR sat unprocessed for 16h while newer PRs took every tick. Emit
EVERY eligible target instead: the address matrix's max-parallel (3) bounds
simultaneity and the per-PR concurrency groups already prevent duplicate
same-PR runs, so one surviving scan drains the whole backlog.
- route used a single shared concurrency group with cancel-in-progress. Under
runner backlog a route job sits QUEUED for minutes, and any newer event
(review submissions arrive constantly) cancelled it — five consecutive
dispatched scans died this way; during event storms no full scan survived
at all. Cron ticks keep deduping through a shared 'route-cron' group, but
dispatches and review/issue events now get unique per-run groups: route is
a seconds-long job, so never cancelling it costs nothing and every trigger
is guaranteed to route.
Contract test updated: fan-out asserted (no single-target break, matrix
max-parallel), new route concurrency expression pinned. 50/50.
* ci(autofix): cap targets emitted per scan (review defense-in-depth)
Review note on the fan-out: bound the scan's output for a pathological
backlog. Clarifications recorded in-thread — the loop lives in review-scan
(timeout 15m), not route (5m), and the pre-change worst case already walked
the full candidate list (break fired on the first ELIGIBLE PR, not the first
candidate) — but an explicit bound is good hygiene: emit at most
MAX_TARGETS_PER_SCAN (10) targets, LOG the deferral (never a silent cap),
and let the next scan pick up the remainder since their signals persist.
Contract test pins the cap, the deferral log, and the slice.
* ci(autofix): review round 2 — per-target route coalescing, busy-PR skip, in-loop budget
Both criticals and the suggestion from review, each verified against live
campaign observations:
- Route concurrency is now keyed by TARGET: cron ticks still coalesce with
each other; review events coalesce PER PR (near-simultaneous reviews on one
PR route once — the one useful side effect of the old shared group,
restored — without events on other PRs cancelling this one); issue events
coalesce per issue; dispatches stay unique and are never cancelled. This
keeps the starvation fix while closing the duplicate-forced-scan window the
per-run_id grouping had opened.
- The scan now skips any PR whose review-address job is RUNNING OR QUEUED in
a live autofix run (one runs-list plus a jobs-view per live run). A
fanned-out matrix holds queued jobs past a 10-minute tick and
schedule/dispatch runs never surface in the PR's checks, so without this
the next scan re-emitted the same PRs and per-PR groups accumulated
duplicates that later replayed stale watermarks — the exact duplicate-round
behavior observed live on the fleet.
- The per-scan target budget now BREAKS the candidate loop instead of slicing
after it, so it genuinely bounds scan runtime and API usage (each candidate
costs several serial reads); the deferral is logged and the remainder keeps
its signals for the next scan.
Contract test updated for all three (route expression per target, busy-skip
message + capture regex, in-loop budget break). 50/50.
* ci(autofix): discard stale duplicate targets via live-watermark revalidation
Review: the busy-set closes the queued-matrix window but not the pre-matrix
one — two near-simultaneous same-PR triggers can both scan before either has
emitted a matrix job, so both emit the PR with the same stale watermark, and
the per-PR address group QUEUES (not discards) the duplicate.
That queueing is exactly what makes revalidation sound: address jobs for one
PR run strictly one at a time, so when the duplicate reaches prepare, the
first job's eval marker is already posted. Prepare now recomputes the
watermark from LIVE markers; if it advanced past the matrix watermark and
nothing (reviews / inline / issue comments / failed checks) is newer — and
there is no conflict — the run marks itself stale and the address + verify
steps are skipped entirely: no agent run, no marker, no comment, no push.
Contract test pins the revalidation, both step gates, and the now-three
shared address-carve-out sites. 50/50.
* ci(autofix): filter live runs server-side in the busy-set listing
A client-side status filter over the 15 newest runs loses a long-lived
fanned-out run once cron traffic (~6 runs/hr) pushes it past the
window — its queued review-address PRs silently stop looking busy and
the next scan re-emits them with a stale watermark. Query in_progress
and queued server-side instead, so the limit applies to LIVE runs only
(at most a handful) and the window cannot be starved by completed runs.
One status query failing does not hide the other (|| true per query);
an empty set stays fail-open by design — the address-side live-marker
revalidation is the second line of defense.
* ci(autofix): document route-group cases and busy-set fail-open contract
Review notes (non-blocking, adopted): a four-case summary above the
chained route-concurrency ternary for the next reader hitting it in a
blame, and an explicit contract at the busy-set listing — a double
status-query failure is deliberately fail-open because the skip is an
optimization and the address-side live-marker revalidation is the
correctness gate; if that revalidation is ever removed, this read must
become fail-closed.
* ci(autofix): discard conflict-only duplicates; fix dead recount in the stale gate
Review round: the ts-only revalidation missed a conflict-only
duplicate. Two overlapping scans emit the same conflicted PR with
watermark W; the first serialized job resolves the conflict and — with
no newer feedback — its marker keeps ts=W while its round advances.
The second job then sees CONFLICT=false live but LIVE_EVAL_WM == W, so
the strict > gate never fired and the agent re-ran against resolved
work. The gate now also extracts LIVE_MAX_ROUND and treats
same-ts-with-newer-round (conflict cleared) as a duplicate signature,
still subject to the nothing-newer recount.
The behavioral replay the reviewer asked for immediately caught a
latent bug in the previous fix: the recount jq opened with '((' and
never closed it, so it failed to compile, LIVE_NEW stayed empty, and
the whole stale gate was dead code in production. Fixed to a single
paren; the replay now proves five transitions (conflict-only duplicate
discards, first conflict job proceeds, live conflict always proceeds,
ts-advanced duplicate discards, round-advanced-with-new-feedback
proceeds).
---------
Co-authored-by: wenshao <wenshao@example.com>
|
||
|---|---|---|
| .. | ||
| ai-release-notes-workflow.test.js | ||
| build-and-publish-image-workflow.test.js | ||
| check-build-status.test.js | ||
| check-i18n.test.ts | ||
| chrome-extension-package.test.js | ||
| ci-flaky-rerun-workflow.test.js | ||
| ci-flaky-rerun.test.js | ||
| clean-package-build-artifacts.test.js | ||
| cli-entry.test.js | ||
| comment-attachment-guard-workflow.test.js | ||
| dev.test.js | ||
| generate-changelog.test.js | ||
| generate-release-notes.test.js | ||
| get-release-version-python-sdk.test.js | ||
| get-release-version.test.js | ||
| install-script.test.js | ||
| lint.test.js | ||
| main-ci-failure-issue-workflow.test.js | ||
| no-ak-integration-ci.test.js | ||
| package-assets.test.js | ||
| package-scripts.test.js | ||
| pr-force-push-reminder-workflow.test.js | ||
| qwen-autofix-workflow.test.js | ||
| qwen-resolve-workflow.test.js | ||
| qwen-triage-workflow.test.js | ||
| release-helpers.test.js | ||
| release-sdk-workflow.test.js | ||
| sandbox-command.test.js | ||
| serve-fast-path-bundle-check.test.js | ||
| start.test.js | ||
| test-setup.ts | ||
| upload-aliyun-oss-assets.test.js | ||
| vitest.config.ts | ||
| workspaces.test.js | ||