qwen-code/scripts/tests
Shaojin Wen f9bc8cb250
fix(autofix): re-anchor growth divergence on measurement time and external head moves (#9192)
* fix(autofix): re-anchor growth divergence on measurement time and external head moves

Tightens the growth-divergence comparability window (PR #9104 follow-up,
tracked as #9114):

- measured_at (R2-6): the growth-now marker now carries the prepare-time
  measurement instant, and the divergence read filters on it instead of
  the comment's created_at. The report posts the marker only after the
  agent's ~120-minute run, so a round in flight when a concurrent base
  update landed would otherwise pass a created_at filter while carrying
  sums measured against the old base.

- external head move (R2-8, subsumes R6-3): prior sums are measured
  against origin/main, so any commit an external actor (author push) or a
  stale-base merge added since the bot last evaluated the branch inflates
  this round's sum relative to them. BASE_UPD_AT only tracks the bot's own
  update-branch merge; the new GROWTH_NOW_CUTOFF also re-anchors (drops all
  prior sums) whenever the checked-out head is not the bot's last judged
  head (LIVE_RED_HEAD), covering author pushes and base updates alike.

The reader now dedups/orders per run by measured= (a re-run's fresh
measurement wins). Contract tests cover the measured-based cutoff, the
external-head-move re-anchor (both branches), and the writer→reader
round-trip with the new field. 172/172.

R6-6 (markers don't store the effective budget, so a mid-window budget
raise counts old rounds against the new regime — fail-safe, one round
early) stays tracked in #9114.

* fix(autofix): drop the head-move re-anchor, keep the measurement-time filter

Review found the head-move half of this change broken in three ways
(all probe-verified), so it is withdrawn and returned to #9114 rather
than patched under review:

- R1-1 (regression): `autofix-redcheck` records the head the agent was
  GIVEN, frozen before its push — so after any pushing round the next
  round's head differs and the cutoff was set to now, dropping every
  prior sum. In the push regime OVER_ROUNDS_PRIOR could never reach the
  threshold and the #9104 handoff would never fire at all.
- R1-2: the cut was stateless — the round after a correct re-anchor fell
  back to an empty cutoff and re-admitted every pre-move sum.
- R1-3: with no redcheck marker (a crash round) the `-n` guard skipped
  re-anchoring across a genuine external move.

A correct version needs both a bot-authored-move test and a PERSISTED
cut; that is its own change.

What remains is the measurement-time filter (R2-6), which stands on its
own: the marker carries the prepare-time instant and the divergence read
filters/orders on it instead of the comment's post-agent created_at.

Also from this review:
- R1-4: `measured=` is OPTIONAL in the scan, falling back to the
  comment's created_at, so deploying does not blank an in-flight
  window's census.
- R1-9: the per-run collapse now runs BEFORE the over/window/cutoff
  filters — a re-run whose fresh attempt came back under budget was
  still represented by its stale over=true attempt.
- R1-7: comments corrected — run= is the DEDUP identity, measured= the
  ORDER key (four sites).
- R1-8: recorded as a known residual next to the sibling growth-base
  reader, which still filters on created_at; tracked in #9114.
- R1-5/R1-6: fixtures decouple created_at from measured=, cover a
  legacy marker (with and without the cutoff), and pin the measured_at
  source line in prepare.

* fix(autofix): keep the failure-path growth marker scannable when prepare never ran

* fix(autofix): prefer explicit measured= over created_at fallback in the per-run growth collapse

* test(autofix): pin the explicit-measured preference in the per-run collapse

cbb7186fb6 made the collapse prefer a marker with an explicit measured=
over one falling back to created_at, but nothing pinned it: dropping
`.explicit` from the max_by key shipped green. The failure path posts an
inert over=false marker with no measured= when prepare never ran, and its
fallback timestamp is ~2h later than the prepare-time measured= of the
same run's real attempt — so without the preference the collapse keeps
the inert marker and the over filter then drops the run entirely, losing
the count (#9192 R3-1).

Mutation-checked: reverting to max_by(.measured) turns this fixture from
'false 1' to 'false 0'.

* fix(autofix): gate the measured= stamp on a real measurement (#9192)

An unmeasured re-run attempt still emitted measured_at, so its report
posted an explicit over=false marker that the per-run collapse preferred
over the same run's real over=true measurement — erasing the count the
explicit-measured preference was added to protect. Emit the stamp only
when NET_MEASURED=true and let the push/no-op writers omit measured=
like the failure path already does.

Also pin the strict > cutoff boundary with an equality fixture (the
>= mutant shipped green), drop the per-run-collapse fixture that
duplicated the strictly-stronger one from cbb7186fb6, and document the
three self-limiting deploy-transition residuals: the legacy-straddle
collapse, the two-clock PREV_SUM skew, and the fetch->stamp base-update
window.

---------

Co-authored-by: qwen-code-dev-bot <qwen-code-dev-bot@users.noreply.github.com>
2026-08-16 03:26:58 +00:00
..
ai-release-notes-workflow.test.js fix(release): keep notes anchored and cap the release body (#8199) 2026-07-31 09:55:38 +00:00
audit-runtime-critical.test.js ci: run Windows merge queue tests on ECS (#8386) 2026-08-05 12:14:42 +00:00
build-and-publish-image-workflow.test.js ci(autofix): restore sandbox image flow (#6261) 2026-07-03 15:30:58 +00:00
capture-tmux-ci.test.js ci: install tmux and zip tooling on the Linux test lane, and pin it (#8792) 2026-08-11 05:47:39 +00:00
check-build-status.test.js fix(review): report what the transcripts prove; build the roster in one call (#7033) 2026-07-18 00:43:57 +00:00
check-i18n.test.ts fix(cli): localize approval mode UI labels (#6592) 2026-07-11 00:07:03 +00:00
check-voice-guard-sync.test.js feat(voice): support trusted private ASR base URLs (#8350) 2026-08-06 14:04:57 +00:00
chrome-extension-package.test.js fix(ci): cover release integration regressions (#5994) 2026-06-29 11:54:11 +00:00
ci-flaky-rerun-workflow.test.js fix(ci): stop a slow patrol classifier from killing every flaky rerun (#7358) 2026-07-21 02:33:51 +00:00
ci-flaky-rerun.test.js feat(ci): auto-open a deflake fix issue for confirmed flaky tests (#7231) 2026-07-19 16:49:29 +00:00
clean-package-build-artifacts.test.js test(core): stabilize file history eviction test (#6637) 2026-07-10 06:39:52 +00:00
cli-entry.test.js fix(cli): preserve Qwen Review startup version in footers (#8431) 2026-08-04 14:58:56 +00:00
comment-attachment-guard-workflow.test.js ci: route trusted-author fork PRs and no-checkout jobs to the ECS pool (#8502) 2026-08-04 03:48:24 +00:00
desktop-oss-workflow.test.js fix(desktop): harden release pipeline (#9009) 2026-08-12 16:38:12 +00:00
dev.test.js fix(review): report what the transcripts prove; build the roster in one call (#7033) 2026-07-18 00:43:57 +00:00
e2e-workflow.test.js fix(ci): keep the post-merge E2E signal on main alive (#7795) 2026-07-28 11:54:52 +00:00
generate-changelog.test.js feat(release): generate AI-assisted release notes (#6756) 2026-07-12 13:00:22 +00:00
generate-release-notes.test.js ci: run Windows merge queue tests on ECS (#8386) 2026-08-05 12:14:42 +00:00
get-release-version-python-sdk.test.js
get-release-version.test.js fix(release): bump preview base past published stable (#7978) 2026-07-29 23:21:00 +00:00
install-script.test.js fix(install): avoid Get-FileHash for Windows checksums (#9112) 2026-08-14 01:12:08 +00:00
integration-vitest-config.test.ts fix(tests): apply integration worker limits to forks (#8689) 2026-08-08 00:53:33 +00:00
issue-triage-ownership-workflow.test.js ci: remove broken legacy scheduled PR triage workflow (#8434) 2026-08-03 10:20:42 +00:00
lint.test.js fix(ci): cache downloaded linters on ECS runners (#9001) 2026-08-13 05:13:23 +00:00
live-host-oss-workflow.test.js fix(ci): restore Live Host release mirroring (#8917) 2026-08-11 07:08:26 +00:00
main-ci-failure-issue-workflow.test.js fix(ci): keep the post-merge E2E signal on main alive (#7795) 2026-07-28 11:54:52 +00:00
no-ak-integration-ci.test.js fix(ci): reduce ENOSPC and load-sensitive test flakes (#8982) 2026-08-13 16:44:39 +00:00
package-assets.test.js fix(review): harden the pipeline against four live-run failures (#9086) 2026-08-14 04:38:49 +00:00
package-scripts.test.js fix(tests): apply integration worker limits to forks (#8689) 2026-08-08 00:53:33 +00:00
pr-force-push-reminder-workflow.test.js ci(autofix): run agents on dedicated ECS runners (#6207) 2026-07-03 07:40:07 +00:00
pr-self-report-label.test.js feat(autofix): escalate stopped takeover PRs and age out unanswered pauses (#8960) 2026-08-15 17:32:23 +00:00
qwen-autofix-fork-bridge-workflow.test.js feat(autofix): bridge fork-PR reviews into the credentialed review lane (#8676) 2026-08-07 16:11:48 +00:00
qwen-autofix-workflow.test.js fix(autofix): re-anchor growth divergence on measurement time and external head moves (#9192) 2026-08-16 03:26:58 +00:00
qwen-fleet-shepherd-workflow.test.js feat(autofix): escalate stopped takeover PRs and age out unanswered pauses (#8960) 2026-08-15 17:32:23 +00:00
qwen-pr-review-workflow.test.js fix(ci): keep no-op review requests out of the PR review concurrency group (#9210) 2026-08-15 16:42:57 +00:00
qwen-repo-hygiene-workflow.test.js fix(ci): route workflow label mutations through REST (#8761) 2026-08-09 15:05:15 +00:00
qwen-resolve-workflow.test.js fix(ci): keep the review workflow under the expression-length limit (#8720) 2026-08-08 05:53:35 +00:00
qwen-triage-finalize-workflow.test.js fix(ci): rename triage status marker to avoid duplicate-guard collision (#7723) 2026-07-26 15:51:22 +00:00
qwen-triage-workflow.test.js perf(ci): make the triage budget operator-tunable and raise it (#8810) 2026-08-10 12:40:21 +00:00
release-helpers.test.js
release-sdk-workflow.test.js fix(ci): skip empty SDK release PR (#6861) 2026-07-14 13:19:42 +00:00
release-workflow.test.js feat(review): say so when the bundle is older than the review it runs (#8390) 2026-08-07 03:21:26 +00:00
review-source-digest.test.ts feat(review): say so when the bundle is older than the review it runs (#8390) 2026-08-07 03:21:26 +00:00
review-worktree-cleanup-workflow.test.js fix(ci): clean review worktrees after cancellation (#8474) 2026-08-05 02:39:36 +00:00
sandbox-command.test.js fix(scripts): avoid shell injection in sandbox command detection (#6108) 2026-07-01 16:20:40 +08:00
sdk-java-workflow.test.js ci: route trusted-author fork PRs and no-checkout jobs to the ECS pool (#8502) 2026-08-04 03:48:24 +00:00
sdk-node-exporter-stub.test.js perf(telemetry): lazy-load the SDK and split OTLP exporter chains by protocol (#7276) 2026-07-21 07:35:30 +00:00
security-workflows.test.js chore(ci): Add security hygiene: CODEOWNERS for release workflows, least-privilege permissions, security checks and Scorecard (#9008) 2026-08-14 01:22:53 +00:00
serve-fast-path-bundle-check.test.js feat(ci): fail the startup bundle check when the CLI entry is hoisted into a chunk (#8203) 2026-07-31 08:57:57 +00:00
start.test.js fix(review): report what the transcripts prove; build the roster in one call (#7033) 2026-07-18 00:43:57 +00:00
test-setup.ts
update-ecs-runner-qwen-workflow.test.js fix(ci): reconcile ECS runner updater on workflow changes (#8373) 2026-08-02 09:25:17 +00:00
upload-aliyun-oss-assets.test.js feat(installer): add standalone hosted install and uninstall flow (#3828) 2026-05-21 11:57:10 +08:00
verify-capture.test.js fix(ci): avoid verify capture color conflict (#8236) 2026-07-31 14:15:40 +00:00
vitest.config.ts ci: run Windows merge queue tests on ECS (#8386) 2026-08-05 12:14:42 +00:00
workflow-helpers.js ci: run Windows merge queue tests on ECS (#8386) 2026-08-05 12:14:42 +00:00
workspaces.test.js feat(desktop): Add desktop app package with Qwen ACP SDK integration (#3778) 2026-06-11 21:57:20 +08:00