qwen-code/scripts/tests
Shaojin Wen afd56117e4
fix(autofix): retry a model API error instead of stranding the PR (#7247)
* fix(autofix): retry a model API error instead of stranding the PR

When the agent's qwen subprocess dies on a model-side [API Error]
(403 access denied, a 429 quota, a 5xx), run-agent.mjs wrote a
handoff/failure.md, so the handoff step treated it as an EVALUATED
handoff — it advanced the watermark and the next scan saw 'nothing
new', stranding the PR until a manual re-arm. But the agent never
actually evaluated the feedback; the model was unreachable.

#7220 hit exactly this: fork-takeover engaged and ran the agent, the
model returned '[API Error: 403 Model access denied]' (the autofix key
lacks access to qwen3.8-max-preview), and the PR was left with an
advanced watermark that will not retry.

Fix, mirroring #7229's no-output-crash handling:
- run-agent.mjs extracts a [API Error: 4xx/5xx] from the captured
  output tail, includes it in failure.md, and drops an
  marker file.
- The handoff step reads that marker and routes the failure to the
  sentinel-ts (retry) path — the watermark does NOT advance, so the
  next scan retries; the round still increments so a PERSISTENT model
  failure is bounded by MAX_ROUNDS. The headline names the model error
  and, on the final attempt, tells the maintainer to check the autofix
  model key/access and re-arm — instead of a generic crash message.

Tests: run-agent.mjs flags a model [API Error] (marker + failure.md)
and does NOT flag a generic failure; the handoff replay treats an
API-error handoff as sentinel|retry (not a watermark advance) with a
model-aware, cause-specific headline. 62/62 + 12/12.

* fix(autofix): scope + broaden the retryable model-API detection (review)

Addresses wenshao's review on #7247:

- Behavioral (1): the agent-api-error marker was written on ANY non-zero
  exit whose output tail contained an API-error string — so a loop
  guard, a timeout, or an agent-written failure.md (a real verdict)
  would wrongly retry and, worst case, silently discard a verdict. The
  write is now scoped to the bare-failure branch and guarded by
  !timedOut, so only an un-evaluated model failure retries.
- Coverage (2): the old regex only matched a LEADING status digit, so
  it missed the canonical rate-limit render, the (Status: …) form, the
  bad-key 401, the Chinese quota text, and the unwrapped Qwen OAuth
  quota — i.e. most real errors this targets. Detection is now a
  whitelist of RECOVERABLE errors (401/402/403/429/5xx + rate-limit /
  quota / api-key / RESOURCE_EXHAUSTED / overloaded phrasings, plus the
  standalone OAuth-quota form); a 400/404 stays terminal.
- Test gap (3): a writer↔reader contract test now runs the REAL
  run-agent.mjs to write the marker, then the extracted workflow reader
  block against that same workdir — a rename on either side (proven
  with the YAML-only mutation) now fails the suite.
- Smaller: API_ERROR_DETAIL is comment-escaped (sed) and capped
  (cut -c1-200) since it derives from agent stdout; the marker match is
  single-line ([^]\n]) so a multi-line render can't smuggle a newline;
  agent-api-error is added to the run-artifacts list.

Non-recoverable 4xx (400/404) deliberately stay terminal; the live
401/403 config cases retry and self-heal once the key/access is fixed.
79/79 across both suites.

* test(autofix): cover the timeout guard and the OAuth-quota fallback (review)

Two coverage gaps from the ci-bot review on #7247:
- The !result.timedOut guard was only asserted indirectly — no test
  emitted an [API Error] AND timed out. Added a case (spawnSync +
  QWEN_TIMEOUT_MS=100): qwen streams [API Error: 503] then hangs past
  the budget → killed → no marker. A refactor to !loopDetected now
  fails here.
- The standalone Qwen-OAuth-quota fallback (unwrapped, no [API Error:])
  had no test. Added a case emitting bare 'Qwen OAuth quota exceeded
  (limit: 100/min)' → marker written, wrapped as
  '[API Error: Qwen OAuth quota exceeded …]'.

* fix(autofix): anchor the API-error code, split retry budget by cause, keep the headline UTF-8

Addresses the review on #7247.

Classifier (points 2 and 4): the status code is now read from its POSITION in
the render (`[API Error: <code>`) instead of matched anywhere in the message.
Matching anywhere retried permanent failures forever — `400 Invalid value for
max_tokens: must be <= 512` matched a bare \b5\d\d\b and `400 context length
exceeded` matched a bare `exceeded`. `exceeded` now only counts as part of
`quota`. A 404 whose message says the model "does not exist or you do not have
access to it" — the OpenAI-compatible render of what a 403 reports — is no
longer terminal.

Retry budget (point 3): the marker now carries the cause class. A transient
429/5xx self-heals and keeps the full round budget; an auth/access error that
only a maintainer can fix is capped at API_AUTH_MAX_ROUNDS (3) and then goes
terminal with the "check the autofix model key/access, then re-arm" headline —
instead of ~100 agent runs and ~100 PR comments over ~17h on a takeover PR.
The terminal round is stamped so the scan's round gate skips the PR while the
sentinel ts keeps the feedback live for a re-arm.

Headline (point 1): `cut -c` counts bytes under GNU coreutils and the
classifier deliberately matches CJK renders, so the 200-byte cap could split a
multi-byte character and emit invalid UTF-8. Guarded with
`iconv -f utf-8 -t utf-8 -c || true`, matching the sibling publish site (the
`|| true` is required — iconv -c exits 1 when it discards).

Minor (point 5): documented that detection is best-effort because apiError is
derived from the last 20 KB of output; `head -1` -> `head -n 1`; tests added
for a permanent 400 carrying a 3-digit number >= 500 and for a >200-byte CJK
render staying valid UTF-8.

* test(autofix): cover the auth-capped retry budget and Chinese API-error patterns (#7247)

* fix(autofix): short-circuit 400 as terminal and classify only the last API error (#7247)

* fix(autofix): treat transport-level API failures as retryable

#7365 stranded at round 2/100 on this render:

    [API Error: terminated (cause: read ECONNRESET)]

The connection to the model dropped mid-run. That is as transient as a 429, but
the classifier never saw it that way: a transport failure never got far enough
to have an HTTP status, so it fell through to the keyword arm, and the keyword
arm only knew about rate limits and quotas. It was classified terminal, the
watermark advanced, and a PR that needed nothing but a re-run was handed to a
human.

Verified against the shipped classifier before the fix — every transport render
came back terminal:

    terminated (cause: read ECONNRESET)   -> terminal
    fetch failed                          -> terminal
    socket hang up                        -> terminal
    connect ETIMEDOUT                     -> terminal

Adds a transport arm to the code-less branch: ECONNRESET, ECONNREFUSED,
ETIMEDOUT, EPIPE, EAI_AGAIN, socket hang up, fetch failed, terminated.

ENOTFOUND is deliberately excluded. A hostname that does not resolve is a
misconfigured endpoint, which repeats forever — the same reasoning that keeps a
bad model name terminal.

Coded errors are unaffected: the arm sits after the status-code branch, so the
400 short-circuit added in 719991a3b still runs first.

* fix(autofix): address review — OAuth fallback override, comment accuracy, display clamp (#7247)

---------

Co-authored-by: wenshao <wenshao@example.com>
Co-authored-by: 易良 <1204183885@qq.com>
Co-authored-by: qwen-code-ci-bot <qwen-code-ci-bot@users.noreply.github.com>
2026-07-21 07:01:56 +00:00
..
ai-release-notes-workflow.test.js ci(release): finalize stable releases asynchronously (#6868) 2026-07-15 00:30:51 +00:00
build-and-publish-image-workflow.test.js ci(autofix): restore sandbox image flow (#6261) 2026-07-03 15:30:58 +00:00
check-build-status.test.js fix(review): report what the transcripts prove; build the roster in one call (#7033) 2026-07-18 00:43:57 +00:00
check-i18n.test.ts fix(cli): localize approval mode UI labels (#6592) 2026-07-11 00:07:03 +00:00
chrome-extension-package.test.js fix(ci): cover release integration regressions (#5994) 2026-06-29 11:54:11 +00:00
ci-flaky-rerun-workflow.test.js fix(ci): stop a slow patrol classifier from killing every flaky rerun (#7358) 2026-07-21 02:33:51 +00:00
ci-flaky-rerun.test.js feat(ci): auto-open a deflake fix issue for confirmed flaky tests (#7231) 2026-07-19 16:49:29 +00:00
clean-package-build-artifacts.test.js test(core): stabilize file history eviction test (#6637) 2026-07-10 06:39:52 +00:00
cli-entry.test.js fix(cli): update npm installs safely in background (#7322) 2026-07-21 06:32:24 +00:00
comment-attachment-guard-workflow.test.js ci: add suspicious comment attachment guard (#6599) 2026-07-10 09:40:54 +00:00
dev.test.js fix(review): report what the transcripts prove; build the roster in one call (#7033) 2026-07-18 00:43:57 +00:00
generate-changelog.test.js feat(release): generate AI-assisted release notes (#6756) 2026-07-12 13:00:22 +00:00
generate-release-notes.test.js feat(release): generate AI-assisted release notes (#6756) 2026-07-12 13:00:22 +00:00
get-release-version-python-sdk.test.js feat(sdk-python): add network timeouts to release version helper (#3833) 2026-05-05 19:25:00 +08:00
get-release-version.test.js fix(scripts): handle missing NPM dist-tags gracefully in release versioning (#6476) (#6481) 2026-07-08 12:21:14 +00:00
install-script.test.js fix(cli): avoid updating active CLI processes (#6874) 2026-07-15 00:33:17 +00:00
issue-triage-ownership-workflow.test.js fix(ci): consolidate issue triage ownership (#7180) 2026-07-19 10:58:40 +00:00
lint.test.js revert: remove local PR verification gate (#7031) 2026-07-16 11:24:38 +00:00
main-ci-failure-issue-workflow.test.js feat(release): generate AI-assisted release notes (#6756) 2026-07-12 13:00:22 +00:00
no-ak-integration-ci.test.js fix(ci): avoid apt on self-hosted Playwright smoke (#6865) 2026-07-14 13:19:40 +00:00
package-assets.test.js fix(release): raise prepared package size limit to 96 MB (#6687) (#6691) 2026-07-11 01:01:59 +00:00
package-scripts.test.js fix(autofix): resolve owning package for nested paths; report verify-failed handoffs as not pushed (#7330) 2026-07-20 14:39:56 +00:00
pr-force-push-reminder-workflow.test.js ci(autofix): run agents on dedicated ECS runners (#6207) 2026-07-03 07:40:07 +00:00
qwen-autofix-workflow.test.js fix(autofix): retry a model API error instead of stranding the PR (#7247) 2026-07-21 07:01:56 +00:00
qwen-fleet-shepherd-workflow.test.js feat(autofix): label-driven takeover and release; fix forced-dispatch green no-op (#7165) 2026-07-19 07:35:13 +00:00
qwen-pr-review-workflow.test.js fix(ci): tighten API error detection to avoid false positive on review prose (#7328) 2026-07-20 12:41:25 +00:00
qwen-resolve-workflow.test.js feat(review): retry transient API failures once; surface quota clearly (#7233) 2026-07-19 15:27:40 +00:00
qwen-triage-workflow.test.js fix(ci): notify silent triage re-runs (#7079) 2026-07-17 16:07:10 +00:00
release-helpers.test.js refactor: extract shared release helper utilities (#3834) 2026-05-05 10:15:17 +08:00
release-sdk-workflow.test.js fix(ci): skip empty SDK release PR (#6861) 2026-07-14 13:19:42 +00:00
sandbox-command.test.js fix(scripts): avoid shell injection in sandbox command detection (#6108) 2026-07-01 16:20:40 +08:00
serve-fast-path-bundle-check.test.js perf(cli): Defer TUI runtime from ACP startup (#7182) 2026-07-19 00:10:15 +00:00
start.test.js fix(review): report what the transcripts prove; build the roster in one call (#7033) 2026-07-18 00:43:57 +00:00
test-setup.ts feat(installer): add standalone archive installation (#3776) 2026-05-11 13:25:48 +08:00
upload-aliyun-oss-assets.test.js feat(installer): add standalone hosted install and uninstall flow (#3828) 2026-05-21 11:57:10 +08:00
vitest.config.ts refactor(auth): unify provider config in core, simplify /auth as "Connect a Provider" (#4287) 2026-05-20 23:48:52 +08:00
workspaces.test.js feat(desktop): Add desktop app package with Qwen ACP SDK integration (#3778) 2026-06-11 21:57:20 +08:00