mirror of
https://github.com/QwenLM/qwen-code.git
synced 2026-08-04 22:00:58 +00:00
* fix(autofix): retry a model API error instead of stranding the PR
When the agent's qwen subprocess dies on a model-side [API Error]
(403 access denied, a 429 quota, a 5xx), run-agent.mjs wrote a
handoff/failure.md, so the handoff step treated it as an EVALUATED
handoff — it advanced the watermark and the next scan saw 'nothing
new', stranding the PR until a manual re-arm. But the agent never
actually evaluated the feedback; the model was unreachable.
#7220 hit exactly this: fork-takeover engaged and ran the agent, the
model returned '[API Error: 403 Model access denied]' (the autofix key
lacks access to qwen3.8-max-preview), and the PR was left with an
advanced watermark that will not retry.
Fix, mirroring #7229's no-output-crash handling:
- run-agent.mjs extracts a [API Error: 4xx/5xx] from the captured
output tail, includes it in failure.md, and drops an
marker file.
- The handoff step reads that marker and routes the failure to the
sentinel-ts (retry) path — the watermark does NOT advance, so the
next scan retries; the round still increments so a PERSISTENT model
failure is bounded by MAX_ROUNDS. The headline names the model error
and, on the final attempt, tells the maintainer to check the autofix
model key/access and re-arm — instead of a generic crash message.
Tests: run-agent.mjs flags a model [API Error] (marker + failure.md)
and does NOT flag a generic failure; the handoff replay treats an
API-error handoff as sentinel|retry (not a watermark advance) with a
model-aware, cause-specific headline. 62/62 + 12/12.
* fix(autofix): scope + broaden the retryable model-API detection (review)
Addresses wenshao's review on #7247:
- Behavioral (1): the agent-api-error marker was written on ANY non-zero
exit whose output tail contained an API-error string — so a loop
guard, a timeout, or an agent-written failure.md (a real verdict)
would wrongly retry and, worst case, silently discard a verdict. The
write is now scoped to the bare-failure branch and guarded by
!timedOut, so only an un-evaluated model failure retries.
- Coverage (2): the old regex only matched a LEADING status digit, so
it missed the canonical rate-limit render, the (Status: …) form, the
bad-key 401, the Chinese quota text, and the unwrapped Qwen OAuth
quota — i.e. most real errors this targets. Detection is now a
whitelist of RECOVERABLE errors (401/402/403/429/5xx + rate-limit /
quota / api-key / RESOURCE_EXHAUSTED / overloaded phrasings, plus the
standalone OAuth-quota form); a 400/404 stays terminal.
- Test gap (3): a writer↔reader contract test now runs the REAL
run-agent.mjs to write the marker, then the extracted workflow reader
block against that same workdir — a rename on either side (proven
with the YAML-only mutation) now fails the suite.
- Smaller: API_ERROR_DETAIL is comment-escaped (sed) and capped
(cut -c1-200) since it derives from agent stdout; the marker match is
single-line ([^]\n]) so a multi-line render can't smuggle a newline;
agent-api-error is added to the run-artifacts list.
Non-recoverable 4xx (400/404) deliberately stay terminal; the live
401/403 config cases retry and self-heal once the key/access is fixed.
79/79 across both suites.
* test(autofix): cover the timeout guard and the OAuth-quota fallback (review)
Two coverage gaps from the ci-bot review on #7247:
- The !result.timedOut guard was only asserted indirectly — no test
emitted an [API Error] AND timed out. Added a case (spawnSync +
QWEN_TIMEOUT_MS=100): qwen streams [API Error: 503] then hangs past
the budget → killed → no marker. A refactor to !loopDetected now
fails here.
- The standalone Qwen-OAuth-quota fallback (unwrapped, no [API Error:])
had no test. Added a case emitting bare 'Qwen OAuth quota exceeded
(limit: 100/min)' → marker written, wrapped as
'[API Error: Qwen OAuth quota exceeded …]'.
* fix(autofix): anchor the API-error code, split retry budget by cause, keep the headline UTF-8
Addresses the review on #7247.
Classifier (points 2 and 4): the status code is now read from its POSITION in
the render (`[API Error: <code>`) instead of matched anywhere in the message.
Matching anywhere retried permanent failures forever — `400 Invalid value for
max_tokens: must be <= 512` matched a bare \b5\d\d\b and `400 context length
exceeded` matched a bare `exceeded`. `exceeded` now only counts as part of
`quota`. A 404 whose message says the model "does not exist or you do not have
access to it" — the OpenAI-compatible render of what a 403 reports — is no
longer terminal.
Retry budget (point 3): the marker now carries the cause class. A transient
429/5xx self-heals and keeps the full round budget; an auth/access error that
only a maintainer can fix is capped at API_AUTH_MAX_ROUNDS (3) and then goes
terminal with the "check the autofix model key/access, then re-arm" headline —
instead of ~100 agent runs and ~100 PR comments over ~17h on a takeover PR.
The terminal round is stamped so the scan's round gate skips the PR while the
sentinel ts keeps the feedback live for a re-arm.
Headline (point 1): `cut -c` counts bytes under GNU coreutils and the
classifier deliberately matches CJK renders, so the 200-byte cap could split a
multi-byte character and emit invalid UTF-8. Guarded with
`iconv -f utf-8 -t utf-8 -c || true`, matching the sibling publish site (the
`|| true` is required — iconv -c exits 1 when it discards).
Minor (point 5): documented that detection is best-effort because apiError is
derived from the last 20 KB of output; `head -1` -> `head -n 1`; tests added
for a permanent 400 carrying a 3-digit number >= 500 and for a >200-byte CJK
render staying valid UTF-8.
* test(autofix): cover the auth-capped retry budget and Chinese API-error patterns (#7247)
* fix(autofix): short-circuit 400 as terminal and classify only the last API error (#7247)
* fix(autofix): treat transport-level API failures as retryable
#7365 stranded at round 2/100 on this render:
[API Error: terminated (cause: read ECONNRESET)]
The connection to the model dropped mid-run. That is as transient as a 429, but
the classifier never saw it that way: a transport failure never got far enough
to have an HTTP status, so it fell through to the keyword arm, and the keyword
arm only knew about rate limits and quotas. It was classified terminal, the
watermark advanced, and a PR that needed nothing but a re-run was handed to a
human.
Verified against the shipped classifier before the fix — every transport render
came back terminal:
terminated (cause: read ECONNRESET) -> terminal
fetch failed -> terminal
socket hang up -> terminal
connect ETIMEDOUT -> terminal
Adds a transport arm to the code-less branch: ECONNRESET, ECONNREFUSED,
ETIMEDOUT, EPIPE, EAI_AGAIN, socket hang up, fetch failed, terminated.
ENOTFOUND is deliberately excluded. A hostname that does not resolve is a
misconfigured endpoint, which repeats forever — the same reasoning that keeps a
bad model name terminal.
Coded errors are unaffected: the arm sits after the status-code branch, so the
400 short-circuit added in
|
||
|---|---|---|
| .. | ||
| installation | ||
| lib | ||
| tests | ||
| acp-http-smoke.mjs | ||
| benchmark-api-latency.mjs | ||
| build-hosted-installation-assets.js | ||
| build-standalone-release.js | ||
| build.js | ||
| build_package.js | ||
| build_sandbox.js | ||
| build_vscode_companion.js | ||
| check-build-status.js | ||
| check-desktop-isolation.js | ||
| check-i18n.ts | ||
| check-lockfile.js | ||
| check-serve-fast-path-bundle.js | ||
| clean-package-build-artifacts.js | ||
| clean.js | ||
| cli-entry.js | ||
| copy_bundle_assets.js | ||
| copy_files.js | ||
| create-standalone-package.js | ||
| create_alias.sh | ||
| daemon-dev.js | ||
| desktop-openwork-sync.ts | ||
| dev.js | ||
| esbuild-shims.js | ||
| generate-changelog.js | ||
| generate-git-commit-info.js | ||
| generate-release-notes.js | ||
| generate-settings-schema.ts | ||
| get-release-version.js | ||
| lint.js | ||
| local_telemetry.js | ||
| measure-flicker.mjs | ||
| pre-commit.js | ||
| prepare-package.js | ||
| prepare.js | ||
| release-script-utils.js | ||
| sandbox_command.js | ||
| sign-release.sh | ||
| start.js | ||
| sync-computer-use-schemas.ts | ||
| telemetry.js | ||
| telemetry_gcp.js | ||
| telemetry_utils.js | ||
| test-rewind-e2e.sh | ||
| test-windows-paths.js | ||
| unused-keys-only-in-locales.json | ||
| upload-aliyun-oss-assets.js | ||
| verify-installation-release.js | ||
| version.js | ||
| workspaces.js | ||