fix(autofix): give the repair pass a budget it can finish in (#9691)

* fix(autofix): give the repair pass a budget it can finish in

The repair attempt ran on a hardcoded 18-minute agent budget while the
primary attempt gets 120 minutes from a configurable default. Raise the
repair budget to 45 minutes and carry the step and job caps that bound it.

The repair attempt is handed strictly less to work with than the primary
one: a deterministic rejection is an opaque check failure, not the
structured review feedback the primary attempt receives, so it must first
re-derive which change caused the rejection before it can amend anything.
Giving that 15% of the primary budget inverted the difficulty and the
allowance.

Measured on four takeover PRs over nine rounds on 2026-08-21: the primary
attempt reported `Autofix agent completed address-review successfully.` in
9 of 9 rounds, and the repair attempt hit `timeout (1080000ms)` in 9 of 9.
Every one of those rounds discarded work the primary attempt had already
finished — on #9340 a completed `origin/main` conflict resolution across
three files with two mutation probes and `vitest run src/commands/review/`
green at 97 files / 4335 tests. Three such rounds tripped
TIMEOUT_WINDOW_CAP and parked the PR at its round cap with
`autofix/needs-human`.

The rejections themselves were a mix — a flaky unrelated test (#9648), a
genuine defect in the PR, and a scope violation — so this is not a
substitute for fixing any one of them. It is the step they all funnel
through: whatever the gate rejects on, the repair attempt has to be able
to finish before the round can push.

Carried bounds, each preserving its documented margin:

- repair step cap 20m → 55m (budget + the same 10-minute margin the
  primary attempt keeps, so the internal kill path still writes
  `agent-timeout` before the step cap fires)
- review-address job cap 300m → 330m (the four long steps now sum to 305m
  plus the 25m setup/report reserve)
- PENDING_STALE_MIN 330 → 360 (its 30-minute margin over the job cap, so a
  live review-address run is never aged out mid-flight)

45 minutes is deliberately a fraction of the primary budget: a repair that
cannot land in 45m is a handoff, not a longer retry.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VgTjRF91xANQh6SY9YGyCf

* fix(autofix): carry the raised repair bounds through sibling prose

* fix(autofix): revert design-record edits outside this PR's footprint (#9691)

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: qwen-code-dev-bot <qwen-code-dev-bot@users.noreply.github.com>
Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>
This commit is contained in:
qqqys 2026-08-23 00:32:36 +00:00 committed by GitHub
parent 1007bcacfc
commit acc46e58cb
No known key found for this signature in database
GPG key ID: B5690EEEBB952194
3 changed files with 27 additions and 16 deletions

View file

@ -597,9 +597,9 @@ describe('qwen-autofix workflow', () => {
expect(reviewScanJob).toContain('echo "targets=[]" >> "${GITHUB_OUTPUT}"');
expect(reviewScanJob).toContain('active checks in flight; skipping until');
// Staleness bound must sit above legitimate check runtimes (a review-address
// job runs up to its 300-minute cap) so an active run is never aged out
// job runs up to its 330-minute cap) so an active run is never aged out
// mid-flight.
expect(reviewScanJob).toContain('PENDING_STALE_MIN=330');
expect(reviewScanJob).toContain('PENDING_STALE_MIN=360');
// The staleness filter itself, including the comparison operator: a check only
// blocks if its start is newer than the cutoff. Asserting `> $cut` too means a
// flipped comparison (which would age out live checks → double-processing) is
@ -4178,7 +4178,7 @@ describe('qwen-autofix workflow', () => {
// No rollup entries → dispatchable.
expect(runMarkerCheck([])).toBe('pass');
// A stranded marker must NOT keep blocking through the 330-minute
// A stranded marker must NOT keep blocking through the 360-minute
// HAS_PENDING_CHECKS gate after its TTL expired: replay the gate's jq
// over fixture rollups.
const pendingGate = reviewScanJob.match(
@ -4228,7 +4228,7 @@ describe('qwen-autofix workflow', () => {
checkRun('build', 'IN_PROGRESS', '2026-08-17T07:50:00Z'),
]),
).toBe('true');
// ...a check stuck past the 330-minute horizon is aged out...
// ...a check stuck past the 360-minute horizon is aged out...
expect(
runPendingGate([
checkRun('build', 'IN_PROGRESS', '2026-08-17T01:00:00Z'),
@ -14672,9 +14672,9 @@ exit 1
expect(repairDeterministicRejectionStep).toContain(
"steps.verify.outputs.retryable == 'true'",
);
expect(repairDeterministicRejectionStep).toContain('timeout-minutes: 20');
expect(repairDeterministicRejectionStep).toContain('timeout-minutes: 55');
expect(repairDeterministicRejectionStep).toContain(
"QWEN_TIMEOUT_MS: '1080000'",
"QWEN_TIMEOUT_MS: '2700000'",
);
const settingsJson = (step) =>
step.match(/SETTINGS_JSON: \|-\n([\s\S]*?)\n {8}run: \|-/)?.[1] ?? '';