- Remove nonexistent --all-files/-a and --show-memory-usage flags from the
CLI arguments and headless option tables (no longer defined in the yargs
parser in packages/cli/src/config/config.ts).
- Add the commonly-needed --model/-m flag to the headless options table and
fix the --approval-mode example to use the valid choice auto-edit (the
parser rejects the underscore form auto_edit).
- Drop the stale Meta+Enter alias from the external-editor shortcut; that
chord is bound to NEWLINE, while OPEN_EXTERNAL_EDITOR binds only Ctrl+X.
- Document the model.reasoningEffort setting (set via /effort), which is
exposed in the settings dialog but was missing from the settings reference.
Co-authored-by: Claude <noreply@anthropic.com>
* docs: document model/auth settings, /model --vision, and --safe-mode
Refresh user docs to match the current codebase:
- commands.md: add the /model --vision override (vision-bridge model)
- settings.md: add model.baseUrl, model.sessionTokenLimit, visionModel,
and voiceModel; document the deprecated security.auth.apiKey and
security.auth.baseUrl keys with a pointer to modelProviders
- troubleshooting.md: document the --safe-mode flag for isolating
customization issues
* docs: address review feedback on sessionTokenLimit, safe-mode, deprecation notes
- model.sessionTokenLimit: correct default to -1 (runtime fallback in
core/config.ts) and clarify breach behavior (current send dropped, not
session abort) per client.ts SessionTokenLimitExceeded handling.
- --safe-mode: expand the disabled-customizations list to also cover
permission rules, approval mode overrides, memory features, and sandbox
settings, matching cli/config.ts.
- security.auth.apiKey/baseUrl: align deprecation wording with the existing
tools.* entries (**Deprecated.**) and drop the unsubstantiated
'(slated for removal)' qualifier.
* docs: note QWEN_CODE_SAFE_MODE env var as a safe-mode alternative
Document the QWEN_CODE_SAFE_MODE=true environment variable as an
alternative activation path for safe mode, for cases where the CLI
cannot accept flags (verified against isSafeModeEnv in
packages/core/src/utils/safe-mode.ts).
* docs: clarify model.baseUrl, sessionTokenLimit=0, and safe-mode subagents
- model.baseUrl: describe it as a picker-managed disambiguator, not a
hand-editable override (stale values can misroute to a same-id provider).
- model.sessionTokenLimit: note that 0 is treated as unlimited (same as -1),
unlike model.maxToolCalls where 0 disallows all calls.
- --safe-mode: include custom subagents in the list of disabled
customizations (only built-in subagents load in safe mode).
* docs: clarify sessionTokenLimit semantics and add --safe-mode to headless flags
- settings.md: reword model.sessionTokenLimit to reflect that the gate
compares the last recorded prompt token count before the next send
(not a per-send preflight cap), and that the next send is dropped.
- headless.md: add a --safe-mode row to the CLI flags table so the
diagnostic flag is discoverable there, cross-referencing Troubleshooting.
* docs: align safe-mode sandbox wording to 'sandbox settings'
Safe mode passes an empty Settings object to loadSandboxConfig
(packages/cli/src/config/config.ts:1793), so it strips settings-sourced
sandbox config while the --sandbox flag and QWEN_SANDBOX env still apply.
Match headless.md to troubleshooting.md's accurate 'sandbox settings'.
* docs: correct safe-mode approval-mode wording and align both lists
Safe mode only strips settings-sourced approval mode; the --yolo and
--approval-mode CLI flags are evaluated before the safeMode guard
(packages/cli/src/config/config.ts:1521-1528) and still take effect.
Reword to 'settings-sourced approval mode overrides' and note the CLI
flags in troubleshooting.md and headless.md, and make the enumerated
safe-mode disable list identical (same items and order) across both.
---------
Co-authored-by: Claude <noreply@anthropic.com>
* feat(cli): headless runaway-protection guardrails (#4103)
Adds two opt-in run-level budgets and a startup safety warning for
non-interactive / CI / SDK runs. All defaults preserve existing
behavior; the budgets only fire when the user explicitly sets a limit.
Phase 1 — surface unsafe configs and fix doc drift
- New `--yolo`-without-sandbox stderr warning at startup of every
non-interactive run, emitted by `getHeadlessYoloSafetyWarning` in
`packages/cli/src/utils/headlessSafetyWarnings.ts`. Suppressible
via `QWEN_CODE_SUPPRESS_YOLO_WARNING=1` (strict `1`/`true` match so
`=0` / `=false` don't silence it). Strict env match also applied to
the `SANDBOX` check so values like `SANDBOX=0` don't accidentally
bypass the warning.
- Gated on `!config.isInteractive()` at the gemini.tsx call site so
TUI users aren't nagged.
- `docs/users/configuration/settings.md`: corrected
`model.skipLoopDetection` default (`true`, not `false`) and reworded
the `--yolo`/sandbox section — `--yolo` does NOT auto-enable a
sandbox; sandboxing must still be opted into explicitly.
Phase 2 — run-level budgets with distinct exit code
- `--max-wall-time` / `model.maxWallTimeSeconds`: wall-clock duration
for the whole run. Flag accepts `90` (s), `30s`, `5m`, `1h`,
`500ms`. Settings is plain seconds.
- `--max-tool-calls` / `model.maxToolCalls`: cumulative tool
executions (success + failure). Ticked BEFORE each `executeToolCall`
so a budget of N caps the run at exactly N executions.
- New `FatalBudgetExceededError` (exit code 55), distinct from
`FatalTurnLimitedError` (53) and `FatalCancellationError` (130) so
CI scripts can branch on the reason. JSON output mirrors the
`handleMaxTurnsExceededError` / `handleCancellationError` envelope
convention.
- Enforced via `RunBudgetEnforcer` in
`packages/cli/src/utils/runBudget.ts`, wired to the same
`AbortController` as SIGINT so existing cancellation plumbing
carries the abort. A `routeAbort` helper distinguishes budget vs.
SIGINT at the abort-check sites and at the outer catch.
Critical correctness fixes (informed by the #4105 review pass)
- Drain-loop fall-through: the inner drain-item `for await` previously
exited via `finalizeAssistantMessage(); return;`, swallowing a
budget abort that fires during the last drain item and surfacing
exit code 0. Now routes through `routeAbort` so exit 55 is
preserved.
- Settings symmetry: `maxWallTimeSeconds: 0` in settings.json is now
rejected (same as `--max-wall-time 0`); the enforcer treats `<=0`
as "no timer" so silent disable would be a foot-gun.
`validateMaxWallTimeSetting` also rejects `Infinity` / `NaN`.
- `setTimeout` overflow: both parser paths reject durations above
`Math.floor((2^31 - 1) / 1000)s` (~24.8 days). Node clamps
oversized delays to 1ms and fires the timer almost immediately;
fail loud at startup instead.
- First-fence-wins + SIGINT race: `markExceeded` no-ops if the
controller was already aborted by a third party, so a budget tick
arriving after user SIGINT doesn't misattribute the abort to exit
code 55.
- Outer catch re-routes mid-stream `AbortError`s through the budget
handler so users see "Run aborted: …" instead of raw "AbortError".
Tests
- `runBudget.test.ts` (32 tests): parser happy / reject paths,
setting validator, post-increment off-by-one, `maxToolCalls=0`
meaning "disallowed", `-1` meaning unlimited, wall-clock under
fake timers, `stop()` cancels pending timer, idempotent `start()`,
first-fence-wins, SIGINT-race protection.
- `headlessSafetyWarnings.test.ts` (7 tests): YOLO + sandbox / env
matrix; strict-truthy `SANDBOX` check; suppression env.
- Pre-existing suites: `nonInteractiveCli.test.ts` (46),
`gemini.test.tsx` (23), `config/config.test.ts` (220),
`core/utils/errors.test.ts` (12), `core/config/config.test.ts`
(172) all green after picking up the new config getters / CliArgs
fields.
Backward compatibility
- All budgets default to `-1` (unlimited); existing CLI invocations
behave identically.
- New stderr warning only fires in the narrow YOLO-no-sandbox case,
with an explicit suppress env.
- New exit code 55 is purely additive; no existing exit codes change
meaning.
* fix(cli): address audit findings for headless guardrails (#4103, #4502)
Round-1 audit (3 angles × line-by-line + removed-behavior + cross-file)
plus an open-ended design pass surfaced eight correctness issues. This
commit lands all of them; the larger ACP / serve-mode structural items
are documented for follow-up.
Correctness fixes
- headlessSafetyWarnings: `SANDBOX` env check reverted to plain truthy.
The sandbox transport sets `SANDBOX` to `sandbox-exec` (macOS
seatbelt) or the container name (`qwen-code-sandbox`), neither of
which matches `isTruthyEnv`. The PR's strict-`1`/`true` check was
emitting the "no sandbox" warning INSIDE real sandboxes. Match the
rest of the codebase (sandboxConfig.ts, gemini.tsx, Footer.tsx,
prompts.ts, …) which all treat any non-empty value as "sandboxed".
- nonInteractiveCli main-loop abort: add `finalizeAssistantMessage()`
before `routeAbort()`. The drain-item loop already had it (PR #4502
Critical bug #1); the main loop was asymmetric — stream-json
consumers would see an unterminated `message_start` when a budget /
SIGINT abort landed mid-stream.
- nonInteractiveCli drain-loop `routeAbort`: also flush
`flushQueuedNotificationsToSdk(localQueue)` and
`finalizeOneShotMonitors()` before exiting. The old `return`-and-
fall-through path went through the outer holdback loop, which did
this flushing; switching to `routeAbort()` skipped it, so
`task_started` envelopes lost their paired `task_notification`.
- nonInteractiveCli catch handler: emit `adapter.emitResult({...})`
BEFORE `handleBudgetExceededError`, with the budget message as
`errorMessage` when budget tripped. Previously the budget handler
`process.exit(55)`ed before the adapter could emit a terminal
`result` envelope, so STREAM_JSON consumers never saw a stream
terminator on budget exits and hung waiting for one.
- runBudget: new `validateMaxToolCalls` mirrors
`validateMaxWallTimeSetting`. yargs coerces non-numeric flag values
(`--max-tool-calls abc`) to `NaN`, and the enforcer's `>= 0` gate
treats `NaN` and negatives as "no limit", silently disabling the
budget. Reject `NaN`, `Infinity`, fractional, and negative-other-
than-`-1` values at both flag and settings layers. `0` remains
legal (`first tick aborts`), unlike wall-time where 0 is fatal.
- runBudget: new `MIN_WALL_TIME_SECONDS = 1` floor. Previously
`--max-wall-time 500ms` parsed cleanly and aborted on the next
event-loop tick before any model round-trip — almost certainly a
typo (`5m`?) and not a useful guardrail at any rate.
- nonInteractiveCli `tickToolCall`: exempt `ToolNames.STRUCTURED_OUTPUT`.
Under `--json-schema` this is the terminal "I'm done" contract tool,
not real work. Without the exemption a budget-edge completion is
aborted as a false positive (model used N tools then emitted
structured_output as call N+1 → exit 55 instead of success).
- commands/serve.ts: emit the YOLO-no-sandbox warning at daemon
startup when settings.json statically configures
`tools.approvalMode: 'yolo'` with no `tools.sandbox` /
`SANDBOX` env. The daemon can't use `getHeadlessYoloSafetyWarning`
(no Config yet — sessions get their own) so we re-derive the
predicate from settings. Per-session ACP override is documented as
out of scope.
Documentation
- `docs/users/features/headless.md`: new "Scope" subsection under
Run-level budgets explaining (a) `--max-tool-calls` counts top-level
dispatches only — subagent / `agent` tool inner calls are not
counted, (b) `structured_output` is exempt, (c) stream-json input
mode resets budgets per user message, (d) `qwen serve` / ACP
sessions do not currently consult budgets from settings.json.
Tests
- `runBudget.test.ts` grows from 32 → 41 tests: `validateMaxToolCalls`
(NaN / Infinity / negatives / fractional), `parseDurationSeconds`
sub-second rejection, `validateMaxWallTimeSetting` sub-second
rejection.
- `headlessSafetyWarnings.test.ts`: replaced the "still warns when
SANDBOX is 0/false/no" case (which encoded the strict-check bug) with
positive coverage for the real sandbox-set values
(`sandbox-exec`, `qwen-code-sandbox`).
All previously-green suites still green: cli/nonInteractiveCli (46),
cli/gemini.test (23), cli/config/config.test (220), core/utils/errors
(12), core/config/config.test (172). 337 tests across the touched suites.
Won't-fix (out of scope, documented or pre-existing)
- Unpaired `tool_use` in stream-json when a tool is aborted mid-execution
— pre-existing structural gap (SIGINT mid-tool has the same outcome);
PR amplifies it but doesn't introduce it.
- Narrow SIGINT-vs-budget-timer race — already mitigated by
`markExceeded`'s `signal.aborted` check.
- `tickToolCall` increments past abort (cosmetic; only affects the
`observed` value in the error envelope for a pathological caller).
* fix(cli): round-2 audit fixes for headless guardrails (#4103, #4502)
Round-2 audit (after round-1 commit 40ae6dd0f) surfaced two NEW
correctness issues introduced by the round-1 catch-handler restructure,
plus a handful of polish items from a parallel design pass.
Correctness fixes (new bugs from R1)
- nonInteractiveCli catch handler: wrap `adapter.emitResult` in
try/catch. R1 moved the emit BEFORE `handleBudgetExceededError` so
STREAM_JSON consumers see a terminal envelope first. But emitResult
eventually hits `stdout.write`, which throws on EPIPE /
ERR_STREAM_WRITE_AFTER_END when a piped consumer closes early
(`qwen -p ... | head -n 1` is the common CI case). Letting that
throw bubble out skipped both `handleBudgetExceededError` and
`handleError`, dropping the documented exit-code-55 contract
precisely when stdout was in trouble. Best-effort emit and continue
to the exit handler.
- nonInteractiveCli `structured_output` exemption: also require
`config.getJsonSchema?.() !== undefined`. Without that guard, an
MCP server registering an unrelated tool literally named
`structured_output` would silently bypass `--max-tool-calls`. Also
documents (in `headless.md` "Scope") the related caveat that failed
Ajv-validation retries skip the tick too, so a malformed-output
retry loop is NOT bounded by `--max-tool-calls` — combine with
`--max-session-turns` or `--max-wall-time`.
Polish
- runBudget `validateMaxToolCalls` upper bound: cap at 1_000_000.
`1e10` (typo for `1e1`) would otherwise parse cleanly, pass the
`>= 0` gate forever, and silently disable the budget — the exact
foot-gun `MAX_WALL_TIME_SECONDS` was built to prevent. Symmetry.
- runBudget `parseDurationSeconds` sub-second hint: only append the
"did you mean Ns?" suggestion when the input actually contained
`ms`. Bare `0.5` would otherwise produce a useless "did you mean
0.5s?" suggestion.
- nonInteractiveCli `routeAbort`: the `throw 'unreachable'` is only
hit if `handleBudgetExceededError` / `handleCancellationError` ever
becomes resumable (e.g. mocked `process.exit` in a test). Carry
the original exceeded.message into the thrown Error so the outer
catch's `errorMessage` field stays actionable instead of degrading
to a literal "unreachable" string.
- commands/serve.ts: compare `approvalMode` against `ApprovalMode.YOLO`
enum instead of the string literal `'yolo'`. If the enum value is
ever renamed, the startup warning stays in sync with the helper at
`headlessSafetyWarnings.ts` instead of silently going dead.
Documentation
- `headless.md` "Scope": clarify the `structured_output` exemption is
unconditional (including failed validations); add explicit note
that `--max-session-turns` does NOT exempt `structured_output`, so
size to `N+1` for `N` real-work turns under `--json-schema`.
- `headless.md` flag table: add `1.5h` to the accepted-forms hint for
`--max-wall-time` (the parser already accepts fractional units).
Tests
- `runBudget.test.ts`: new coverage for the `validateMaxToolCalls`
ceiling. Total 42 tests across `runBudget.test.ts` (was 41), all
green. cli/nonInteractiveCli, gemini.test, config/config all
unchanged and still green.
Won't-fix (documented above or out of scope)
- ACP per-session approval-mode escalation (mid-session flip to YOLO)
doesn't print the warning — daemon-level wiring; out of scope for
this PR.
- 1s wall-time floor vs higher (5–10s) — debatable, keeping 1s with
loud sub-second rejection; can raise later without semver impact.
- Integration test for the full budget-trip → catch → emitResult →
exit 55 path — requires a process-exit-mocking harness; tracked as
follow-up.
* docs: align headless guardrails examples with R1 sub-second floor
Round-3 audit caught two stale doc surfaces that R1's 1-second wall-time
floor (and R2's `1.5h` fractional-unit addition) didn't update:
- `docs/users/features/headless.md` budget table: replace stale `500ms`
example with `1.5h`, add explicit "minimum 1s — sub-second values are
rejected as typos" note.
- `docs/users/configuration/settings.md` `model.maxWallTimeSeconds` row:
same fix. Also extend `model.maxToolCalls` row with the structured_output
exemption note, the `0` semantic, and the 1,000,000 ceiling that R2
added.
A user copying the documented `--max-wall-time 500ms` example from either
surface would hit a startup error after R1.
Known follow-up (not addressed in this commit)
- No test exercises the R2 `isStructuredOutputExempt` predicate end-to-end.
Adding one needs the same process-exit-mocking harness called out in the
R2 commit as a separate follow-up.
* docs: align JSDoc / schema / CLI help with R1+R2 validation rules
Round-4 final-pass audit caught four schema/help-text/JSDoc surfaces
that drifted from the validators introduced in R1 (1s wall-time floor,
24-day ceiling) and R2 (1M tool-call ceiling, structured_output
exemption, `0` sentinel).
- `runBudget.ts` `parseDurationSeconds` JSDoc: replace stale claim
that `500ms` is accepted and "sub-second precision is preserved"
with the actual contract — `[MIN_WALL_TIME_SECONDS, MAX_WALL_TIME_SECONDS]`,
ms suffix only legal when value resolves to >= 1s. Adds `1.5h` to
the accepted-forms list.
- `settingsSchema.ts` `model.maxWallTimeSeconds` description: now
documents the 1s minimum and ~24-day ceiling.
- `settingsSchema.ts` `model.maxToolCalls` description: documents the
structured_output exemption, the `0` sentinel ("no tool calls
allowed"), and the 1,000,000 ceiling.
- `vscode-ide-companion/schemas/settings.schema.json`: mirrors both
schema descriptions above so the VS Code settings UI auto-completion
matches.
- `config.ts` yargs `--max-wall-time` description: documents the 1s
floor and the ~24-day max.
- `config.ts` yargs `--max-tool-calls` description: documents the
structured_output exemption, the `0` sentinel, and the 1M ceiling.
`qwen --help` is the most-read surface for these flags; matches the
prose docs in headless.md and settings.md.
No code changes — pure doc/help-text alignment.
---------
Co-authored-by: 克竟 <dingbingzhi.dbz@alibaba-inc.com>
* feat(retry): add persistent retry mode for unattended CI/CD environments
When running in CI/CD pipelines or background daemon mode, transient API
capacity errors (429/529) should not terminate long-running tasks after a
fixed number of retries. This adds an environment-aware persistent retry
mode that retries indefinitely for transient errors, with exponential
backoff capped at 5 minutes and heartbeat keepalives every 30 seconds to
prevent CI runner timeouts.
* docs: add persistent retry mode documentation
Add environment variable entries (QWEN_CODE_UNATTENDED_RETRY, QWEN_CODE_BG)
to the settings reference, and a new "Persistent Retry Mode" section to the
headless mode docs covering activation, behavior, and CI/CD usage examples.
* refactor(retry): simplify to single explicit env var QWEN_CODE_UNATTENDED_RETRY
Remove QWEN_CODE_BG and CI=true as activation triggers for persistent retry.
Having multiple env vars with identical behavior adds confusion, and silently
activating infinite retry on CI=true is dangerous — a regular CI test hitting
a 429 would hang forever instead of failing fast.
* fix(retry): address PR review feedback
- Forward caller's abortSignal into retryWithBackoff in both
baseLlmClient.ts and geminiChat.ts so persistent waits remain
cancellable (wenshao)
- Re-apply maxBackoff and capMs after jitter so delays strictly
respect stated caps (Copilot)
- Respect shouldRetryOnError in persistent mode so callers can
force fast-fail even for transient 429/529 errors (Copilot)
- Guard sleepWithHeartbeat against infinite loop when heartbeat
interval is <= 0 via Math.max(1, ...) (Copilot)
- Normalize isEnvTruthy with trim/toLowerCase for robust env
var parsing across CI conventions (Copilot)
* test(retry): add missing UT for shouldRetryOnError override and heartbeat zero-interval guard
* fix(retry): do not cap Retry-After delays at maxBackoff
Server-specified Retry-After values should only be limited by the
absolute cap (capMs/6h), not the exponential backoff cap (maxBackoff/5min).
Jitter is also skipped for Retry-After since the server already specified
the exact wait time.
* refactor(retry): align isUnattendedMode with project env parsing convention
Replace custom isEnvTruthy (trim + toLowerCase) with strict matching
(val === 'true' || val === '1') to match parseBooleanEnvFlag used
elsewhere in the codebase. Prevents inconsistent behavior where
'TRUE' or ' 1 ' would activate persistent retry here but not in
telemetry or other env-driven features.
* test(retry): add Retry-After handling tests for persistent mode
Cover three key behaviors:
- Retry-After is NOT capped at maxBackoff (only at capMs)
- Retry-After IS capped at persistentCapMs absolute limit
- Retry-After delays have no jitter applied
* fix(test): add isUnattendedMode to retry.js mock in baseLlmClient tests
The existing vi.mock for retry.js only exported retryWithBackoff.
After adding isUnattendedMode to the retry module, baseLlmClient.ts
imports it, causing all 10 generateJson tests to fail with
'No "isUnattendedMode" export is defined on the mock'.
* fix(retry): wire persistent retry mode into client.ts generateContent
Forward persistentMode and abortSignal to retryWithBackoff() in
GeminiClient.generateContent(), matching the existing wiring in
baseLlmClient.ts and geminiChat.ts.