* fix(kosong): fail fast on quota-exhausted 429 instead of retrying
A 429 caused by an exhausted account quota or insufficient balance
(Moonshot error.type "exceeded_current_quota_error", OpenAI
"insufficient_quota") can never succeed on retry, yet it was classified
as APIProviderRateLimitError and silently retried for the whole budget
(10 attempts, ~3 minutes of backoff) with no UI feedback — the session
appeared frozen on every request.
Introduce APIProviderQuotaExhaustedError, minted in
normalizeAPIStatusError from the structured body error.type/error.code
forwarded by convertOpenAIError, with billing-anchored message patterns
as a fallback for gateways that flatten the body to text. The new class
is excluded from isRetryableGenerateError (fail fast, even when a
retry-after header is present) and from isProviderRateLimitError (no
swarm requeue/suspend). toKimiErrorPayload and translateProviderError
map it to provider.api_error (retryable: false) instead of
provider.rate_limit, and classifyApiError reports it as
quota_exhausted in telemetry. agent-core-v2 mirrors the same fix.
Transient rate-limit 429s keep the existing retry, backoff, and
Retry-After behavior (verified end-to-end against a mock provider:
quota body fails after attempt 1/10; rate-limit body still walks the
full 10-attempt ladder).
Behavior changes to note: quota-failed swarm subagents now fail
instead of suspending indefinitely as "Rate limited...", and quota
errors cross the wire as provider.api_error rather than
provider.rate_limit.
* fix(kosong): classify quota exhaustion in OpenAI Responses stream errors
Responses response.failed / error SSE events carry no HTTP status and
were minted by errorFromOpenAIResponsesEvent as either a rate-limit
error (rate_limit_exceeded / embedded status_code=429) or a base
ChatProviderError — and the base class falls into the retryable
unclassified-failure fallback, so an insufficient_quota event still
burned the whole retry budget on the openai_responses path. Route the
event code and message through the same quota-exhausted check before
the rate-limit branch, in kosong and the agent-core-v2 mirror. Covers
all three entry paths (error events, response.failed, nested gateway
frames) since they share the single converter.
* style(agent-core-v2): drop inline comments per AGENTS.md header-only rule
agent-core-v2 comments live solely in the top-of-file block, never
beside functions or statements; the kosong twins keep the full
rationale.
* refactor(kosong,agent-core-v2): move quota-429 checks to vendor hook
Per review on #1857: the knowledge of how a backend signals quota
exhaustion is vendor-specific and must not run for every
OpenAI-compatible provider from the shared conversion layer.
- Add a convertError hook: ProtocolTrait.convertError in agent-core-v2
(single-value, last-declarer-wins, bound by composeOpenAIChatHooks /
composeAnthropicHooks / traitConvertError) and an equivalent optional
hook parameter on convertOpenAIError / convertAnthropicError. Bases
consult it with the raw failure (SDK error on HTTP paths, raw event on
the Responses in-stream path) after the abort guard, before their own
rules.
- Declare Moonshot's quota signals (exceeded_current_quota_error,
billing wordings) on the Kimi side: kimiOpenAITrait and
kimiAnthropicTrait in v2, the KimiChatProvider and KimiFiles catch
sites in kosong, all through the new classifyKimiQuotaError.
- Drop the options parameter from normalizeAPIStatusError and the
shared quota code/pattern tables: the contract layer keeps only the
vendor-neutral APIProviderQuotaExhaustedError type and its retry /
rate-limit / wire-mapping semantics.
- The OpenAI bases keep recognizing only OpenAI's own documented
insufficient_quota code (HTTP and Responses stream events) as
protocol knowledge of that wire.
Behavior: kimi and openai provider types classify exactly as before;
an unregistered vendor speaking Moonshot billing wordings through a
plain openai transport now stays a retryable rate limit by design.
* fix(kosong,agent-core,agent-core-v2): wire kimi quota hook fully
Follow-up to the second review round on #1857, all four findings:
- Kimi-over-Anthropic (legacy engine): AnthropicOptions gains the same
optional convertError hook as the OpenAI bases, threaded through
AnthropicStreamedMessage and every catch site, and the provider
manager's anthropic route now passes classifyKimiQuotaError for
provider type kimi — a quota-exhausted 429 over this transport
previously still burned the retry budget. classifyKimiQuotaError now
also walks error -> .error -> .error.error for the code/type, since
the Anthropic SDK keeps the full body on .error instead of hoisting.
- v2 telemetry: ApiErrorKind gains 'quota_exhausted' and
classifyApiError checks APIProviderQuotaExhaustedError before the
generic 429 branch, matching the legacy engine's reporting.
- Hook contract: converted ChatProviderErrors now pass through before
the vendor hook is consulted in convertOpenAIError /
convertAnthropicError (both engines), so the hook sees each raw
failure exactly once even when a stream-minted error crosses an
outer catch; tests assert the single consult.
- protocolTrait: the convertError member doc shrinks to the concise
style and the consult contract moves into the file header's
composition rules.
* test(kosong,agent-core,agent-core-v2): lock quota hook assembly paths
Third review round on #1857:
- Fix the v2 anthropic base header and AnthropicHooks doc still claiming
withThinking is the only hook.
- Drop the two remaining non-header JSDoc blocks in protocolTrait.ts per
the AGENTS.md header-only rule; the consult contract already lives in
the file header.
- Update the ProtocolTrait contract test to the seventeen-hook shape
(convertError included) and cover the traitConvertError binding.
- Add real-assembly regression probes: the v2 registry composes a
(kimi, anthropic) provider whose mocked SDK client throws a Moonshot
quota 429 and generate rejects with the non-retryable
APIProviderQuotaExhaustedError (a plain anthropic composition keeps
the same 429 retryable); the legacy ProviderManager routing test
asserts convertError is classifyKimiQuotaError on the kimi-anthropic
route and absent for plain anthropic; the legacy provider threads
options.convertError to its generate catch.
* test(kosong,agent-core-v2): cover KimiFiles quota 429 and drop stale docs
Fourth review round on #1857:
- Drop the AnthropicHooks member JSDoc (its content already lives in the
anthropic.ts and anthropicHooks.ts file headers) and fix the anthropic
contrib header still calling the hook set single-hook.
- Add the missing KimiFiles regression in both engines: a mocked files
client rejecting with a Moonshot quota 429 makes uploadVideo reject
with the non-retryable APIProviderQuotaExhaustedError, locking the
classifyKimiQuotaError argument at the upload catch sites.
When a provider returns an HTML error page (e.g. nginx 413 Request Entity Too Large), the error message carried CRLF line endings and raw HTML. The trailing carriage returns made the TUI render the error line as blank. Extract the page title for the wire message and strip carriage returns before rendering.