* feat(agent-core): align coder subagent tools with v2 and drain background tasks
- align the bundled coder subagent profile tools with agent-core-v2
CODER_TOOLS (Skill, Agent, AgentSwarm, Task* trio, plan-mode tools,
TodoList); cron tools stay declared-but-undelivered on sub agents,
matching v2
- rebuild builtin tools in setActiveTools: profile-gated capabilities
(Bash/Agent allowBackground via the Task* trio) were baked at
construction with an empty enabled set, so profiles applied later
(every subagent) silently lost them
- hold subagent completion until the child agent's background tasks
settle (print-mode drain semantics) and suppress their terminal
notifications, so no unobserved follow-up turn runs on a finished
subagent; the run's timeout/cancel signal bounds the drain
- cover with real-Session e2e (tool execution, nested Agent/AgentSwarm,
drain blocking and cancel) and update profile/subagent-host tests
* chore: add changeset for coder subagent tool alignment
* docs(agents): sync sub-agent capabilities with expanded coder tool set
- coder sub-agents can now dispatch nested sub-agents and use background
tasks, todo lists, Plan mode, and skills; drop the stale claim that
sub-agents cannot schedule nested sub-agents
- note that a sub-agent run reports completion only after its background
tasks settle
* fix(agent-core): close drain race for tasks settled before completion
A background task that terminated before the drain's suppression pass was
excluded from the active-only list, so its terminal notification escaped
suppression and could still steer an orphan turn onto the finished
subagent. Suppress every child task (including settled ones whose
notification may still be in flight) and run the pass both before and
after the settle wait; notification delivery re-checks suppression after
its async output snapshot, which makes the block deterministic. Cover the
mid-turn delivery path with an e2e case.
* feat: raise default per-step LLM retry budget to 10 attempts
The default of 3 attempts only produced ~1.5s of backoff (0.5s/1s), so
sustained provider overload (429) surfaced to the user almost
immediately. With 10 attempts the exponential ramp (500ms base, x2,
32s cap, 25% jitter) waits out multi-minute overload windows before
failing the turn. Applies to both agent-core (chatWithRetry) and
agent-core-v2 (stepRetry service); loop_control.max_retries_per_step
still overrides.
* test(agent-core): pin retry-dependent tests to the new default budget
- error-paths: the warn-log attempt field now reads 1/10 with the
default 10-attempt budget
- goal-session pause tests: every LLM call throws a retryable error, so
the default 10-attempt backoff (~2min) blew the 5s test timeout; pin
maxRetriesPerStep=1 since the pause behavior does not depend on the
retry count
* feat(agent-core): drop default timeouts for print-mode background tasks
- add `bash_task_timeout_s` config under [background] (0 = no timeout),
covering the background Bash default and the re-arm after a foreground
command is moved to the background on timeout
- allow `[subagent] timeout_ms = 0` (no timeout) and thread it through
Agent/AgentSwarm task registration and the swarm batch timer
- fill both with 0 in print-mode config defaults so `kimi -p` never kills
background work by wall-clock and only the model stops a task;
interactive defaults are unchanged
- sync the Bash tool description/parameter text with the effective
default and update user docs (en/zh) plus changeset
* test(agent-core): satisfy KimiConfig providers requirement in print-defaults tests
* fix(agent-core): clear the foreground deadline on detach when the background timeout is disabled
`detach()` only re-arms when `detachTimeoutMs` is defined, so passing
`undefined` with `bash_task_timeout_s = 0` kept the armed foreground
deadline and a manually detached command was still killed at its
foreground timeout. Pass `0` instead so the reset clears the timer.
Caught by Codex review on #1737.
* fix: preserve provider thinking effort values
* fix: resolve Kimi thinking effort fallbacks
* fix: use resolved Kimi effort in v2 requests
* fix: align Kimi effort resolution paths
* fix: synchronize forced Kimi effort state
* fix: synchronize forced Kimi effort in v2 state
* fix: tolerate unresolved models in v2 status
* fix(agent-core): mark auto-approved plan exits as not user-reviewed
In auto permission mode ExitPlanMode is approved without user
involvement, but its result read "## Approved Plan:" exactly like a
genuine user approval and the transcript showed a green "Approved"
chip. The model treated that as a signal to start executing, even
when the user had asked it to stop after planning.
- Emit an "auto-approved, not user-reviewed" result with a note that
execution follows the user's original instructions (v1 + v2).
- Extend the auto-mode reminder: plan approvals are automatic and are
not a user signal to proceed (v1 + v2).
- Render an "Auto-approved" warning-toned chip in the TUI, keeping
backward compatibility with the old result marker.
- Document the Auto mode plan-exit behavior in interaction guides.
* fix(agent-core): only mark plan exits as auto-approved in auto permission mode
Address review feedback on the auto-approved plan exit change:
- Branch the direct-execution output on the permission mode in both
engines: only auto mode yields the "auto-approved, not
user-reviewed" result and telemetry outcome. In manual / yolo modes
the direct path means a configured or session allow/ask rule let the
call through — an explicit user decision that keeps the user-approved
output, the Approved transcript chip, and the approved outcome.
- Move the v2 rationale out of the method body into the file header,
per the agent-core-v2 comment convention.
- Drop the absolute "only Auto mode" wording from the interaction
guides (en/zh).
* fix(agent-core-v2): fix wire migration rewrite race and reject unversioned logs (8 files)
- await the migration rewrite in wireRecordService.restore() so restore only resolves once the migrated log is durable and rewrite failures surface
- serialize AppendLogStore.rewrite() behind the log flush so appends arriving during the whole-file replace land after the new content
- reject wire logs missing the metadata envelope (missingWireMetadataError) in restore, session fork, and agent create instead of fabricating a current-version envelope over legacy records
* fix(agent-core-v2): stop background tasks on session close (10 files)
- add IAgentTaskService.stopAllOnExit: suppress terminal notifications, then SIGTERM → grace → SIGKILL for every active task (v1 stopBackgroundTasksOnExit parity, gated by keepAliveOnExit)
- call stopAllOnExit from agentLifecycle.remove before scope disposal, and abort live tasks in AgentTaskService.dispose as a last resort for disposal paths that bypass the graceful close
- honor [task] killGracePeriodMs in the stop grace window and bind the v1 KIMI_CODE_BACKGROUND_KEEP_ALIVE_ON_EXIT env override for keepAliveOnExit
* fix(agent-core-v2): root task persistence at the owning agent's scope
- persist task records under <sessionScope>/agents/<agentId>/tasks/ (v1's
per-agent layout) so tasks written by older versions are found on resume
- stop one agent's restore from loading, marking lost, or re-notifying
another agent's tasks
- add a per-agent isolation test covering loadFromDisk + reconcile
* fix(agent-core-v2): address task lifecycle review findings
* fix(agent-core-v2): preserve queued appends across rewrites
* test(agent-core-v2): cover rewrite cutover flush
* fix(agent-core-v2): surface append durability failures
* fix(agent-core-v2): retire released append log buffers
* fix(agent-core-v2): preserve session-level task records
* fix(agent-core-v2): harden append log cutover recovery
* fix(agent-core-v2): suppress only background exit notifications
---------
Co-authored-by: qer <wbxl2000@outlook.com>
* feat(agent-core): make subagent timeout configurable and raise default to 2h
- add `[subagent] timeout_ms` config (env `KIMI_SUBAGENT_TIMEOUT_MS` overrides) to replace the hardcoded 30-minute cap for Agent / AgentSwarm subagents
- raise the default subagent timeout from 30 minutes to 2 hours
- thread the value through tool construction so foreground and background subagents use it, with the timeout message reflecting the effective value
* feat(background): add print_background_mode with steer for multi-turn -p runs
- add `[background].print_background_mode` (`exit`/`drain`/`steer`) and
`print_max_turns`; when unset, falls back to `keep_alive_on_exit = true`
mapping to `drain`, preserving existing behavior
- core: `Session.handlePrintMainTurnCompleted()` returns `finish`/`continue`;
in `steer` mode the run stays alive so a background-task completion
`turn.steer`s the main agent into a new turn (matching background
subagents), bounded by `print_wait_ceiling_s` and `print_max_turns`
- cli: print driver follows every main turn instead of only the first and
defers `finish()` until the run quiesces or a limit is hit
- plumb `handlePrintMainTurnCompleted` through node-sdk RPC; docs + tests
* docs(changelog): sync 0.23.5 from apps/kimi-code/CHANGELOG.md
* chore: update config model doc
* docs: update config-files example with new models and services
---------
Co-authored-by: liruifengv <liruifeng1024@gmail.com>
* feat(kosong): classify HTTP 413 request-body-too-large as a dedicated error type
* feat(agent-core): lower default image downscale cap to 2000px and make it configurable
* feat(agent-core): strip media to text markers and retry when the compaction request is too large
* feat(agent-core): cap model-initiated image reads with a configurable byte budget
* feat(agent-core): resend with degraded media when the provider rejects the request body as too large
* test(agent-core): add explicit timeouts to encode-heavy image budget tests
* feat: add WebP decoding support with wasm integration
- Introduced a new WebP decoding module using @jsquash/webp's wasm decoder.
- Implemented functions to decode WebP images and check for animated WebP formats.
- Updated image compression tests to include scenarios for WebP handling, including encoding and decoding.
- Enhanced error handling for API request size limits to accommodate various error messages.
- Updated pnpm lockfile to include new dependencies for WebP encoding and decoding.
* chore(changeset): consolidate this PR's entries into one
* fix(nix): update pnpmDeps hash for merged lockfile
* feat(agent-core): refuse HEIC/HEIF reads with platform-matched conversion guidance
* feat(kosong): honor explicit anthropic max output override
Add claude-opus-4-8 output ceiling and treat explicit max_output_size/KIMI_MODEL_MAX_OUTPUT_SIZE as the final Anthropic max_tokens value. Sync configuration and environment variable docs.
* feat(kosong): drop claude-opus-4-8 ceiling
Revert the newly added claude-opus-4-8 default output ceiling while keeping explicit max_output_size overrides for Anthropic.
* feat(agent-core): record llm request trace in wire.jsonl
Add three observability record types so every request sent to the model
can be reconstructed from the wire log at the logical-request level:
- llm.tools_snapshot: content-addressed snapshot of the top-level tools
table as sent (post deferred-strip), written once per unique table
- llm.request: one record per outbound request (retries, strict resends,
and compaction rounds included) carrying the effective request params
and hash links to the system prompt and tools snapshot
- mcp.tools_discovered: the server's verbatim tools/list result plus the
agent's gating (allow-list, collisions), deduplicated by content hash
Observability records never feed state rebuild; replay only restores the
write-dedup cursors. The records/types.ts contract now documents the two
record classes explicitly (persisted is not the same as replayed).
Recording happens at the single Agent.generate choke point. The
LLMRequestLogFields side channel gains kind/projection/maxTokens/
droppedCount, chatWithRetry preserves caller-set fields, and compaction
tags its requests. The vis wire view renders the new record kinds.
* fix(agent-core): record the provider-clamped completion cap in the request trace
The llm.request trace recorded the client-requested budget cap, but
chat-completions providers tighten the actual wire value inside
withMaxCompletionTokens (remaining-context sizing, transport ceilings,
model-default resolution) — with the default budget the clamp is active
on nearly every non-empty-context request, so the recorded value did not
match what was sent.
Providers now expose the effective cap they computed as a readonly
maxCompletionTokens field on the clone, and the recorder reads it from
the effective provider at the Agent.generate choke point. This replaces
the side-channel recomputation, which is removed along with the
appliedCompletionBudgetCap helper.
* fix(agent-core): park pre-replay MCP discovery records and hash the collision outcome
Two wire-hygiene fixes for the mcp.tools_discovered trace:
Parking: the real Session ordering connects MCP servers concurrently with
agent construction, so ToolManager can observe a connected server before
agent.resume() has replayed the wire. Recording at that point bypassed
the restored dedup cursor (duplicating a 1-50KB record on every resume)
and appended a stray metadata record ahead of replay. AgentRecords now
exposes a one-shot opened latch — set when replay completes (after the
migration rewrite flushes) or when the first live record is logged — and
ToolManager parks discoveries until then, re-running the dedup check at
drain time. A frozen range-limited replay never opens; those agents are
transient previews.
Collision hashing: the dedup hash now covers the collision outcome, not
just the raw list and allow-list. Collisions depend on which other
servers hold a sanitized qualified name at registration time, so a
server can re-register with identical tools but a flipped outcome; that
gating change must produce a new record instead of being suppressed.
* fix(agent-core): skip the request trace for pre-flight-aborted calls
Mirror kosong generate()'s pre-flight abort check at the Agent.generate
choke point: a call whose signal is already aborted never reaches the
wire (generate throws before dispatching), so it must not leave an
llm.request/llm.tools_snapshot trace or a diagnostic log line claiming a
request was sent. Recording stays before dispatch for every call that
passes the gate, preserving the crash-safety of the trace.
* chore(agent-core): remove a leftover adaptive-thinking override hook
The adaptiveThinkingOverride option was a temporary local hook explicitly
marked for removal before commit. Nothing passes it, so resolution falls
back to the alias-level adaptiveThinking value in all cases; drop the
option and the dead indirection.
* fix(kosong): derive the exposed completion cap from generation kwargs
maxCompletionTokens was a field stored only by withMaxCompletionTokens,
so caps that reach the wire through other paths were invisible to the
request trace: with completion budgeting disabled via env, Anthropic
still sends the constructor-resolved max_tokens (required by the
Messages API), and constructor-level kwargs like OpenAILegacyOptions
maxTokens were likewise unreported.
Replace the stored field with a getter derived from each provider's
generation kwargs — the single source the request body reads — covering
constructor defaults, direct withGenerationKwargs configuration, and
budget application in one place. Kimi mirrors its request-time legacy
max_tokens alias normalization; openai-legacy reuses the same
normalizeGenerationKwargs the request path uses.
* feat(agent-core): add thinkingKeep passthrough for Kimi providers and update tests
* feat(agent-core): enable Preserved Thinking by default on the Anthropic provider
Default thinking.keep to "all" for the Anthropic provider (Claude and Kimi in Anthropic-compatible mode) while Thinking is on, via a context_management clear_thinking_20251015 edit, mirroring the Kimi default. Reuses [thinking] keep and KIMI_MODEL_THINKING_KEEP (env > config > default "all"); off-values disable it.
* feat(kosong): route Anthropic Preserved Thinking through the beta Messages API
Force the beta endpoint (client.beta.messages.create) when thinking.keep is enabled, since clear_thinking_20251015 is only honored there. Also prepend clear_thinking to any existing context-management edits (for example clear_tool_uses) instead of replacing them, keeping it first as Anthropic requires when combining edits.
* docs: clarify Anthropic beta endpoint and compaction keep behavior
Note in code comments and bilingual docs that enabling Anthropic Preserved Thinking routes requests to the beta Messages API (client.beta.messages.create), with keep=off as the escape hatch back to the standard endpoint. Correct the resolveThinkingKeep comment to reflect that compaction shares ConfigState.provider and intentionally carries the same keep.
* test(kosong): cover Anthropic beta endpoint (streaming and forced betaApi)
Add a streaming beta-endpoint capture and a test that withThinkingKeep forces the beta endpoint even when constructed with betaApi: false, pinning down the documented behavior.
Default `thinking.keep` to "all" when Thinking is on so prior `reasoning_content` is kept across turns. Add `[thinking] keep` to config.toml and keep `KIMI_MODEL_THINKING_KEEP` as an override (env > config > default); off-values disable it.
Add the two kimi server run flags introduced in #1368 to the CLI reference (en + zh), including a danger callout for the auth-bypass flag, and add the changeset so the next release notes the feature.
* feat(agent-core): guide the model away from repeating denied or failed tool calls
- system.md: add a diagnose-before-retrying paragraph next to the existing
permission-denial guidance, covering failed tool calls
- permission: when the user rejects an approval on the main agent, tell the
model not to re-attempt the exact same call (sub agents already had an
equivalent hint)
* fix(agent-core): close abandoned tool exchanges and dedupe duplicate tool_use ids
A turn that dies between a recorded tool.call and its paired tool.result
(e.g. a transcript write failure mid-batch) used to leave
pendingToolResultIds open forever: every later message was stranded in
deferredMessages and user input was silently swallowed.
- runOneTurn now defensively closes any dangling tool calls when a turn
ends (completed, cancelled, or failed), synthesizing an error result
that names the cause, with a warn log and a tool_exchange_abandoned
telemetry event
- the projector drops assistant tool calls whose id already appeared
earlier (first occurrence wins): a duplicate id is wire-invalid on
strict providers and not repairable by the strict resend; reported via
the existing projection-repair log and telemetry
- resume-side closePendingToolResults now logs what it closes (warn for
a mid-history gap, info for the routine trailing interruption)
* chore: add changesets for tool exchange fixes
* fix(agent-core): scope duplicate tool_use id dedup to the strict resend
Unconditional dedup regressed providers that emit per-response counter
ids (e.g. call_0 in every step) and accept their own duplicates: later
tool exchanges silently vanished from the projected history, and a
duplicate call's own recorded result was left dangling.
- the dedupe pass is now opt-in via dedupeDuplicateToolCalls and enabled
only in strictMessages, so the normal projection keeps the history the
provider produced
- the pass also drops every tool result after the first for an id, so no
dangling tool message survives; when the kept call has no result of
its own, the surviving one is reattached by the adjacency repair
- kosong now classifies the Anthropic "tool_use ids must be unique" 400
as a recoverable request-structure error so it triggers the strict
resend
* feat(cli): wait for background subagents before exiting kimi -p
When `background.keep_alive_on_exit` is enabled, `kimi -p` now waits for
all background subagents to reach a terminal state before exiting, bounded
by `background.print_wait_ceiling_s` (default 3600s). This lets concurrent
background subagents run to completion in single-turn runs instead of being
torn down when the main agent's turn ends.
* docs: use kimi-for-coding in model overrides example
kimi-for-coding is the stable public model ID users actually configure;
kimi-k2 is the underlying model name and shouldn't appear in the config
example. Demonstrate overrides with max_context_size and display_name.
* docs: drop dangling references to commented-out experimental section
The `## experimental` section was commented out when micro_compaction
was removed, but the top-level fields table and the intro sentence still
linked to the now-dead #experimental anchor. Remove those references.
* feat(tui): include shell commands in input history
Shell commands entered through the `!` prompt are now saved to input history. Recalling one restores bash mode, and in bash mode Up only cycles through previous shell commands while a normal prompt browses all history.
* docs(interaction): document shell command recall in input history
Note that shell commands are now saved to input history and can be recalled in Shell mode, in both the English and Chinese interaction guides.
* feat(pi-tui): add setHistoryFilter and onRecall to editor history
Add two first-class hooks to the editor's history navigation: setHistoryFilter to limit which entries Up/Down visit, and onRecall to decorate a recalled entry before it is shown. Draft restore, direction-aware cursor placement, and undo behavior are unchanged.
* refactor(tui): use pi-tui history filter for shell command recall
Replace the CustomEditor navigateHistory shadow with pi-tui's setHistoryFilter + onRecall hooks, wired in the editor-keyboard controller. This keeps pi-tui's draft-restore and direction-aware cursor behavior intact (the shadow dropped both) and moves the shell/prompt filtering and mode-restore logic into the business layer.
* feat(pi-tui): save and restore host state with the history draft
Add onHistoryDraftSave/onHistoryDraftRestore hooks so hosts can stash their own state when entering history browsing and restore it when the user navigates back to the draft. The saved host state is discarded when browsing ends any other way (typing, submit), mirroring the editor draft lifecycle.
* fix(tui): restore input mode when returning to the history draft
Wire pi-tui's history draft save/restore hooks to the editor input mode. Without this, recalling a shell entry and then pressing Down back to an empty draft left the editor in bash mode, so the next typed message was submitted as a shell command.
* fix(pi-tui): capture host draft state before running the history filter
Fire onHistoryDraftSave before the history filter runs when entering browse, so the host's filter can read the browse-entry mode rather than a mode that changes as entries are recalled. The captured state is still only committed once a matching entry is found.
* fix(tui): lock history filter to the browse-entry mode
Lock the history filter to the input mode captured when entering browse. Previously the filter read inputMode live, so after recalling a shell entry (which flips to bash mode) a second Up would only show shell commands.
* feat: add KIMI_MODEL_THINKING_EFFORT to force a thinking effort
Send thinking effort only when the model declares it in support_efforts, and add the KIMI_MODEL_THINKING_EFFORT environment variable as an escape hatch to force a specific effort regardless of declared support.
* test: align thinking effort expectations with support_efforts gating
Update the kimi adapter e2e and compaction tests that asserted the previous pass-through behavior on models without support_efforts.
* feat: add model alias overrides
Preserve user model overrides across provider catalog refreshes and resolve effective model metadata for runtime, TUI, protocol, and ACP consumers.
* fix: apply model display name overrides
Show overridden model display names in the footer, welcome panel, status output, and model switch confirmations.
* fix: pass through kimi effort when undeclared
Keep support_efforts authoritative when declared, but pass requested Kimi thinking effort through when the model does not declare support_efforts.
* fix: honor model overrides in effort commands
Use effective model metadata for /effort choices and for always_thinking clamping when resolving thinking effort.
* fix(provider): honor base_url for google-genai and vertexai providers
The google-genai and vertexai provider types silently ignored a configured
base_url and always hit generativelanguage.googleapis.com (e.g. a Gemini-
compatible proxy URL + key could not be used). Plumb the endpoint through to
the @google/genai SDK via httpOptions.baseUrl:
- kosong: add baseUrl to GoogleGenAIOptions and inject it into the client's
httpOptions alongside the existing headers (the SDK merges headers and
overrides the base host).
- agent-core: forward provider.baseUrl in the google-genai and vertexai
branches, with GOOGLE_GEMINI_BASE_URL / GOOGLE_VERTEX_BASE_URL env
fallback. vertexai keeps deriving location from an aiplatform host.
- docs: document base_url for both providers, noting the host root only
must be given because the SDK appends the API version itself.
Covered by unit tests asserting the URL reaches the kosong config and the
SDK client's httpOptions.
* fix(provider): use the effective base_url for vertex location detection
The vertexai branch forwarded the endpoint from config `base_url` OR the
GOOGLE_VERTEX_BASE_URL env fallback, but service-account detection
(`hasVertexAIServiceEnv` / `vertexAILocation`) still derived the region from
`provider.baseUrl` only. Supplying the regional endpoint via the env fallback
(with a project but no explicit GOOGLE_CLOUD_LOCATION) therefore left location
undefined and silently downgraded Vertex ADC to API-key Gemini routing.
Resolve the effective base URL once and use it for both forwarding and location
derivation, so the env fallback behaves exactly like `base_url`. Add a
changeset for the kosong + agent-core patch release.
* feat(agent-core): slim WebSearch to query-only; fetch page content via FetchURL
The coding search backend no longer honors `limit`/`enable_page_crawling` and
always returns full page content, which was token-heavy and frequently
truncated. Realign the web tools around a search+fetch split:
- WebSearch request sends only `text_query`; the tool exposes only `query`.
- Search results drop inline page content and add the source site; full page
content is now fetched on demand via FetchURL.
- Add a citation reminder to both WebSearch and FetchURL results (in the
FetchURL front note so it survives body truncation).
- Update tool descriptions and reference docs accordingly.
* fix(agent-core): satisfy no-base-to-string lint and drop redundant | undefined
- Assert the exact serialized WebSearch request body instead of String()-coercing
a BodyInit, fixing the type-aware no-base-to-string lint error.
- Drop redundant `| undefined` from WebSearchResult optional fields per AGENTS.md.
- Add changeset.
* chore(changeset): bump agent-core to minor for WebSearch input-contract change
Removing the `limit`/`include_content` tool inputs tightens a closed
(`additionalProperties: false`) schema, so previously-valid args are now
rejected — an incompatible change for a released package. Bump minor rather
than patch.