mirror of
https://github.com/MoonshotAI/kimi-code.git
synced 2026-08-15 11:45:52 +00:00
371 commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
2cd78e73ef | ci: release packages | ||
|
|
d96cd03770
|
fix(cli): pre-send warning and clearer error for over-long /goal objectives (#2928)
Some checks are pending
CI / test (2) (push) Waiting to run
CI / test (3) (push) Waiting to run
CI / test (4) (push) Waiting to run
CI / test (5) (push) Waiting to run
CI / test-pi-tui (push) Waiting to run
CI / test-vscode-legacy (push) Waiting to run
CI / test-windows (push) Waiting to run
CI / lint (push) Waiting to run
CI / typecheck (push) Waiting to run
CI / build (push) Waiting to run
CI / test (1) (push) Waiting to run
Nix Build / Check flake.nix workspace sync (push) Waiting to run
Nix Build / nix build .#kimi-code (push) Blocked by required conditions
Release / Release (push) Waiting to run
Release / Deploy docs (push) Blocked by required conditions
Release / Native release artifact (push) Blocked by required conditions
Release / Publish native release assets (push) Blocked by required conditions
* fix(cli): warn on over-long /goal objectives before sending and keep the input - Show a live footer warning in the TUI while a typed /goal objective exceeds the 4000-character limit, measuring paste-expanded text only when the input can be a /goal command. - Restore the rejected /goal input into the editor instead of losing it. - Include the file-reference workaround in the GOAL_OBJECTIVE_TOO_LONG error messages (TUI, goal queue, agent-core, agent-core-v2). * fix(cli): keep the /goal length warning in its own footer slot A transient hint (exit confirm, detach, image paste) that displaced the warning previously left the footer blank after clearing, because no editor change re-applied it. The footer now renders the warning from a dedicated slot whenever no transient hint is active, so the warning returns on its own. * fix(cli): gate the /goal length warning on trimmed text Submitted text is trimmed before slash-command dispatch, so leading whitespace still runs /goal — normalize with trimStart in the gate and in the length check to match. * fix(cli): restore input rejected by the slash-command busy gate An idle-only command submitted while streaming/compacting was rejected after the editor buffer had already been cleared, losing hand-typed input (e.g. an over-long /goal objective that never reached the local validation). * fix(cli): restore input at the post-creation busy re-check The lazy-session race rejects an idle-only command after a first prompt has already started a turn; the editor buffer is long cleared by then, so give the submitted input back like the dispatch blocked branch does. * fix(cli): close the remaining input-loss and gate gaps around /goal - Restore the submitted input when lazy session creation fails before a session-requiring command runs. - Restore only into a still-empty editor after async gates, so a draft typed while creation was pending is never overwritten. - Expand pastes that can complete a partially typed /goal command (e.g. /go[paste #1 …]) in the length-warning gate. * fix(cli): never displace newer UI state with a delayed input restore A session-less /goal submission restores its input only after an async gap (lazy session creation). If the user opened an editor-replacement panel meanwhile, restoring would tear it down (and leave activeDialog inconsistent). Track editorReplacementMounted in TUIState and gate all delayed restores through canRestoreSubmittedInput. * refactor(cli): move canRestoreSubmittedInput into commands/resolve Avoids the goal.ts <-> dispatch.ts runtime import cycle flagged by import/no-cycle; the helper takes a structural host shape instead. * fix(cli): match the slash parser's delimiter in the /goal length warning parseSlashInput splits the command name at a literal space only, so a newline or tab after /goal dispatches as a plain message — the warning must not fire for inputs the dispatcher will not treat as a goal. |
||
|
|
13d86f8b7b
|
ci: release packages (#2881)
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> |
||
|
|
102984aa66
|
fix: settle cancelled MCP OAuth callbacks (#2899)
Some checks are pending
CI / build (push) Waiting to run
CI / test (1) (push) Waiting to run
CI / test (2) (push) Waiting to run
CI / test (3) (push) Waiting to run
CI / test (4) (push) Waiting to run
CI / test (5) (push) Waiting to run
CI / test-pi-tui (push) Waiting to run
CI / test-windows (push) Waiting to run
CI / lint (push) Waiting to run
CI / typecheck (push) Waiting to run
Nix Build / Check flake.nix workspace sync (push) Waiting to run
Nix Build / nix build .#kimi-code (push) Blocked by required conditions
Release / Native release artifact (push) Blocked by required conditions
Release / Publish native release assets (push) Blocked by required conditions
Release / Release (push) Waiting to run
Release / Deploy docs (push) Blocked by required conditions
Co-authored-by: yuchengzhen <yuchengzhen@moonshot.cn> |
||
|
|
c9bfe8b2c8
|
feat: replace the secondary-model experiment with a declarative subagent model pool (#2700)
* feat: replace secondary-model experiment with [subagent.models] pool
Add a declarative subagent model pool to agent-core-v2: [subagent.models]
maps [models] entry ids to selection hints rendered in the Agent/AgentSwarm
tool descriptions, and [subagent].default_model picks the spawn model when
the caller passes none. The tools' model parameter becomes a free-form
alias string (stripped when no pool is configured), description rendering
is caller-aware (primary (alias) [main model]), and a session-start
validation service fails fast with CONFIG_INVALID on a missing/invalid
default_model or an unresolvable pool alias.
Remove the secondary-model experiment from the v2 engine, node-sdk,
kap-server, and the TUI (the /secondary_model command), and drop the
agent-profile modelPreference / model_preference frontmatter field on v2.
The legacy v1 engine keeps the experiment unchanged; v2 ignores leftover
[secondary_model] config silently.
* fix(agent-core-v2): harden subagent model-pool validation and error/picker mapping
Deep-review follow-ups to the [subagent.models] pool:
- validate the pool before session materialization (after config.ready)
and before the fork file copy, so a broken pool no longer leaves
orphaned session dirs or leaked MCP overlay connections; the
Session-scope validation service stays as a backstop
- reject the reserved "primary" pool alias at startup, and again
defensively in resolveSubagentBinding so a pool broken by a runtime
config edit fails loudly at spawn instead of binding the wrong model
- keep the [default] marker when the caller's own model is the pool
default (primary (alias) [main model] [default])
- recompile the cached tool-args validator when a tool advertises a new
schema object (mid-session pool edits no longer hit a stale validator)
- map config.invalid to VALIDATION_FAILED in kap-server's session routes,
the debug transport mapper, and the catch-all error handler
- hide the v1-synthesized __secondary__ entry from the /model and
/provider pickers again
- fold per-export doc blocks into file headers per package comment
conventions; add pre-flight/reserved-key/validator/mapping tests and
document that create/resume/fork all fail on a broken pool
* feat: re-add /secondary_model and accept a lone subagent default_model
- v2 engine: a pool-less [subagent] default_model forms an implicit
single-entry pool — validated at session create/resume/fork like an
explicit pool, and advertised through the Agent/AgentSwarm model
parameter.
- Tool descriptions: the caller's own alias is a normal pool entry
marked [main model]; the primary line stays distinct because only it
inherits the caller's thinking level.
- TUI: /secondary_model returns, persisting [subagent] default_model
(merging into an existing pool with an empty description); the picker
hides the no-op Thinking footer and rejects the reserved primary
alias.
- kap-server: /api/v1/config accepts and echoes subagent; the
snake-to-camel patch conversion preserves user-defined map keys under
providers/models/experimental/raw without leaking preserve mode into
a colliding alias's own fields.
- v1 config schema learns subagent.defaultModel/models so the shared
config.toml round-trips; the v1 engine still ignores them at runtime.
- Docs (en/zh) and changesets updated.
* docs: use public model identifiers in the subagent model pool examples
* refactor: rename /secondary_model to /secondary-model
* test: cover the /secondary-model command name resolution
* Revert "test: cover the /secondary-model command name resolution"
This reverts commit
|
||
|
|
504e6292ed
|
feat(mcp): inspect effective authorization state in v1 (#2856)
* feat(mcp): inspect effective authorization state * test(agent-core-v2): register MCP auth coordinator fixture * fix(mcp): validate runtime names against full catalog * fix(mcp): reconnect after pending auth updates * docs(mcp): describe auth coordinator collaborator * fix(mcp): ignore disabled runtime name collisions * fix(mcp): serialize OAuth token refresh * test(mcp): await OAuth credential writes * fix(mcp): queue trailing credential reconnect * fix(oauth): preserve access-only refresh winners * fix(mcp): preserve legacy offline auth state * fix(mcp): redact inspection credentials * refactor(mcp): keep app inspection on v1 * fix(mcp): guard legacy auth status mutations * fix(mcp): avoid deterministic legacy auth probes * fix(mcp): cover initialization credential updates --------- Co-authored-by: 刘仲诺 <liuzhongnuo@dev.msh.team> |
||
|
|
3b0936d8e0
|
fix(agent-core): only discover the root SKILL.md for plugin skills fallback (#2847)
When a plugin manifest omits `skills` and the plugin root contains a SKILL.md, the fallback treated the whole plugin root as a generic skill scan directory, so sibling Markdown files such as CHANGELOG.md were misidentified as skills and inflated the plugin skill count. Mark the fallback root as root-skill-only so discovery parses only the root SKILL.md; explicit `skills` entries (including "./") keep the directory scan semantics. Applied to both agent-core and agent-core-v2. |
||
|
|
101c4d1997
|
feat(agent-core-v2): remove Agent and AgentSwarm from builtin profile tool lists (#2837)
* feat(agent-core-v2): remove Agent and AgentSwarm from builtin profile tool lists The builtin agent and coder profiles no longer expose the Agent and AgentSwarm tools, so sessions on the v2 engine do not offer subagent delegation by default. The tools themselves remain registered; profiles that list them explicitly can still opt in. * feat(agent-core): remove Agent and AgentSwarm from builtin profile tool lists Align the v1 builtin agent/coder profiles with the v2 change: the default profiles no longer offer subagent delegation, while the tools stay registered for profiles that list them explicitly. The parity projection drops v1's inactive Agent/AgentSwarm roster entries: v1 reports registered-but-inactive builtin tools where v2 only registers the tools a profile lists, so an inactive entry has no v2 counterpart. Active entries still compare in full. * fix: keep Agent and AgentSwarm in the builtin agent profile Scope the removal to the coder subagent profile on both engines: the main agent keeps Agent/AgentSwarm so default sessions can still delegate, while coder subagents no longer spawn nested subagents by default. Snapshots and token counts shift only for the embedded coder tool list; the v1 parity projection needs no change since the main agent rosters match again. |
||
|
|
01c74e9372
|
fix(agent-core): isolate builtin profile catalogs per session (#2740)
Some checks failed
CI / build (push) Has been cancelled
CI / test (1) (push) Has been cancelled
CI / test (2) (push) Has been cancelled
CI / test (3) (push) Has been cancelled
CI / test (4) (push) Has been cancelled
CI / test (5) (push) Has been cancelled
CI / test-pi-tui (push) Has been cancelled
CI / test-windows (push) Has been cancelled
CI / lint (push) Has been cancelled
CI / typecheck (push) Has been cancelled
Nix Build / Check flake.nix workspace sync (push) Has been cancelled
Release / Release (push) Has been cancelled
Release / Native release artifact (push) Has been cancelled
Nix Build / nix build .#kimi-code (push) Has been cancelled
Release / Deploy docs (push) Has been cancelled
Release / Publish native release assets (push) Has been cancelled
|
||
|
|
437a1b8ba1
|
fix(sdk): probe MCP auth status through connection (#2731)
Some checks are pending
CI / build (push) Waiting to run
CI / test (1) (push) Waiting to run
CI / test (2) (push) Waiting to run
CI / test (3) (push) Waiting to run
CI / test (4) (push) Waiting to run
CI / test (5) (push) Waiting to run
CI / test-pi-tui (push) Waiting to run
CI / test-windows (push) Waiting to run
CI / lint (push) Waiting to run
CI / typecheck (push) Waiting to run
Nix Build / Check flake.nix workspace sync (push) Waiting to run
Nix Build / nix build .#kimi-code (push) Blocked by required conditions
Release / Release (push) Waiting to run
Release / Deploy docs (push) Blocked by required conditions
Release / Native release artifact (push) Blocked by required conditions
Release / Publish native release assets (push) Blocked by required conditions
|
||
|
|
0b2e803d5e
|
feat(sdk): expose global MCP auth status (#2706)
Some checks are pending
CI / build (push) Waiting to run
CI / test (1) (push) Waiting to run
CI / test (2) (push) Waiting to run
CI / test (3) (push) Waiting to run
CI / test (4) (push) Waiting to run
CI / test (5) (push) Waiting to run
CI / test-pi-tui (push) Waiting to run
CI / test-windows (push) Waiting to run
CI / lint (push) Waiting to run
CI / typecheck (push) Waiting to run
Nix Build / Check flake.nix workspace sync (push) Waiting to run
Nix Build / nix build .#kimi-code (push) Blocked by required conditions
Release / Native release artifact (push) Blocked by required conditions
Release / Publish native release assets (push) Blocked by required conditions
Release / Release (push) Waiting to run
Release / Deploy docs (push) Blocked by required conditions
|
||
|
|
02c026d487
|
feat(agent-core-v2): tombstone removed MCP servers and freeze plugin prompt inputs (#2694)
* feat(mcp): tombstone removed MCP servers and apply plugin changes immediately (20 files) - add 'removed' MCP server status: workspace config removals call markRemoved instead of remove, keeping tool registrations alive while short-circuiting calls with a removal notice - fire onDidReload after every plugin mutation (install/enable/disable/remove) so workspace consumers refresh contributions immediately - TUI renders the removed status in the MCP panel/startup summary and shows an apply-immediately hint on the v2 engine * feat(agent-core-v2): freeze plugin prompt inputs for live agents (2 files) - snapshot the model skill listing and plugin system-prompt sections on the first successful prompt build and reuse the frozen values for the agent's lifetime, so plugin install / enable / disable / remove / reload never rewrites a live agent's prompt (same keep-live-sessions-stable philosophy as the MCP tombstone) - freeze only on success: a not-yet-ready skill catalog or a failed enabledSystemPrompts() read must not pin empty values for the agent's lifetime - refreshSystemPrompt still rebuilds on catalog change events but reuses the frozen values, so the prompt only moves when non-plugin inputs change (AGENTS.md, [tools] section, session tool policy, compaction); new agents snapshot the then-current state * chore(changeset): add changesets for MCP tombstone and frozen plugin prompt inputs * docs: describe immediate plugin changes and the removed MCP status on the v2 engine * fix(klient): mirror the removed MCP server status in the wire contract * docs: drop the legacy-engine behavior notes from the plugin and MCP pages * fix(agent-core-v2): freeze plugin sections only on a loaded snapshot - enabledSystemPrompts() resolves to its consumption fallback (never rejects) while the initial plugin load has failed; freezing that empty read locked plugin sections out of the live agent even after a later successful reload - expose hasLoadedSnapshot() on IPluginService so resolvePluginSections can tell a real empty snapshot from the fallback before freezing |
||
|
|
7b2784b9b7
|
feat: surface the bound model and thinking effort on subagent UIs (#2679)
* feat: surface the bound model on subagent UIs The subagent.spawned event now carries the display-normalized model alias (the derived __secondary__ entry resolves to its base alias), so clients can show which model a subagent is bound to. The TUI subagent card, swarm panel header, and background-agent entry show it at spawn; the WS snapshot roster and REST /tasks (background/detached subagents) carry it too, keeping the model visible across client reconnects. * feat: carry the subagent thinking effort alongside the model The spawned event, snapshot roster, and REST /tasks now also carry the child's effective thinking effort (read from the child profile at spawn, the same vocabulary as agent.status.updated). UIs show it only when it diverges from the main session's current effort — an inherited level adds no information, and 'off' is never shown. * feat(tui): show the bound model and effort in the /tasks browser The task browser's Detail pane renders Model and Effort rows for agent tasks (raw alias and level — it is the inspector surface, so no diff filtering), and its minimum height grows to fit the new rows. The values were already persisted on SubagentTaskInfo; the TaskInfo union, its zod schemas (protocol, kap-server, klient contract), and the v1 type declaration now carry them so nothing strips them in transit. * feat(tui): show concrete subagent effort levels unconditionally Display rule simplified: any concrete effort tier (low/high/max/…) is shown next to the model — including when it matches the main session's level. Only the boolean states stay hidden: 'off' (no thinking) and 'on' (generic thinking) carry no level information. * docs: trim the changeset entry * fix(tui): keep the model and effort on background-agent entries across resume replayBackgroundProjection only copied agentId/parentToolCallId/ description, so a background subagent that outlived a resume lost its model/effort on the later terminal transcript entry. The projection now threads the persisted values (catalog-mapped model; boolean effort states dropped), and session replay passes the loaded model catalog through. * fix(agent-core-v2): normalize the derived secondary alias regardless of the flag A child bound while the secondary-model experiment was on keeps __secondary__ in its persisted binding; if the flag is later switched off with the recipe still configured, resolveSecondaryModel() gated the normalization and the sentinel leaked back onto resumed subagents. subagentDisplayModel now reads the recipe straight from config (the flag gates new bindings, not the interpretation of existing ones), which also drops SessionSwarmService's now-unused IFlagService dependency. Also adds the SDK package to the release: the new SubagentSpawnedEvent/AgentTaskInfo fields are SDK-visible types. * fix(agent-core-v2): normalize the status-frame model at the source A derived-bound child republishes agent.status.updated right after spawn with its raw modelAlias, which overwrote the spawned event's normalized display model on single-subagent cards (swarm headers were first-wins and escaped). emitStatusUpdated now maps through subagentDisplayModel, a no-op for the never-derived main agent. Also moves the inline comments added by this branch into top-of-file headers per the v2 comment convention. * fix(tui): clamp the /tasks detail frame to the available body At terminals near the minimum height the forced 10-row detail frame overflowed the body and truncated the preview frame's border. The detail height now caps out at whatever leaves the preview its borders plus one content row, with a regression test at exactly MIN_HEIGHT. * fix: normalize inherited derived aliases and keep model/effort on replayed terminal entries - resolveSubagentBinding's caller-fallback branch also maps through subagentDisplayModel: a caller itself bound to the derived entry (a resumed subagent making a nested Agent call) no longer publishes __secondary__. - The replayed background-task terminal notification builds its metadata with the persisted model (catalog-mapped) and concrete effort, matching the live completion path. - Drops the inline comments this branch added inside v2 test bodies; the scenario context lives in the source file headers. |
||
|
|
f881cdd970
|
feat(cli): default CLI surfaces to the agent-core-v2 engine (#2627)
Some checks are pending
CI / lint (push) Waiting to run
CI / typecheck (push) Waiting to run
CI / build (push) Waiting to run
CI / test (1) (push) Waiting to run
CI / test (2) (push) Waiting to run
CI / test (3) (push) Waiting to run
CI / test (4) (push) Waiting to run
Nix Build / Check flake.nix workspace sync (push) Waiting to run
Nix Build / nix build .#kimi-code (push) Blocked by required conditions
Release / Release (push) Waiting to run
Release / Deploy docs (push) Blocked by required conditions
Release / Native release artifact (push) Blocked by required conditions
Release / Publish native release assets (push) Blocked by required conditions
CI / test (5) (push) Waiting to run
CI / test-pi-tui (push) Waiting to run
CI / test-windows (push) Waiting to run
* feat(cli): default to agent-core-v2 engine with KIMI_CODE_LEGACY_FLAG opt-out - invert the engine gate: isKimiV2Enabled() now returns true unless KIMI_CODE_LEGACY_FLAG is truthy; KIMI_CODE_EXPERIMENTAL_FLAG no longer selects the engine - replace the experimental `kimi acp-v2` command with the native v2 implementation as the default `kimi acp`; the legacy acp-adapter path remains under the legacy flag - drop the acp-v2 experimental flag from the registry - rename the dev:cli:v2 script to dev:cli:legacy - update en/zh docs for the new default engine and the legacy flag * feat(cli): route export and provider through the engine gate - select the harness via isKimiV2Enabled(): agent-core-v2 by default, the legacy harness when KIMI_CODE_LEGACY_FLAG is truthy - close the harness after each one-shot command so the v2 engine's watchers do not keep the process alive - document both commands in the KIMI_CODE_LEGACY_FLAG env-var entry |
||
|
|
2a4990182d
|
docs: slim root AGENTS.md, move deep package docs to package guides (#2626)
- move the kimi-inspect, kap-server, transcript, and minidb project-map entries into new package-level AGENTS.md files - move the standalone Agent class rule to packages/agent-core/AGENTS.md - merge root-only agent-core-v2 facts (MCP persistence, trust routes, seed contract examples) into packages/agent-core-v2/AGENTS.md - compress the remaining long project-map entries to summaries with pointers, and add an anti-bloat rule for map entries - drop the stale server-e2e entry; the package no longer exists |
||
|
|
2ee6e43124
|
fix(mcp): re-register the OAuth client when its redirect URI no longer matches (#2620)
The callback listener binds a random port per flow, while DCR registration records the redirect URI of the flow that created it — so every interactive authorization after the first was rejected with "Invalid redirect URI", an error rendered only in the user's browser while the client waited for a callback that never came. Detect the mismatch before invoking auth() and drop the stale registration so the flow re-registers with the current callback URI (v1 + v2). Resolve #2606 Co-authored-by: zouying <zouying@moonshot.cn> |
||
|
|
74c321e4c6
|
fix(mcp): drop protocol-reserved _meta keys from model-visible output (#2600)
* fix(mcp): drop protocol-reserved _meta keys from model-visible output Follow-up to #2596. The MCP spec reserves _meta key prefixes whose labels include "modelcontextprotocol" or "mcp" for protocol use; those entries carry host/protocol plumbing rather than model-facing data, so filter them out before serializing the <mcp-structured-result> block. Unprefixed and vendor-prefixed keys still pass through — their semantics belong to the server. Also moves the v2 implementation commentary into the module header per the agent-core-v2 comment convention. * fix(mcp): reserve _meta prefixes only when a label follows mcp/modelcontextprotocol Per the spec's key-name rules a prefix is reserved when a modelcontextprotocol or mcp label is followed by at least one more label; a trailing reserved word (com.example.mcp/) is a legitimate vendor namespace and now passes through. --------- Co-authored-by: zouying <zouying@moonshot.cn> |
||
|
|
3126422757
|
feat(mcp): carry an absolute expiresAt on the OAuth authorization-url update (#2609)
* feat(mcp): carry an absolute expiresAt on the OAuth authorization-url update The authenticate flow waits for the OAuth callback with a fixed budget, but the authorization-url tool update surfaced to embedding hosts did not say when that window ends — hosts had to hardcode a mirror of the 15-minute constant to render countdowns. Include the absolute deadline (now + effective wait timeout) in the update payload for v1 and v2. Resolve #2607 * fix(protocol,kap-server): accept expiresAt in the OAuth authorization-url update schemas The zod validators mirrored the pre-expiresAt payload shape and would strip the new field at the kap-server boundary. --------- Co-authored-by: zouying <zouying@moonshot.cn> |
||
|
|
c32e661faa
|
fix(mcp): pass structuredContent and _meta through to model-visible tool output (#2596)
MCP tool results were narrowed to {content, isError}, dropping the
spec-defined structuredContent field and _meta metadata. Servers that
return structured contracts in these fields (validated against
outputSchema, or namespaced metadata such as browser-handoff payloads)
were invisible to the agent. Surface them as a serialized
<mcp-structured-result> block appended to the tool output, still subject
to the existing text budget.
Co-authored-by: zouying <zouying@moonshot.cn>
|
||
|
|
1328b32037
|
feat(acp): add experimental agent-core-v2 ACP server (kimi acp-v2) (#2571)
* feat(acp): add agent-core-v2 ACP server - add ACP session lifecycle, configuration, permissions, and event bridging - expose the experimental kimi acp-v2 command with terminal authentication - add integration coverage and workspace build configuration * test: use neutral example domains in test fixtures and docs - replace placeholder hostnames (evil.com, foo.com, internal.corp, real.corp) with example.test / example.com in agent-core-v2 and kap-server tests - replace fixture emails (x@y.com, a@x.com) with example addresses in minidb tests and README * fix(acp): align acp-server with agent-core-v2 interfaces and address review - add missing appendText to AcpHostFileSystem (IHostFileSystem drift) - replace IAgentPromptService.prompt with inject - use Turn.cancel() instead of abortController - gate FS reverse-RPCs on client capabilities, fallback to local FS - return PROTOCOL_VERSION constant instead of echoing client version - remove misleading mcpCapabilities from initialize response - dispose old session wrapper before replacing on load/resume - fix object stringification lint error in convert.ts - add acp-v2 to expected CLI sub-command list in test * fix(acp): use enqueue for prompt submission, stop advertising unimplemented builtins - replace IAgentPromptService.inject with enqueue so onBeforeSubmitPrompt hooks (prompt-blocking policy) are not bypassed - stop advertising builtin slash commands (/help, /status, etc.) until builtin command execution is implemented - add comment explaining appendText stays local (ACP has no append RPC) - update skills test to match new availableCommands behavior * fix(acp): filter turn events by turnId, surface auth failures as auth_required - track turnId in driveTurn and ignore events from unrelated turns, preventing queued prompts from settling on the running turn - reject prompt requests with auth_required when turn fails with an auth-related error code, enabling ACP client re-auth flow * fix(acp): gate acp-v2 behind experimental flag, filter sessions by cwd - add acp-v2 experimental flag (KIMI_CODE_EXPERIMENTAL_ACP_V2) and gate CLI command registration behind it - filter session/list results by requested cwd instead of returning sessions from all workspaces - detect hook-blocked prompts via PromptHandle.state and add TODO for streaming block messages once the hook context exposes them * refactor(acp-server): rewire ACP server onto the klient facade - replace direct agent-core-v2 scope/service access (ISessionLifecycleService, ISessionIndex, IEventBus, ISessionInteractionService, etc.) with the Klient facade: klient.global.sessions / klient.session(id) / agent('main') handles - drive turns via agent.prompt() + session-level agent event subscriptions instead of per-prompt IEventBus wiring; settle on turn.ended - route approval/question bridging through session.interactions events - hide the thinking config option and skill catalog behind KLIENT-GAP markers until klient exposes those surfaces - acp-fs: pass realpath through to the local inner backend - klient: session.restore() rejects both null and undefined handles * feat(agent-core-v2): add session delete and ephemeral per-session MCP servers - add ISessionLifecycleService.delete: close a live session first, then remove its persisted data, evict the index read-model entry, and append a deleted tombstone to session_index.jsonl; unknown ids raise session.not_found - add CreateSessionOptions/ResumeSessionOptions.mcpServers: session-owned MCP overlay merged over the workspace manager via MergedMcpConnectionView (an ephemeral name shadows a workspace server), never persisted, released when the session scope tears down - return PromptLaunchResult from activateSkill so callers get the launched turn id and activation failures (unknown skill, busy) surface - add ISessionSkillCatalog.list() as a wire-friendly catalog snapshot - add ISessionIndex.remove for read-model eviction on delete * feat(klient): expose session delete, per-session MCP, skills, and stream events - session lifecycle contract: delete, resume/restore options, and CreateSessionOptions.mcpServers (ephemeral per-session MCP servers) - add the session skills contract and facade accessors for the wire-friendly skill catalog snapshot - register tool.call.delta, tool.progress, and compaction.* agent stream events so consumers can subscribe with typed payloads * feat(acp-server): align ACP v2 server with acp-adapter capabilities - complete the klient-facade rewire: ACP client connection holder and the terminal/* reverse-RPC runner routed through the Agent scope - negotiate the protocol version on initialize instead of pinning v1 - compress oversized prompt images at the ACP ingestion point with a format gate, caption, and persisted originals; a cancel arriving mid-compression settles the prompt as cancelled without a turn - stream tool call args via tool.call.delta (lazy pending create, cumulative replace, started upgrade) and refresh titles via tool.progress status updates - report compaction progress and results after /compact via the compaction.* events - answer unknown slash commands locally instead of sending them to the model - accept legacy "<id>,thinking" model ids and legacy approve / approve_for_session approval option ids - keep sessions without cwd metadata in cwd-filtered session/list - sanitize wire errors: auth codes map to auth_required, turn.agent_busy to invalid_request, everything else to a fixed internal-error message - bump @agentclientprotocol/sdk to ^1.3.0 * fix(cli): drop stale registerServerCommand call and sherif ACP SDK split - commands.ts called registerServerCommand, which no longer exists on current main (the deprecated `kimi server` shim is registered via registerWebCommand), breaking typecheck, build, and every CLI test that builds the program - sherif rejects the @agentclientprotocol/sdk major split between acp-adapter (^0.23.0, production kimi acp) and acp-server (^1.3.0, experimental); the two hosts legitimately target different SDK majors, so ignore the dependency in the sherif invocation * test: update fixtures for acp-v2 flag and domain rename, refresh nix deps hash - kap-server origin.test: two CORS cases still used foo.com after the whitelist moved to foo.example.com, so the origin was no longer whitelisted and the expected CORS headers were withheld - node-sdk config.test: expect the new acp-v2 experimental flag in the harness feature metadata - flake.nix: update the fetchPnpmDeps hash for the @agentclientprotocol/sdk 1.3.0 lockfile change * fix(acp): widen the ACP v2 auth gate beyond OAuth-only providers The gate consulted only auth.summarize(), which iterates providers declaring an oauth section — configurations that authenticate with a plain apiKey or provider env-bag credentials (no OAuth at all) were rejected with auth_required even though the default model is fully usable. - klient: expose authSummaryService.ensureReady on the global auth facade (the contract already declared it) - acp-server: gate on the engine's own readiness probe for the default model — config apiKey / env-bag / OAuth token all count, matching how the model is actually used — and fall back to "any logged-in OAuth provider" (the legacy adapter's first branch) - test: an apiKey-only config passes the gate with auth enforcement on; the OAuth logout regression is unchanged * fix(acp): reject concurrent prompts instead of displacing the in-flight turn A second session/prompt while a turn is running overwrote the session's only TurnDriver: the engine quietly queues plain prompts submitted during an active turn (the launch resolves undefined, indistinguishable from a hook-blocked launch), so the first prompt never settled and both turns' events went unattributed. Guard both model-bound launch paths (plain prompt and skill activation) with a synchronous in-flight check and reject with invalid_request (turn.agent_busy), matching the legacy adapter's busy semantics. Local slash handling (builtins, unknown-command answers) is unaffected. |
||
|
|
98ef0f0b2f
|
fix(agent-core): replay v2 profile.bind records so resumed sessions keep their tools (#2567)
Some checks are pending
CI / build (push) Waiting to run
CI / test (1) (push) Waiting to run
CI / test (2) (push) Waiting to run
CI / test (3) (push) Waiting to run
CI / test (4) (push) Waiting to run
CI / test (5) (push) Waiting to run
CI / test-pi-tui (push) Waiting to run
CI / test-windows (push) Waiting to run
CI / lint (push) Waiting to run
CI / typecheck (push) Waiting to run
Nix Build / Check flake.nix workspace sync (push) Waiting to run
Nix Build / nix build .#kimi-code (push) Blocked by required conditions
Release / Release (push) Waiting to run
Release / Deploy docs (push) Blocked by required conditions
Release / Native release artifact (push) Blocked by required conditions
Release / Publish native release assets (push) Blocked by required conditions
* fix(agent-core): replay v2 profile.bind records so resumed sessions keep their tools Sessions created by the v2 engine (CLI 0.31+, wire protocol 1.5) persist the profile binding, including the tool allowlist, as a profile.bind record. The v1 replay path had no branch for it and silently dropped the record, so a session resumed by a v1 host (e.g. the VS Code extension via kimi-code-sdk) never called setActiveTools and sent requests with no tools at all (observed server-side as tools_count=0; the model emits reasoning only and stops with empty content). v1 replay now maps profile.bind onto config.update + setActiveTools when activeToolNames is an array, skips the record otherwise so the resume-time default-profile fallback still applies, and treats tools.reset_active_tools as a no-op. * fix(vis): handle v2 profile records in context projection * fix(agent-core): avoid synthetic replay for v2 profile binds * fix(vis): render v2 profile wire records |
||
|
|
6ba75a173b
|
feat(config): add deprecation mechanism and rename loop retry limit (#2572)
* feat(config): add deprecation mechanism and rename loop retry limit - agent-core-v2 config: declarative section `deprecations` (deprecated TOML keys are ignored and report a warning diagnostic; the file is never rewritten) and env binding `deprecatedEnv` (old var still resolves as a fallback with a warning), surfaced via the new `IConfigService.onDidChangeDiagnostics` event - loop_control: rename `max_retries_per_step` to `max_attempts_per_step` and `KIMI_LOOP_MAX_RETRIES_PER_STEP` to `KIMI_LOOP_MAX_ATTEMPTS_PER_STEP`; `max_steps_per_run` moves onto the same mechanism (no longer silently mapped) - kap-server: push the global `event.config.warning` WS event to every connection whenever the config warning set changes - TUI: show config diagnostics in warning yellow at startup instead of the dim startup notice - docs: config-files/env-vars (en+zh), regenerated config manifest, and the agent-core-dev config guide * feat(cli): validate config.toml against v2 section registry in doctor - add v2/validate-config.ts: validate config.toml with the agent-core-v2 ConfigRegistry, reporting registered-section schema failures as errors and unknown top-level keys / deprecated keys and env vars as non-fatal warnings - route `kimi doctor` config validation through the v2 validator when the KIMI_CODE_EXPERIMENTAL_FLAG master switch is on (lazy dynamic import, keeping the v2 module graph off the default path) - let doctor checks surface non-fatal warning messages on OK results * chore: downgrade loop-control changeset to patch |
||
|
|
95a656ca61
|
fix(agent-core): strip the no-op subagent model parameter while the secondary-model experiment is off (#2449)
The Agent/AgentSwarm tool schemas always advertised a \`model\` choice parameter, so the secondary-model concept entered the prompt even with the experiment disabled. Gate the advertised JSON schema on the flag in both engines: off (the default) drops the parameter, on keeps it, and spawn-time resolution already falls back to the caller's model either way. Also scrub ambient KIMI_CODE_EXPERIMENTAL_* env vars in both packages' vitest setup so flag-dependent tool schemas in llm.tools_snapshot stay deterministic regardless of the developer shell. |
||
|
|
bc28e9d802
|
ci: release packages (#2342)
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> |
||
|
|
40172c7ca9
|
feat: unify the host identity across OAuth, telemetry, and kap-server (#2382)
* refactor(oauth): make X-Msh-Platform an explicit host identity field X-Msh-Platform was hardcoded to kimi_code_cli in createKimiDeviceHeaders, so non-CLI hosts could not state their own platform and the desktop had to patch the header after the fact. KimiHostIdentity now carries a required platform (every host declares its own value; the CLI constant stays the fallback only for direct createKimiDeviceHeaders callers), and userAgentProduct is renamed to productName so the transport identity uses one name everywhere. All in-repo identity constructions pass platform explicitly; the wire value for CLI and VS Code hosts is unchanged (kimi_code_cli). * feat(agent-core-v2): carry the host identity in the bootstrap snapshot Replace the flat clientVersion field with a required clientIdentity (KimiHostIdentity) so every consumer reads the same host identity object: OAuthToolkitService now passes it to the OAuth toolkit, which means the OAuth device-flow endpoints (device authorization, token polling, refresh) on the kap-server path finally send the full X-Msh-* device headers instead of none, and the telemetry cloud appender reads client_version from the same source. A built-in CLI fallback keeps bare bootstrap() calls in tests working; composition roots must pass their own identity. The session export manifest grows an optional desktopVersion field (payload plumbed through; filled by kap-server in a follow-up). * feat(agent-core): thread the host identity into the managed auth facades The v1 managed auth facade constructed its OAuth toolkit without an identity, so token refreshes from inside the core went out without any X-Msh-* device headers. createManagedAuthFacade now takes an optional KimiHostIdentity and every call site supplies one: CoreProcessService._defaultOAuthTokenResolver forwards the core process's options.identity (the same source _defaultKimiRequestHeaders uses), and the DI-held services (oauth / auth summary / model catalog) read it from a new optional identity field on IEnvironmentService. The library-level "no identity, no device headers" contract is unchanged. * feat(kap-server)!: require the host identity and derive request headers from it ServerStartOptions.hostIdentity is now a required ServerHostIdentity (KimiHostIdentity + optional prompt display fields), replacing both the old optional HostIdentityOverrides (renamed to PromptIdentityOverrides, its productName field now displayName) and the version option (renamed to serverVersion — it is the engine version reported as server_version, while the host product version travels in hostIdentity.version). The server now feeds bootstrap's clientIdentity from hostIdentity and derives the default outbound headers (User-Agent + X-Msh-*) from it via createKimiDefaultHeaders, so kap-server-hosted OAuth flows and model / WebSearch requests carry the real host identity instead of a hardcoded kimi-code-cli fallback UA. Explicit header seeds still win as an escape hatch. Session export manifests record the host product version: kimiCodeVersion now carries hostIdentity.version (the engine version no longer appears), and desktop exports (desktop: true) are additionally stamped with a desktopVersion field. The instance registry keeps its host_version wire field for compatibility (kimi-inspect reads it); only the in-memory name changed to serverVersion. * feat(cli): wire the CLI host identity into the kimi web server kimi web now passes createKimiCodeHostIdentity(version) as the server's hostIdentity, so web-UI OAuth flows and the engine's outbound requests carry the explicit CLI identity (productName + version + platform). The explicit hostRequestHeadersSeed is dropped — kap-server derives the same headers from hostIdentity — and buildKimiDefaultHeaders goes away with its only consumer. * test(klient): drop clientVersion from the bootstrap contract parity list * chore: add changesets for the host identity unification * feat(cli): tag kimi web requests with a (web) User-Agent suffix kimi web shares the CLI product token and platform, so its outbound requests were indistinguishable from direct CLI runs upstream. Its host identity now carries userAgentSuffix 'web', putting web-UI traffic at kimi-code-cli/<version> (web) while X-Msh-Platform stays kimi_code_cli. * fix(klient): keep the env() clientVersion wire field after the bootstrap identity switch The bootstrap snapshot replaced the flat clientVersion scalar with clientIdentity, which broke klient's env() fan-out (RPCError: method not found). The wire surface keeps clientVersion — now sourced from clientIdentity.version — and bootstrapService gains a clientIdentity read (registered in envContract with an object schema) for consumers that want the full identity. * feat(oauth): send the product User-Agent on OAuth requests The OAuth endpoints used to receive only the X-Msh-* device headers (undici's default UA otherwise), which left the OAuth host unable to distinguish runtime surfaces — notably kimi web, whose platform matches the CLI and whose only distinguishing mark is the (web) UA suffix. The toolkit now feeds the full identity headers (User-Agent + X-Msh-*) into every device authorization, token polling, and refresh request; the request-header type widens from DeviceHeaders to OAuthRequestHeaders. * feat(vscode): report kimi_code_vscode as the extension's platform The VS Code extension inherited the CLI's hardcoded X-Msh-Platform value; with platform now an explicit identity field it declares its own, so the managed endpoints and OAuth host can tell extension traffic apart from CLI runs. * refactor(agent-core-v2)!: require the client identity at the composition root The bootstrap fallback identity fabricated a kimi-code-cli/unknown host for any caller that forgot to pass one — the same silent-misreport pattern this series set out to remove, and it made "required" a lie. BootstrapInput.clientIdentity is now required, so a missing identity fails at compile time instead of being papered over. Test and example callers pass a shared fixture (klient examples and test engines get one each); the node-sdk v2 client asserts its host identity with the oauth helper. Also folds DeviceHeaders from an interface into a type alias so it stays assignable to the widened OAuthRequestHeaders record. * feat(oauth)!: require and validate the platform in device headers Drops the quiet CLI fallback in createKimiDeviceHeaders (the same silent-misreport pattern removed from the bootstrap identity): platform is now a required option, validated with the same required-ASCII rule as the version — empty or all-non-ASCII values throw instead of emitting a blank X-Msh-Platform, and header-unsafe characters are stripped rather than sent raw. * fix(node-sdk): seed the host request headers on the v2 client path The interactive v2 engine path (experimental flag) bootstrapped without a hostRequestHeaders seed, so managed vendor calls went out with the SDK's default User-Agent (OpenAI/JS) and no X-Msh-* at all — v1 passes the full identity headers on the same requests. The v2 client now seeds the headers from its asserted host identity, and a test pins the seed. * chore: simplify the CLI changeset wording |
||
|
|
691ec4679e
|
fix: remove the blocking wait from the TaskOutput tool (#2379)
Some checks are pending
CI / build (push) Waiting to run
CI / test (1) (push) Waiting to run
CI / test (2) (push) Waiting to run
CI / test (3) (push) Waiting to run
CI / test (4) (push) Waiting to run
CI / test (5) (push) Waiting to run
CI / test-pi-tui (push) Waiting to run
CI / test-windows (push) Waiting to run
CI / lint (push) Waiting to run
CI / typecheck (push) Waiting to run
Nix Build / Check flake.nix workspace sync (push) Waiting to run
Nix Build / nix build .#kimi-code (push) Blocked by required conditions
Release / Release (push) Waiting to run
Release / Deploy docs (push) Blocked by required conditions
Release / Native release artifact (push) Blocked by required conditions
Release / Publish native release assets (push) Blocked by required conditions
* fix: remove the blocking wait from the TaskOutput tool The block/timeout parameters let a model stall the whole turn waiting for a background task (up to 3600s), even though completion already arrives via automatic notification. Remove both parameters from the v1 and v2 engines (kept in model-facing parity), simplify retrieval_status to success/not_ready, and update the tool, Bash, and Agent prompt wording plus user docs accordingly. Stale callers passing block are silently treated as a non-blocking snapshot. * fix: align background-task prompts with the non-blocking TaskOutput The compaction reminder promised TaskOutput could fetch a task's result for tasks that are still running, where it now returns not_ready — reword it to snapshot semantics and point at the completion notification. Also list AskUserQuestion(background=true) as a task source in the TaskOutput description. * test: exercise stale TaskOutput args through the runtime validator A stale block/timeout argument never reaches the tool: the executor's preflight validates args against the closed tool schema and rejects them immediately, so the old test documented silent-tolerance semantics the runtime never exhibits. Assert the real behavior through compileToolArgsValidator/validateToolArgs instead, and drop statement-adjacent comments to match the package's header-only comment convention. |
||
|
|
fa2c5ce18b
|
feat: support plugin-contributed custom agents (#2365)
* feat: support plugin-contributed custom agents * fix: await plugin loading before agent catalog * fix: refresh plugin agents on v1 reload * test(agent-core-v2): add enabledSystemPrompts to the plugin service stub |
||
|
|
02d77b20d9
|
feat(agent-core-v2): let plugins contribute system prompt instructions via the manifest systemPrompt field (#2314)
* feat(agent-core-v2): let plugins contribute system prompt instructions via the manifest systemPrompt field
* feat(agent-core-v2): add systemPromptPath to load plugin system prompt from a file
* docs: explain plugin system prompt templates
* fix(agent-core-v2): refresh plugin system prompts after changes
* fix(agent-core-v2): freeze restored profile bindings and converge plugin contributions at session scope
- restore no longer re-renders or re-persists prompts: a resumed agent
keeps its replayed profile binding (prompt and tool set) as persisted
- a new Session-level convergence point reloads plugin skills into the
session skill catalog before fanning out to every live agent prompt,
and every catalog-kind plugin mutation awaits the whole pipeline;
MCP-only toggles carry a distinct change kind and skip it
- live refreshes after a restart re-resolve the bound profile by name
and rebind the full slice (prompt, disallowed tools, active tools)
atomically, warning and keeping the persisted state when the profile
is gone; renders reuse the first-render timestamp and unchanged
prompts are not re-persisted, so convergence never churns the wire
- cap plugin system-prompt contributions (32 KB per field/file, 64 KB
aggregate per prompt build) with manifest diagnostics and warnings
- bump the changeset to minor: this is a new user-facing capability
* fix(agent-core-v2): register the new session domain and dedupe the missing-profile warning
- add sessionPluginContribution to the domain-layer registry so
lint:domain stays green
- emit system-prompt-refresh-profile-missing once per profile name,
matching the service's other deduped warnings
- document the convergence timeout escape hatch and the klient
exclusion of enabledSystemPrompts
* fix(agent-core-v2): dedupe the plugin budget warning and surface section read failures
- emit plugin-sections-oversized once per skipped-plugin signature
- let enabledSystemPrompts failures propagate to the refresh catch
(keeps the current prompt and warns) instead of silently rendering
and persisting a prompt without plugin instructions
- cover the convergence timeout cut-off with a fake-timers test
- clarify that the first-render timestamp anchors per process
* fix(agent-core-v2): serialize session convergence and restore onDidReload timing
- run at most one convergence per session and bound each change's wait
by the timeout, so a fan-out emitter never interleaves deliveries
after a timed-out convergence
- fire onDidReload as soon as the reload commits again, keeping hook
reloads independent of prompt convergence
- sign the plugin budget warning with an unambiguous key
* docs(agent-core-v2): align convergence wording with the serialized semantics
- the timeout retry promise only holds once stalled work clears
- note the per-session serial delivery cost model on the plugin change
contract and the dual-queue invariant on the service
* fix(agent-core-v2): keep empty plugin sections byte-neutral in the prompt template
- place ${plugin_sections} on the same template line as
${skills_section} so prompts without either block render exactly as
before this feature
- note on the change contract that waitUntil work must not call back
into plugin mutations, and spell out the per-session convergence
order in the user docs
* fix(agent-core-v2): pin a fork's profile so refresh triggers never rebind it
- applyBindingSnapshot left the fork with no pinned profile, which
routed in-process forks into the post-restart catalog rebind and
could reset an inherited tool set; forks now inherit the source
agent's pinned profile object
- pin the first-render timestamp reuse with a ${now}-embedding test
and document the anchored ${now} semantics
- tighten the plugin docs budget and resume-refresh wording
* fix(agent-core-v2): join in-flight convergence during agent bootstrap
- an agent created while a plugin convergence is in flight now waits
for it, and a restored agent refreshes once after it, so a plugin
mutation never straddles an agent's bootstrap
- warn on a non-string systemPrompt field and strip a UTF-8 BOM from
systemPromptPath files before trimming
- correct the consumption-surface wording (every CLI surface on the
experimental flag, not just kimi -p), the per-session queueing note,
and the single-plugin combined budget clause
* fix(agent-core-v2): bound the bootstrap convergence join by the timeout
A permanently wedged convergence kept convergeTail pending forever,
and the unconditional settled() wait in bindBootstrap would have
blocked every later agent creation in that session; the join now
races the shared convergence timeout and continues (a restored agent
still refreshes once, which never touches the tail), and the timeout
constant moves to the contract for reuse
* fix(agent-core-v2): close the convergence race against in-progress restores
- a convergence fan-out could land while an agent's wire log is still
replaying, dispatching a replay-visible config record whose effect
the rest of the replay then overwrites; refreshSystemPrompt now
skips while the wire restore is in progress
- convergence completion is tracked by a generation counter; bootstrap
compares it (after a bounded join) and refreshes a restored agent
exactly once when a round completed after its creation began,
replacing the wasConverging flag that could miss both windows
* fix(agent-core-v2): bound each convergence so a wedged participant cannot stop the pipeline
- the fan-out now races the convergence timeout, so convergeTail always
settles: a permanently hung refresh delays its round (blocked entries
drain oldest-first on later changes) instead of killing the session's
convergence for good
- warn when agent bootstrap stops waiting on a stalled convergence
- diagnose a blank systemPromptPath and pin the plugin-root escape
guard with traversal, absolute-path, and symlink tests
* fix(agent-core-v2): bound the skill reload, preserve user-tool overlays, roll the prompt clock daily
- the convergence's skill-reload segment now races the same timeout as
the fan-out, so no segment of the pipeline can wedge a session for
good; it continues with the previous catalog and retries next change
- a cold rebind that resets the tool set replays session-added user
tools onto the new base instead of dropping them for the rest of the
process
- the rendered timestamp re-anchors when the UTC date rolls over, so
long-lived processes keep a fresh clock while steady-state renders
stay byte-stable within a day
- the plugin budget warning dedupes per plugin id, and the docs note
that systemPromptPath content is frozen until the next reload
* feat(agent-core-v2): converge cold plugin changes on resume through a drift-free gate
- restore replays the persisted binding untouched, then bootstrap
refreshes only when drift-free inputs changed while the session was
cold: the catalog profile's tool set/denylist, or the plugin-sections
baseline persisted alongside the prompt on the existing bind/update
payloads; directory-listing and date drift wait for live triggers,
so quiet resumes append no replay-visible records
- the rendered timestamp is day-precision (UTC date at 00:00,
re-anchored on rollover), keeping steady-state renders byte-stable
across resumes and sessions on the same day
- consolidate both timeout helpers onto a shared raceOutcome, and drop
the generation counter the gate supersedes
- align the plugin-sections precedence prose with the AGENTS.md
disclaimer (no self-granted authority, system instructions win on
conflict)
* fix(agent-core-v2): bound the restored-prompt gate and land the sections baseline
- the gate's plugin-sections read now races the convergence timeout, so
agent creation never blocks behind an unrelated plugin mutation
- refreshes serialize per agent through a tail, so overlapping triggers
cannot write prompts out of order
- when plugin sections change but a plugin-free custom prompt does not,
the new baseline lands as a sections-only update instead of making
every later resume re-render in vain
- align the system prompt's Date and Time paragraph with the
day-precision anchored timestamp
* Update plugin system-prompt instructions in changeset
Live sessions pick up plugin changes, while the default TUI and `kimi -p` paths ignore these fields.
Signed-off-by: 7Sageer <sag77r@hotmail.com>
* refactor(agent-core-v2): keep plugin skill reload user-driven
Plugin mutations still converge live agent prompts, but the session
skill catalog goes back to refreshing only on explicit plugin reload,
as before: the prompt feature does not need skill convergence, and the
pre-existing manual-reload semantics stay uniform across all plugin
contributions. Removes the convergence-driven skill reload, the
reloadSource de-privatization, and their tests; restores the
PluginSkillSource onDidReload forwarding and its catalog tests.
* refactor(agent-core-v2): apply plugin system-prompt changes only on explicit reload
Drop the live convergence machinery (the plugin onDidChange barrier,
the sessionPluginContribution fan-out, the restored-prompt drift gate,
and the day-precision render clock) so plugin system-prompt sections
take effect at the same point as every other plugin contribution:
/plugins reload or a new session. The profile now refreshes when the
session skill catalog re-pulls its plugin source on reload, reading
both the skill list and the prompt sections fresh.
* feat(agent-core): let plugins contribute system prompt instructions via the manifest systemPrompt field
* chore(agent-core-v2): remove inline implementation comment
* docs: clarify plugin prompt refresh semantics
---------
Signed-off-by: 7Sageer <sag77r@hotmail.com>
|
||
|
|
efac96c8a9
|
feat(agent-core): custom agent files and secondary model on the v1 engine (#2232)
* feat(agent-core): custom agent files and secondary model on the v1 engine
Migrate the custom agentfile and secondary-model capabilities from
agent-core-v2 to the v1 engine so they work in the TUI and plain
kimi -p sessions:
- discover Markdown agent files from user/project/extra/explicit
directories with the v2 precedence rules, a merged session profile
catalog replacing the hardcoded builtin profile lookups, SYSTEM.md
main prompt override, and ${base_prompt} backed by the effective
default
- --agent/--agent-file now work in print mode on the default engine;
CreateSessionOptions gains agentProfile/agentFiles
- [secondary_model] config + KIMI_SECONDARY_MODEL/EFFORT bind newly
spawned subagents to a cheaper model behind the secondary-model
experiment flag, with primary/secondary model params on Agent and
AgentSwarm and upfront session warnings
- full disallowedTools deny semantics (exact names + mcp__ globs)
evaluated by the tool manager and persisted in the agent wire
* fix(cli): guard optional agentFiles in the prompt runner
runPrompt is also driven programmatically (headless goal flow) with
options that never pass through the CLI parser defaults, so agentFiles
can be undefined; mirror the addDirs optional-chaining pattern. Also
extend the SDK experimental-feature assertion with the secondary-model
flag.
* fix(agent-core): preserve custom agent bindings on v1
* fix(agent-core): narrow secondary model error hints
* fix(agent-core): persist custom agent profile bindings
* Delete .changeset/sdk-agent-profile-options.md
Signed-off-by: 7Sageer <sag77r@hotmail.com>
* Update v1-custom-agent-files.md
Signed-off-by: 7Sageer <sag77r@hotmail.com>
* Update v1-secondary-model.md
Signed-off-by: 7Sageer <sag77r@hotmail.com>
* Update v1-custom-agent-files.md
Signed-off-by: 7Sageer <sag77r@hotmail.com>
* fix(agent-core): keep SYSTEM.md a prompt-only overlay for delegation
* docs: update agent file and secondary model availability wording
* fix(cli): reject --agent-file combined with session resume
The resume path only forwards the agent file's name for the bound-profile
assertion; the file's content is never re-applied (the session keeps its
creation-time catalog snapshot). Previously the combination was silently
accepted, so an edited file (or a same-named one) appeared to apply but did
not. Reject it at option validation and document the constraint.
* refactor(agent-core): share prompt-section prose and note v2 twins in agentfile headers
The Windows notes, additional-dirs and skills prose blocks existed twice:
inline in the builtin default template (system.md) and as constants in the
agent-file renderer (from-file.ts). Extract them to profile/prompt-sections.ts
as the single source: system.md renders them through injected KIMI_* template
variables and from-file.ts imports the same constants. Rendered prompts are
byte-identical for all four builtin profiles across macOS/Windows and
skills/dirs on/off; a new test pins system.md to the shared constants.
Also mark each profile/agentfile file with the path of its agent-core-v2
counterpart so format/semantics changes land in both engines.
* feat(cli): add /secondary_model command for the subagent model
Mirror /model: a picker with a thinking-effort step that persists [secondary_model] and live-applies to the current session via a new Session.setSecondaryModel RPC (node-sdk wrapper included), so newly spawned subagents bind the new model right away. The /model picker now hides the synthesized __secondary__ derived entry; docs and the update-config builtin skill mention the section.
* feat(tui): show the bound model in subagent run stats
Subagents report their model alias via agent.status.updated after spawn; resolve it to a display name and surface it in tool-call subagent stats and agent-group rows.
* fix(agent-core): validate agent profile before session persistence
* fix(agent-core): refresh subagent tools after model switch
* fix(agent-core): show subagent model preferences
* fix(agent-core): preserve secondary model recipe on live apply
* fix(agent-core): make secondary model apply explicit
* fix(tui): refresh secondary model display state
* chore: merge secondary model changesets into one
* Add /secondary_model command for subagent configuration
Show each subagent's model in the subagent card header and agent-group rows. Requires the secondary-model experiment (KIMI_CODE_EXPERIMENTAL_SECONDARY_MODEL=1); run /secondary_model to pick a model and thinking effort, applied to the current session immediately.
Signed-off-by: 7Sageer <sag77r@hotmail.com>
* fix(agent-core): align explicit agent file precedence
* fix(agent-core): let disallowedTools deny select_tools
* chore(cli): drop engine mention from --agent/--agent-file help text
* feat(cli): support --agent/--agent-file in the interactive TUI
Bind the selected agent profile to the startup session when launching
the TUI with --agent/--agent-file, including the session created after
an OAuth login at startup. Sessions created later in the process (/new)
keep the default profile.
Make both flags creation-only in every mode: combining them with
--session/--continue is now rejected in print mode too, since resume
restores the bound agent from the session automatically.
* fix(agent-core): persist new secondary-model selections under env overrides
stripSecondaryModelConfig restored secondary_model.model/default_effort
from raw whenever KIMI_SECONDARY_MODEL/KIMI_SECONDARY_EFFORT was set, so
a /secondary_model pick made under the env vars was silently discarded
on write. Restore from raw only when the value being written still
equals the env value (an overlay round-trip), mirroring the pointer
check in stripEnvModelConfig; a genuinely different selection now
reaches config.toml.
* fix(cli): report the effective secondary model when env overrides the pick
/secondary_model toasted the picked alias even when
KIMI_SECONDARY_MODEL/KIMI_SECONDARY_EFFORT made the session bind a
different model. Read the effective binding back from the reloaded
config (as /model does from session status) and warn with the
env-overridden values instead.
* feat(tui): show the bound model name in the AgentSwarm panel header
---------
Signed-off-by: 7Sageer <sag77r@hotmail.com>
|
||
|
|
de0ba9d065
|
fix(agent-core): count validation-rejected tool calls toward the repeat breaker (#2313)
* fix(agent-core): count validation-rejected tool calls toward the repeat breaker
Args-rejected calls returned before prepareToolExecution, so the breaker
never counted them and the model could re-issue the same invalid call
until maxSteps. Register them in finalizeToolResult so reminders fire at
3/5/8 and the turn force-stops at 12.
* fix(agent-core): key parse-failed repeats on raw argument text
Malformed JSON arguments normalize to {} on parse failure, which keyed
every malformed-but-different attempt identically and could force-stop a
turn whose calls were evolving rather than identical. Register skipped
calls on the raw arguments text when parsing failed.
---------
Co-authored-by: fengchenchen <fengchenchen@moonshot.ai>
|
||
|
|
cdbd33c13c
|
fix(kosong): fail fast on quota-exhausted 429 instead of retrying (#1857)
* fix(kosong): fail fast on quota-exhausted 429 instead of retrying A 429 caused by an exhausted account quota or insufficient balance (Moonshot error.type "exceeded_current_quota_error", OpenAI "insufficient_quota") can never succeed on retry, yet it was classified as APIProviderRateLimitError and silently retried for the whole budget (10 attempts, ~3 minutes of backoff) with no UI feedback — the session appeared frozen on every request. Introduce APIProviderQuotaExhaustedError, minted in normalizeAPIStatusError from the structured body error.type/error.code forwarded by convertOpenAIError, with billing-anchored message patterns as a fallback for gateways that flatten the body to text. The new class is excluded from isRetryableGenerateError (fail fast, even when a retry-after header is present) and from isProviderRateLimitError (no swarm requeue/suspend). toKimiErrorPayload and translateProviderError map it to provider.api_error (retryable: false) instead of provider.rate_limit, and classifyApiError reports it as quota_exhausted in telemetry. agent-core-v2 mirrors the same fix. Transient rate-limit 429s keep the existing retry, backoff, and Retry-After behavior (verified end-to-end against a mock provider: quota body fails after attempt 1/10; rate-limit body still walks the full 10-attempt ladder). Behavior changes to note: quota-failed swarm subagents now fail instead of suspending indefinitely as "Rate limited...", and quota errors cross the wire as provider.api_error rather than provider.rate_limit. * fix(kosong): classify quota exhaustion in OpenAI Responses stream errors Responses response.failed / error SSE events carry no HTTP status and were minted by errorFromOpenAIResponsesEvent as either a rate-limit error (rate_limit_exceeded / embedded status_code=429) or a base ChatProviderError — and the base class falls into the retryable unclassified-failure fallback, so an insufficient_quota event still burned the whole retry budget on the openai_responses path. Route the event code and message through the same quota-exhausted check before the rate-limit branch, in kosong and the agent-core-v2 mirror. Covers all three entry paths (error events, response.failed, nested gateway frames) since they share the single converter. * style(agent-core-v2): drop inline comments per AGENTS.md header-only rule agent-core-v2 comments live solely in the top-of-file block, never beside functions or statements; the kosong twins keep the full rationale. * refactor(kosong,agent-core-v2): move quota-429 checks to vendor hook Per review on #1857: the knowledge of how a backend signals quota exhaustion is vendor-specific and must not run for every OpenAI-compatible provider from the shared conversion layer. - Add a convertError hook: ProtocolTrait.convertError in agent-core-v2 (single-value, last-declarer-wins, bound by composeOpenAIChatHooks / composeAnthropicHooks / traitConvertError) and an equivalent optional hook parameter on convertOpenAIError / convertAnthropicError. Bases consult it with the raw failure (SDK error on HTTP paths, raw event on the Responses in-stream path) after the abort guard, before their own rules. - Declare Moonshot's quota signals (exceeded_current_quota_error, billing wordings) on the Kimi side: kimiOpenAITrait and kimiAnthropicTrait in v2, the KimiChatProvider and KimiFiles catch sites in kosong, all through the new classifyKimiQuotaError. - Drop the options parameter from normalizeAPIStatusError and the shared quota code/pattern tables: the contract layer keeps only the vendor-neutral APIProviderQuotaExhaustedError type and its retry / rate-limit / wire-mapping semantics. - The OpenAI bases keep recognizing only OpenAI's own documented insufficient_quota code (HTTP and Responses stream events) as protocol knowledge of that wire. Behavior: kimi and openai provider types classify exactly as before; an unregistered vendor speaking Moonshot billing wordings through a plain openai transport now stays a retryable rate limit by design. * fix(kosong,agent-core,agent-core-v2): wire kimi quota hook fully Follow-up to the second review round on #1857, all four findings: - Kimi-over-Anthropic (legacy engine): AnthropicOptions gains the same optional convertError hook as the OpenAI bases, threaded through AnthropicStreamedMessage and every catch site, and the provider manager's anthropic route now passes classifyKimiQuotaError for provider type kimi — a quota-exhausted 429 over this transport previously still burned the retry budget. classifyKimiQuotaError now also walks error -> .error -> .error.error for the code/type, since the Anthropic SDK keeps the full body on .error instead of hoisting. - v2 telemetry: ApiErrorKind gains 'quota_exhausted' and classifyApiError checks APIProviderQuotaExhaustedError before the generic 429 branch, matching the legacy engine's reporting. - Hook contract: converted ChatProviderErrors now pass through before the vendor hook is consulted in convertOpenAIError / convertAnthropicError (both engines), so the hook sees each raw failure exactly once even when a stream-minted error crosses an outer catch; tests assert the single consult. - protocolTrait: the convertError member doc shrinks to the concise style and the consult contract moves into the file header's composition rules. * test(kosong,agent-core,agent-core-v2): lock quota hook assembly paths Third review round on #1857: - Fix the v2 anthropic base header and AnthropicHooks doc still claiming withThinking is the only hook. - Drop the two remaining non-header JSDoc blocks in protocolTrait.ts per the AGENTS.md header-only rule; the consult contract already lives in the file header. - Update the ProtocolTrait contract test to the seventeen-hook shape (convertError included) and cover the traitConvertError binding. - Add real-assembly regression probes: the v2 registry composes a (kimi, anthropic) provider whose mocked SDK client throws a Moonshot quota 429 and generate rejects with the non-retryable APIProviderQuotaExhaustedError (a plain anthropic composition keeps the same 429 retryable); the legacy ProviderManager routing test asserts convertError is classifyKimiQuotaError on the kimi-anthropic route and absent for plain anthropic; the legacy provider threads options.convertError to its generate catch. * test(kosong,agent-core-v2): cover KimiFiles quota 429 and drop stale docs Fourth review round on #1857: - Drop the AnthropicHooks member JSDoc (its content already lives in the anthropic.ts and anthropicHooks.ts file headers) and fix the anthropic contrib header still calling the hook set single-hook. - Add the missing KimiFiles regression in both engines: a mocked files client rejecting with a Moonshot quota 429 makes uploadVideo reject with the non-retryable APIProviderQuotaExhaustedError, locking the classifyKimiQuotaError argument at the upload catch sites. |
||
|
|
0cef160c4b
|
fix(agent-core): continue goal pursuit when a goal turn hits the per-turn step limit (#2210)
Some checks are pending
CI / build (push) Waiting to run
CI / test (1) (push) Waiting to run
CI / test (2) (push) Waiting to run
CI / test (3) (push) Waiting to run
CI / test (4) (push) Waiting to run
CI / test (5) (push) Waiting to run
Release / Native release artifact (push) Blocked by required conditions
CI / test-pi-tui (push) Waiting to run
CI / test-windows (push) Waiting to run
CI / lint (push) Waiting to run
CI / typecheck (push) Waiting to run
Nix Build / Check flake.nix workspace sync (push) Waiting to run
Nix Build / nix build .#kimi-code (push) Blocked by required conditions
Release / Release (push) Waiting to run
Release / Deploy docs (push) Blocked by required conditions
Release / Publish native release assets (push) Blocked by required conditions
* fix(agent-core): continue goal pursuit when a goal turn hits the per-turn step limit * chore(agent-core-v2): regenerate state manifest * fix(agent-core): start goal pursuit when the goal-creating turn hits the step limit * refactor(agent-core-v2): drop step-cap narration from goal module and test |
||
|
|
c497af60e6
|
fix(tui): steer user messages into the running turn while a goal is active (#2153)
Some checks failed
CI / build (push) Has been cancelled
CI / test (1) (push) Has been cancelled
CI / test (2) (push) Has been cancelled
CI / test (3) (push) Has been cancelled
CI / test (4) (push) Has been cancelled
CI / test (5) (push) Has been cancelled
CI / test-pi-tui (push) Has been cancelled
CI / test-windows (push) Has been cancelled
CI / lint (push) Has been cancelled
CI / typecheck (push) Has been cancelled
Nix Build / Check flake.nix workspace sync (push) Has been cancelled
Release / Release (push) Has been cancelled
Nix Build / nix build .#kimi-code (push) Has been cancelled
Release / Deploy docs (push) Has been cancelled
Release / Native release artifact (push) Has been cancelled
Release / Publish native release assets (push) Has been cancelled
* fix(tui): steer user messages into the running turn while a goal is active * fix(tui): reset request state when steering mid-goal input fails * fix(agent-core): update session prompt metadata on steer |
||
|
|
f06eb5c60e
|
feat(agent-core): defer registered user tools (#2119)
Some checks are pending
CI / test-pi-tui (push) Waiting to run
CI / test-windows (push) Waiting to run
CI / lint (push) Waiting to run
CI / typecheck (push) Waiting to run
CI / test (1) (push) Waiting to run
CI / test (2) (push) Waiting to run
CI / test (3) (push) Waiting to run
CI / test (4) (push) Waiting to run
CI / test (5) (push) Waiting to run
CI / build (push) Waiting to run
Nix Build / Check flake.nix workspace sync (push) Waiting to run
Nix Build / nix build .#kimi-code (push) Blocked by required conditions
Release / Release (push) Waiting to run
Release / Deploy docs (push) Blocked by required conditions
Release / Native release artifact (push) Blocked by required conditions
Release / Publish native release assets (push) Blocked by required conditions
* feat(agent-core): defer registered user tools Allow hosts to mark registered user tools as deferred so select_tools can discover and load their schemas on demand in both agent-core implementations. * fix(agent-core): hide unregistered deferred tools Filter loaded deferred user-tool schemas from provider history and loaded-state checks after unregister while preserving canonical history and disconnected MCP behavior. * fix(agent-core-v2): drop stale inline schemas Stop treating previously loaded user-tool schemas as active after the same tool is re-registered for inline disclosure. |
||
|
|
66f611aae9
|
fix: echo thinking under the reasoning field the endpoint actually uses (#2104)
Some checks are pending
CI / lint (push) Waiting to run
CI / build (push) Waiting to run
CI / test (1) (push) Waiting to run
CI / test (2) (push) Waiting to run
CI / test (3) (push) Waiting to run
CI / test (4) (push) Waiting to run
CI / test (5) (push) Waiting to run
CI / test-pi-tui (push) Waiting to run
CI / test-windows (push) Waiting to run
CI / typecheck (push) Waiting to run
Nix Build / Check flake.nix workspace sync (push) Waiting to run
Nix Build / nix build .#kimi-code (push) Blocked by required conditions
Release / Release (push) Waiting to run
Release / Deploy docs (push) Blocked by required conditions
Release / Native release artifact (push) Blocked by required conditions
Release / Publish native release assets (push) Blocked by required conditions
OpenAI-compatible endpoints disagree on the wire field for reasoning content: `reasoning_content` (DeepSeek/Moonshot convention, older vLLM) vs `reasoning` (OpenAI GPT-OSS guidance, current vLLM — which also only accepts `reasoning` on the request side, vllm-project/vllm#38488). Reading only `reasoning_content` silently dropped thinking, and the lost thinking then never made it back into later requests. Add per-endpoint dialect detection: inbound responses are scanned in priority order (reasoning_content, reasoning_details, reasoning; first string value wins), the carrying key is remembered, and outbound messages echo thinking under the same key (default reasoning_content). An explicit `reasoningKey` / `reasoning_key` config still pins the dialect. - kosong: shared reasoning-key module; wire the kimi and openai-legacy providers; the dialect cell is shared across per-step provider clones. - agent-core-v2: same mechanism in the vendored openai-legacy base; the kimi trait no longer pins reasoning_content (the default already is), keeping Moonshot wire behavior byte-identical while adapting to vLLM. - agent-core: memoize the base provider behind ConfigState.provider so the detected dialect survives across turns instead of being rebuilt per access. |
||
|
|
5fdbdb4a22
|
feat: configure web search/fetch services via KIMI_WEB_* env vars (#2096)
* feat: configure web search/fetch services via KIMI_WEB_* env vars KIMI_WEB_SEARCH_BASE_URL / KIMI_WEB_SEARCH_API_KEY and KIMI_WEB_FETCH_BASE_URL / KIMI_WEB_FETCH_API_KEY overlay the [services] config section field by field in both engines (env wins over config.toml), so the WebSearch and FetchURL backends can be pointed at a Moonshot service without OAuth login. The v2 engine now also honors the explicit [services.moonshot_fetch] section (config > managed OAuth > local), which it previously parsed but never consumed. * fix(agent-core-v2): guard stripServicesEnv against clearing the services section config.replace(SERVICES_SECTION, undefined) — the logout deprovisioning path in authService — passes undefined straight into the section strip, and the composed stripServicesEnv dereferenced it (value[key]), throwing TypeError and aborting logout mid-cleanup. Add the same isPlainObject guard the other strip implementations (stripProvidersEnv, stripEnvBoundFields) already have, plus a regression test that also locks in the env overlay staying effective after the file value is cleared. * style(agent-core-v2): follow header-only comment convention * fix: isolate env web service credentials * Add env vars for web search and fetch services The new environment variables `KIMI_WEB_SEARCH_BASE_URL`, `KIMI_WEB_SEARCH_API_KEY`, `KIMI_WEB_FETCH_BASE_URL`, and `KIMI_WEB_FETCH_API_KEY` take priority over the corresponding fields in `config.toml`. The `kimi web` backend now honors the `[services.moonshot_fetch]` config section. Signed-off-by: 7Sageer <sag77r@hotmail.com> --------- Signed-off-by: 7Sageer <sag77r@hotmail.com> |
||
|
|
527d485d92
|
feat: add global default MCP server timeout configs (#2065)
Some checks are pending
CI / build (push) Waiting to run
CI / test (1) (push) Waiting to run
CI / test (2) (push) Waiting to run
CI / test (3) (push) Waiting to run
CI / test (4) (push) Waiting to run
CI / test (5) (push) Waiting to run
CI / test-pi-tui (push) Waiting to run
CI / test-windows (push) Waiting to run
CI / lint (push) Waiting to run
CI / typecheck (push) Waiting to run
Nix Build / Check flake.nix workspace sync (push) Waiting to run
Nix Build / nix build .#kimi-code (push) Blocked by required conditions
Release / Native release artifact (push) Blocked by required conditions
Release / Release (push) Waiting to run
Release / Deploy docs (push) Blocked by required conditions
Release / Publish native release assets (push) Blocked by required conditions
* feat: add global default MCP server startup timeout config Add a `[mcp] startup_timeout_ms` config.toml section with a `KIMI_MCP_STARTUP_TIMEOUT_MS` env override as the global default MCP server connection (startup + tool discovery) timeout. Precedence: per-server `startupTimeoutMs` in mcp.json > env var > config.toml > built-in 30s default. * feat: add global default MCP tool call timeout config Extend the `[mcp]` section with `tool_timeout_ms` and the `KIMI_MCP_TOOL_TIMEOUT_MS` env override as the global default for single MCP tool calls, mirroring the startup timeout: a per-server `toolTimeoutMs` in mcp.json still wins, and unset entries fall back to the SDK built-in 60s default. * chore: shorten the mcp timeouts changeset * feat(agent-core): add global default MCP server timeout configs Port the `[mcp]` section (`startup_timeout_ms` / `tool_timeout_ms`) and the `KIMI_MCP_STARTUP_TIMEOUT_MS` / `KIMI_MCP_TOOL_TIMEOUT_MS` env overrides to agent-core (v1), mirroring the v2 semantics: per-server fields in mcp.json > env vars > config.toml > built-in defaults. The v1 TOML loader gains explicit `mcp` read/write mappings, and both connection-manager construction sites (Session, testGlobalMcpServer RPC) pass the resolved defaults through. * fix(agent-core): validate and apply MCP timeout defaults * fix: propagate MCP startup timeout to SDK requests * refactor(agent-core-v2): resolve MCP default timeouts at connect time Keep the session connection manager synchronously lazy instead of gating its existence on config readiness: resolveDefaultTimeouts is read from the mcp config section at each (re)connect, so AgentMcpService's eager construction stays unconditionally safe and reconnects pick up changed preferences. The initial connect still awaits config.ready for a deterministic snapshot. Also restore the ISessionMcpService method docs, bump the changeset to minor, and fix the v1 env-parse comment. * chore(agent-core-v2): regenerate config manifest for the mcp section * Add global default MCP server timeouts configuration Specify the new global default MCP server timeouts in both the config file and environment variables. Signed-off-by: 7Sageer <sag77r@hotmail.com> --------- Signed-off-by: 7Sageer <sag77r@hotmail.com> |
||
|
|
8bf5bacba9
|
ci: release packages (#1989)
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> |
||
|
|
8250e590f3
|
docs(cron): drop references to the non-existent kimi resume command (#2050)
* docs(cron): drop references to the non-existent `kimi resume` command - point user docs at the real resume command `kimi --session` - reword cron tool descriptions and code comments in agent-core and agent-core-v2 to describe session resume without naming a subcommand * chore: add changeset for cron docs wording fix |
||
|
|
4c763f6763
|
feat: send prompt-attached videos directly with the prompt (#1999)
Some checks are pending
CI / lint (push) Waiting to run
CI / build (push) Waiting to run
CI / test (1) (push) Waiting to run
CI / test (2) (push) Waiting to run
CI / test (3) (push) Waiting to run
CI / test (4) (push) Waiting to run
CI / test (5) (push) Waiting to run
CI / test-pi-tui (push) Waiting to run
CI / typecheck (push) Waiting to run
Nix Build / Check flake.nix workspace sync (push) Waiting to run
Release / Native release artifact (push) Blocked by required conditions
Release / Release (push) Waiting to run
CI / test-windows (push) Waiting to run
Nix Build / nix build .#kimi-code (push) Blocked by required conditions
Release / Deploy docs (push) Blocked by required conditions
Release / Publish native release assets (push) Blocked by required conditions
* feat: send prompt-attached videos directly with the prompt
Videos attached to a prompt (pasted in the TUI, uploaded in the web UI)
previously reached the model only after it opened the file with
ReadMediaFile — an extra tool round trip that could leave the video
unseen when the model never made that call. They are now uploaded
through the model provider's file channel and embedded directly in the
user message as the provider-issued reference; ReadMediaFile stays as
the fallback and as the model's own way to open video files.
- agent-core: new uploadVideo agent RPC and Session.uploadVideo in the
SDK; the TUI uploads pasted videos at submit time, falling back to
the file-tag form on failure
- kap-server: inline file-source video prompt parts at the REST edge
and map provider file ids back to local uploads behind
GET /files/llm/{llm_id}
- kimi-web: play provider-referenced prompt videos after reload/resume
* fix: fall back to inline video when the upload channel fails
ReadMediaFile only used the provider's video upload channel when one
was bound, and surfaced a hard tool error when the upload itself
failed — on providers without a files endpoint (or a transient upload
failure) the model was told the video could not be read at all. Now a
missing or failing upload falls back to delivering the video inline
(base64), the same shape providers without an upload channel already
get. Both engines (agent-core and agent-core-v2) are fixed.
* test: cover prompt video edge cases and failure paths
Stress coverage for the prompt video pipeline:
- agent-core uploadVideo RPC: extension/magic classification
(.txt with video bytes accepted, extension trusted when magic is
absent), exact 100MB boundary, directory and nonexistent paths
- TUI: mixed per-video outcomes (one inlined, one tag fallback),
submission order behind an in-flight upload, queued video messages
carrying final uploaded parts
- kap-server: per-video provider-id mappings, Range requests through
the llm redirect, mapping persistence across a server restart
* fix: surface auth rejections from the video upload channel
The base64 fallback for a failing video upload must not mask auth
rejections: a 401/403 (surfaced as provider.auth_error) drives the
credential force-refresh and a clear auth error, while an inline
payload would just be rejected again by the next request. Only a
missing or broken upload channel (no files endpoint, network/server
errors) falls back to inline delivery.
* fix: keep the by-design no-hook video error from degrading to inline
Main's contract for a provider with no video upload hook is an honest
"does not support video upload" tool error — an inline payload would
be dropped on that protocol's wire anyway. The base64 fallback now only
applies to an upload channel that exists but failed at runtime. The
no-hook throw gets a stable type (VideoUploadUnsupportedError) so the
two cases are told apart without matching message text.
* fix: constrain the llm video route param to a safe alphabet
The provider file id is used as a blob-store key, and the node-fs
backend joins scope and key into a storage path. An id containing an
encoded path separator (%2F) could address a different storage path
than the intended llm-video mapping; the route now only accepts the
provider-id alphabet.
* style(agent-core-v2): move added explanations into top-of-file comment blocks
The package's comment convention keeps comments solely in the top-of-file
block; relocate the new notes (url-source id pairing, video delivery
fallback, VideoUploadUnsupportedError) from functions and schema fields
into their module headers.
* fix: reject prompt video uploads when the model lacks video input
The uploadVideo RPC only checked the provider's upload channel, so an
SDK caller on a text-only or unknown-capability model could obtain a
valid-looking video_url part the current model is not supposed to
accept. The TUI capability guard does not cover this public path, so
the agent now gates on video_in itself.
* fix: sniff uploaded video bytes before inlining them
The inline branch trusted the upload-time content type; now the bytes
are sniffed first, and anything magic-confirmed as a non-video kind
(e.g. an image mislabeled as video/mp4) falls back to the file-tag
form instead of being uploaded. Like the image gate, bytes are
authoritative where the container format allows it — an MPEG-PS
lookalike still rides the extension, matching ReadMediaFile.
* test: keep the generation stub pending until the abort lands
An immediately-answered 404 ends the turn (non-retryable) before the
test's abort call, racing the cleanup into a 409; the stub now hangs
generation until the abort cancels it.
* fix: validate prompt file references before mutating session controls
A stale file_id failed inside media resolution — after the model,
thinking and permission overrides had already been applied — so a
rejected prompt still changed the session's controls. File references
are now checked up front, keeping failed submits side-effect free.
* fix: serialize prompt submissions per session
A slow provider video upload let a later text-only request reach the
queue ahead of an earlier video one, silently reordering the
conversation for REST clients and multiple tabs. Submissions to the
same session now chain, matching the ordering the TUI already
guarantees locally.
* fix: play reloaded provider videos through the authenticated fetch path
A video recovered from an ms:// reference carried only the bare redirect
URL, which 401s under daemon auth when loaded natively. The attachment
now keeps the provider file id (llmFileId) end to end, and AuthMedia
fetches the bytes with the Bearer credential through the daemon's llm
redirect — the same blob-URL path uploaded files already use.
* fix: fence pending video submits against session and model switches
A slow paste-upload left the TUI idle, so /new, the session picker or
/model could fire mid-upload; the continuation then dispatched the old
session's provider reference into the newly selected session or a
model that cannot resolve it. The dispatch now re-checks that the
session and model are unchanged and asks for a resend instead.
* fix: preserve the caller's video path verbatim in uploadVideo
Trimming the path changed the filesystem target before validation and
upload, so a name that legitimately starts or ends with whitespace
resolved to the wrong file; the trim is now only the emptiness check.
* fix: forward provider-issued ids on image URL prompt parts too
The shared url-source schema accepts id for image and video parts, but
only the video path forwarded it, dropping provider-keyed image ids
between prompt acceptance and the model request.
* fix(web): reconcile inlined video echoes into the optimistic user message
The loose user-message matcher counted media parts and <video path>
tags but not the [video:ms://…] text shape, so a racing server echo of
an inlined upload slipped through as a duplicate user bubble.
* fix: queue bash submits behind a pending video upload
A bash-mode submit could start while a pasted video was still
uploading, recording shell context before the earlier prompt was
dispatched and reordering user actions; it now chains behind the
upload like normal submits.
* test: cast the driver through unknown for the session-switch fence test
* fix: validate prompt file references before resolving the prompt agent
A stale file_id posted to a fresh or cold session materialized the main
agent (registering it in session metadata and igniting agent-scoped
services) before the request was rejected. The file check now runs
first, so failed submits create nothing and mutate nothing.
* fix: key the playback mapping by the id embedded in the reference URL
Projections and clients read the provider id from the ms:// URL, not
the id field, so an upload that returns an id-less or mismatched
reference would inline fine but 404 on playback. The mapping now
derives its key from the URL, falling back to the explicit id.
* fix: sanitize provider video ids on the write side of the playback map
The read route got the safe-alphabet guard, but recordLlmVideoRef still
used the provider-returned id verbatim as a blob-store key; a crafted
provider response could write the mapping outside the llm-video
namespace. Out-of-alphabet ids are now dropped on both put and get.
* fix(web): keep recovered provider videos resendable through the edit path
The composer reload path only honored fileId and fell back to a
bare fetch(url) without the Bearer token, so editing a reloaded
provider-video turn dropped the chip after a 401. llmFileId is now
threaded through and the bytes are re-uploaded via the authenticated
llm redirect.
* fix: fall back to inline video for no-hook providers whose wire carries it
The by-design no-hook error is only the honest answer when the wire
would drop an inline payload anyway (the OpenAI family). Protocols
that convert video_url (kimi, anthropic, google-genai, vertex) now
take the base64 fallback instead of failing every video read; the
registrar computes the flag from the model's protocol.
* test: satisfy the full IBlobStore shape in the playback-map stub
The tsgo typecheck job rejects a structural stub missing _serviceBrand
and list even though plain tsc accepted it.
* fix: queue prompt-producing slash commands behind pending uploads
Skill activations and plugin commands started their turn immediately
while a pasted video was still uploading, so the earlier video prompt
queued behind them and user actions ran out of order. Both paths now
chain behind inputSubmitChain (and re-check session/model at dispatch)
like normal and bash submits.
* fix: validate media kinds in the prompt file-reference preflight
The preflight only proved a referenced file exists; a real upload used
with the wrong kind (e.g. a PDF submitted as video) still passed it and
mutated session controls before assertMediaFile rejected the request.
The kind assertion now runs up front with the existence check.
* fix(web): play sent and recovered videos in the file preview
The media preview returned early for every kind except image, so a
user-turn video chip's play action was a no-op even with llmFileId
threaded through. The preview now handles video: bytes come from the
authenticated file/llm fetch into a blob URL and render in a native
player.
* fix(web): preview recovered videos with the authenticated blob URL
The llm re-upload branch fetched bytes with auth but kept the protected
redirect URL as the chip preview, which 401s as a native video src; the
fetched blob now becomes the preview URL, mirroring the fileId branch.
* style(agent-core-v2): move the inline-fallback note into the module header
Same package comment convention as before: rationale lives in the
top-of-file block, not beside the registration call.
* fix: fence delayed bash submits to the originating session
The chained bash callback ran runShellCommandFromInput against whatever
session was active at dispatch time, so a command submitted in session
A could execute in session B's workspace and be recorded there after a
mid-upload switch. The originating session is now captured at submit
and re-verified at dispatch, like the prompt and skill paths.
* fix: emit prompt video telemetry from the agent scope
video_upload is an agent-level event requiring ambient agent identity,
but the uploader was built with the Core-scoped session view, leaving
prompt-upload events unattributable. The route now resolves telemetry
from the target agent for the uploader while image compression keeps
the session-scoped view.
* fix: preserve provider image ids through legacy projections and web mappers
The url-source id accepted by the prompt schema was dropped again by
the legacy message projection and the web wire mapper, so provider-
keyed image references lost their id across messages, snapshots, and
undo responses. It now flows through both directions.
* fix(web): revoke recovered video blob URLs before dropping attachments
A failed llm re-upload removed the attachment without revoking the
freshly created preview blob URL, pinning the whole video in the
browser blob store until page unload on every failed edit attempt.
* fix: serialize foreground slash commands behind pending uploads
/compact and /init started their turn while a pasted video was still
uploading, so the earlier message landed after them — and compaction
summarized the context without it. Prompt-producing builtin commands
now share the same queueBehindPendingUploads chain (with the
session/model dispatch fence) as skill, plugin, and bash submits.
* fix: defer session controls until media preparation succeeds
Media resolution now runs before any profile/model/thinking/permission/
denylist mutation, with the uploader resolved transiently from the
requested (or currently bound) model — a failed submission leaves the
session's controls untouched. A concurrent model switch during
preparation is rejected with session.busy instead of enqueueing a
reference uploaded for the previous model.
* chore: bump the new SDK video upload API as a minor release
Session.uploadVideo is new public API surface, not a patch-level tweak.
* fix: re-check busy state before draining upload-queued commands
A slash command deferred behind a video upload ran the moment the video
prompt dispatched, landing on an already-running turn: beginSessionRequest
wiped the active turn's live pane, and /init or /compact started on top
of it. The deferred callbacks for skills, plugin commands, /compact and
/init now re-run the resolver's busy check at dispatch and show the same
blocked message the user would get when typing while streaming.
* fix: resolve profile-bound models before choosing the uploader
A first prompt carrying "profile" without "model" resolved the upload
model from the still-unbound alias, so no uploader was installed and
every attached video fell back to a tool-read tag even when the
configured default model supports provider upload. The transient
resolution now mirrors AgentProfileService.bind: an explicit body model
wins, a profile bind falls back to the configured default model, and
only otherwise does the currently bound alias apply.
* refactor: resolve prompt videos at request time inside the engine
Move prompt-video delivery out of the submission edge: the TUI submits
synchronously with a local file:// part the v1 turn resolves before the
message enters history, and kap-server carries an internal kimi-file://
reference the v2 requester resolves against the effective model with an
app-scoped upload cache. History keeps the durable local file id, the
/messages projection emits structured video parts, and the web plays
videos back through the authenticated /files channel - deleting the
submit fences, per-session serialization, provider-id reverse mapping,
redirect endpoint, and the unreleased SDK upload API.
* fix: propagate abort through video upload delivery instead of degrading
A turn cancelled mid-upload used to be treated as an ordinary upload
failure: v1 fell back to an inline base64 part and appended the degraded
message to history, and the v2 resolver memoized the tag fallback for the
rest of the agent's lifetime. Both catch sites now check the delivery
signal itself - abort rejections vary in shape by provider - and re-throw
so cancellation ends the turn (v1, classified as cancelled via the abort
reason) or the request (v2, not memoized, so the next turn uploads).
* fix: keep the tag form for no-upload providers whose wire drops inline video
An OpenAI-family model configured with video_in but no provider upload
channel used to receive prompt videos as an inline base64 part - which
chat completions rejects and the Responses adapter degrades to an
omitted-video placeholder, persisting ~4/3x the file size in history for
bytes the model never sees. The prompt path now mirrors the v2 resolver's
protocol gate and degrades to the <video path> tag instead; ReadMediaFile's
own delivery is unchanged. Also merges the prompt-video changesets into a
single user-facing entry.
* fix: escape the NUL separator in the video upload cache key
The cache-key template literal contained a literal NUL byte instead of
the \0 escape, which made Git classify the whole source file as binary
- no inline diffs, unreliable text tooling. The escape produces the
byte-identical runtime string, so hashed cache keys are unchanged.
* fix: check the abort signal before the inline video fallback
The no-uploader inline path (and the post-upload-failure fall-through)
never consulted the delivery signal, so cancelling a turn while the video
bytes were being read still base64-encoded the file and appended the
degraded message to history. The inline branch now re-throws the abort
reason first, matching the upload catch.
* fix: retry transient prompt video upload failures on later steps
A generic upload failure used to memoize its tag fallback for the rest of
the agent's lifetime, freezing a transient files-endpoint error into a
permanently degraded video. The resolver now marks failure-born fallbacks
as non-memoizable: the current request keeps the lightweight tag form and
the next step retries the upload. Structural outcomes (successful uploads,
capability and sniff fallbacks, no-hook inline) stay memoized for
step-retry stability.
|
||
|
|
ba921ca531
|
fix: gate always-thinking inference to OpenAI wires, plus catalog review follow-ups (#2036)
* fix: gate always-thinking inference to OpenAI wires, plus review follow-ups
- catalog: strip the inferred alwaysThinking marker on non-OpenAI wires so
Claude/Gemini keep their native off; verified against live models.dev
- v1 provider-manager: honor per-alias baseUrl on kimi/google-genai/vertexai
- thinking: rewrite the PHASE-6 contract comment, drop three dead
cannot-disable warning branches, normalize requested effort in v1 to
match v2
- tests: input-cap compaction preference in both engines, kap-server WS
status cap, v2 legacy status input cap
- docs + changeset: new model fields, catalog-refresh behavior, kap-server
* fix: strip always-thinking only where the wire encodes a true off
The previous gate kept the marker only on the OpenAI wires, which wrongly
stripped it from Gemini 3 on the Google wires: its floor is
thinkingLevel MINIMAL with suppressed thoughts — still reasoning — so an
Off option there would be a lie. The criterion is now the wire's encoding,
not its family: strip only on anthropic and kimi, the two wires with a
protocol-level `thinking: {type: 'disabled'}` that the catalog's effort
list can never show. Verified against live models.dev data: google now
marks exactly the gemini-3 family (10), anthropic and moonshotai stay 0.
* docs(agent-core-v2): move the strict-validation contract into the thinking.ts file header
The scoped guide keeps comments in the top-of-file block only; the
function JSDoc shrinks to a short what-it-answers note matching its
neighbors. No behavior change.
|
||
|
|
ec88d352e8
|
fix: five correctness follow-ups to the catalog metadata work (#2030)
Some checks are pending
CI / build (push) Waiting to run
CI / test (1) (push) Waiting to run
CI / test (2) (push) Waiting to run
CI / test (3) (push) Waiting to run
CI / test (4) (push) Waiting to run
CI / test (5) (push) Waiting to run
CI / test-pi-tui (push) Waiting to run
CI / test-windows (push) Waiting to run
CI / lint (push) Waiting to run
CI / typecheck (push) Waiting to run
Nix Build / Check flake.nix workspace sync (push) Waiting to run
Nix Build / nix build .#kimi-code (push) Blocked by required conditions
Release / Native release artifact (push) Blocked by required conditions
Release / Release (push) Waiting to run
Release / Deploy docs (push) Blocked by required conditions
Release / Publish native release assets (push) Blocked by required conditions
* fix: five correctness follow-ups to the catalog metadata work - Normalize a configured effort value (case/whitespace) before thinking resolution on both engines, so "OFF" is read as off instead of being sent upstream as an invalid effort. - Clamp a declared input cap to the effective context window in the effective model resolution, so an override lowering max_context_size cannot leave a stale larger cap behind. - Report context usage against the same effective cap (max_input_tokens ?? max_context_tokens) in the SDK getStatus and the v2 legacy status projections, matching the session-event and service surfaces. - Preserve concretely declared per-model endpoints when the override npm is unrecognized, via the same OpenAI-compatible fallback used for top-level entries. - Attribute resolved.maxInputSize and resolved.capabilities .max_input_tokens in the model inspector to their config, override, or clamp provenance. * fix: honor observed context caps and publish the effective cap in v2 status - The overflow-learned provider window is now written into both max_context_tokens and max_input_tokens of the effective compaction context, so the strategy cannot bypass it by re-selecting the raw catalog input cap during overflow recovery. - agent.status.updated publishes max_input_tokens ?? max_context_tokens, matching the other v1/v2 status surfaces. * fix: never mutate config on clamp, cap usage ratio, and refuse proprietary override SDKs - effectiveModelAlias / effectiveModelConfig now build a copy before clamping maxInputSize to the effective window instead of rewriting the caller's config record in place. - context_usage is clamped to 1 in the v2 legacy status, the v1 session service, and the SDK getStatus; the default-model status fallback also resolves the input cap. - Overrides naming known proprietary SDKs (Bedrock, Cohere) are refused before the OpenAI-compatible fallback, matching top-level import behavior. - The inspector attributes a clamped maxInputSize to the clamp even when the raw value came from models.*.overrides. * fix: normalize forced efforts, prefer input cap in WS status, keep raw event ratio - Whitespace-only configured efforts now read as absent, and the KIMI_MODEL_THINKING_EFFORT override is lowercased on both engines. - kap-server's WS legacy status publishes max_input_tokens ?? max_context_tokens like the other status surfaces. - SDK getStatus keeps the ratio unclamped on purpose (>100% is the documented overflow signal on that path); schema-bounded REST status stays clamped to 1. - Pin the provider-observed window beating a declared input cap with a v1 full-compaction regression test. * fix: attribute resolved.maxInputSize provenance and drop test-body comments in v2 |
||
|
|
b5efba7abc
|
fix: consume the model metadata declared by the models.dev catalog (#2015)
* fix: stop advertising Claude thinking efforts for non-Claude models Models served over the Anthropic protocol whose names carry no Claude marker (e.g. a catalog-imported Kimi K3) no longer inherit the latest Opus effort list, so the model selector stops offering levels the model does not accept. The models.dev catalog import now also parses reasoning_options and records the declared effort levels on the model alias, so K3 offers its real levels (low / high / max). * fix: consume deprecated, override, and input-limit metadata from the models.dev catalog - Models declared status=deprecated in the catalog are no longer offered for import. - Per-model provider overrides on gateway providers (an npm package targeting an Anthropic SDK plus a usable endpoint) now land as alias protocol and base_url, so those models are served over the right protocol and endpoint; overrides without a usable URL are skipped. - A declared limit.input now sizes the context budget instead of the larger total context window (e.g. gpt-5: 272k instead of 400k). The model alias schema gains an optional base_url field (not accepted in overrides) that Anthropic wire resolution prefers over the provider-level base URL. * fix: honor thinking-disable semantics and the OpenAI-compatible fallback in catalog imports - reasoning_options 'none' is the model's off encoding: off_effort flows from the catalog through the model alias to the OpenAI wire providers, so turning thinking off sends 'none' instead of omitting the effort field; models with effort levels but no way to disable thinking are imported as always_thinking and no longer offer an Off option. - Bare Claude family aliases (e.g. sonnet-latest) recover the inferred Anthropic effort profile; v2 comment conventions restored. - Providers whose SDK the catalog does not type now fall back to the OpenAI-compatible wire (with a visible "guessed" note) instead of being refused; imports lacking a usable endpoint ask for one (--base-url on the CLI, a prompt in the TUI). Proprietary SDKs (Amazon Bedrock), unrecognized explicit types, and env-placeholder URLs are refused with a clear reason. * fix: align catalog imports with the reference models.dev consumer - A JSON null tier in declared effort values is now read as the 'none' off-encoding (previously such models were wrongly imported as always-thinking with no way to turn reasoning off). - Alpha-status models are filtered out alongside deprecated ones. - Models whose per-model provider override targets a wire that cannot be expressed per-model (e.g. Claude on google-vertex, whose wire here is Gemini-mode Vertex, or gpt entries on an Anthropic provider) are skipped instead of being imported under the silently wrong protocol. - interleaved: true no longer pins reasoning_content: the provider's default three-field scan is wider and the pinned key only narrowed reasoning parsing for gateways answering with another field name. * fix: require endpoints for Anthropic-compatible catalog imports and honor --base-url - catalogProviderNeedsBaseUrl now covers the Anthropic wire: a non-official Anthropic-compatible vendor without a concrete catalog endpoint (e.g. google-vertex-anthropic) must supply --base-url / the TUI prompt instead of silently falling back to the default Anthropic endpoint. - --base-url now takes precedence over the catalog-declared endpoint, and an empty --base-url is rejected instead of persisting a blank endpoint. * fix: enforce always-on thinking on every wire and refuse Cohere at import A model that declares always_thinking (e.g. a catalog-imported gpt-5) no longer resolves to a dishonest off state via thinking.enabled=false or an SDK/ACP off request: resolution clamps to the model's default effort on every wire instead of letting upstream keep reasoning while the UI reports Off. The Anthropic warn-and-send path for unlisted effort levels is unchanged. Cohere's proprietary SDK joins Amazon Bedrock on the import-refusal list instead of being guessed as OpenAI-compatible. * fix: harden catalog import edge cases - An explicit but unrecognized catalog type is now refused before npm/id inference, so a future catalog protocol is never silently miswired through the OpenAI fallback. - User-supplied --base-url values for Anthropic-wire providers get the same trailing-/v1 normalization as catalog endpoints, avoiding /v1/v1/messages requests. - The TUI import prompt rejects env-placeholder base URLs like the CLI does. * fix: await the floating assertion promise in the catalog add CLI test * refactor: unify catalog import resolution into a single decision function Wire-type inference, the OpenAI-compatible fallback, proprietary-SDK refusal, endpoint adaptation, and the base-URL requirement are now produced together by resolveCatalogImport, one pure resolver consumed by both the CLI and the TUI — replacing the cooperating predicates (inferWireType, isGuessedWireType, catalogProviderNeedsBaseUrl) whose permutations kept producing edge cases. No behavior change. * fix: close configured-off clamp hole, keep inferWireType compat, carry same-wire override endpoints - A configured thinking.effort = "off" no longer bypasses the always-on clamp: it is treated as absent and the model default applies, mirrored on both engines. - The previously public inferWireType stays as a deprecated compatibility wrapper over resolveCatalogImport so existing SDK consumers do not break on a patch release. - Catalog model overrides that stay on the provider's wire but declare their own endpoint now persist it on the alias (and the v1 OpenAI wire branches honor alias-level base URLs like the Anthropic branch). * fix: split total window from input cap and close override/endpoint gaps - max_context_tokens once again means the total context window (used by completion budgeting); a model's declared input limit is tracked as max_input_tokens, which compaction, context-splice and usage-ratio checks prefer — fixing the over-clamping introduced when the input cap was stored as the context budget. - A catalog endpoint declared only as an env placeholder now always produces needs-base-url (official SDK included), so credentials are never sent to the public vendor host by default. - api-only per-model overrides are honored as same-wire endpoint changes; overrides targeting another known but inexpressible wire (e.g. google-genai on an OpenAI gateway) are skipped; same-wire models whose declared endpoint is an unusable placeholder are skipped instead of silently rerouted. * style: drop a function-level comment from the v2 thinking resolver * chore: consolidate the PR's changesets into two user-facing entries |
||
|
|
37eda4e59a
|
feat(config): add env overrides for loop control and background task limits (#1993)
* feat(config): add env overrides for loop control and background task limits
Add three operational environment overrides, resolved as env > config.toml
> default in both engines (agent-core and agent-core-v2), matching the
existing KIMI_IMAGE_MAX_EDGE_PX / KIMI_SUBAGENT_TIMEOUT_MS pattern:
- KIMI_LOOP_MAX_STEPS_PER_TURN overrides loop_control.max_steps_per_turn
- KIMI_LOOP_MAX_RETRIES_PER_STEP overrides loop_control.max_retries_per_step
- KIMI_CODE_BACKGROUND_MAX_RUNNING_TASKS overrides background.max_running_tasks
Invalid values are ignored and fall back to the config value. The v2 engine
resolves them through config section env bindings (effective-only, never
persisted); the v1 engine resolves them at the consumption point.
* fix(config): strip env-bound fields before persisting config writes
Environment overrides resolved into the effective config could be echoed
back through IConfigService.set/replace (e.g. GET then POST
/api/v1/config) and persisted into config.toml, outliving the env var.
This affected the pre-existing image / subagent / keep-alive bindings as
well as the new loop control and max-running-tasks bindings.
Extend ConfigStripEnv with a getEnv parameter and add a shared
stripEnvBoundFields helper: while a field's env var is set, writes
restore the field's raw on-disk value (or drop it) instead of persisting
an echoed env value; when unset, normal writes persist. Register it for
the loopControl, task/background, image, and subagent sections.
Also fix ConfigService.stripEnv looking up rawSnake by the camelCase
domain key; on-disk sections are keyed snake_case.
* fix(config): honor binding parsers when stripping env fields on persist
An invalid env value (e.g. KIMI_LOOP_MAX_STEPS_PER_TURN=abc) is ignored
on the read path but still marked the field env-owned on the write path,
so a config write for that field was silently dropped.
stripEnvBoundFields now derives the guard from the section's envBindings
and skips fields whose env value fails the binding's parse, so invalid
env values are ignored on both paths — and the duplicated field/env
descriptor list is gone.
Also drop function-level comments added beside helpers; agent-core-v2
keeps comments solely in the top-of-file block, so the strip semantics
now live in the config.ts / configService.ts headers.
* fix(config): re-apply env overlays from the env-free base on every read
Two follow-ups from review:
- ConfigService.get()/getAll() re-applied section env bindings on the
already-overlaid effective cache (and get() mutated it in place), so a
valid override degraded to invalid or unset kept serving the stale
value until the next reload — and an echoed stale value could then be
persisted by a config write. Reads now recompute from a cached
env-free validated base, so degraded or removed env values fall back
to the file immediately.
- stripEnvBoundFields restored env-owned fields from the raw snake
sub-object, missing values persisted under legacy keys (e.g.
max_steps_per_run). Section stripEnv now receives the env-free,
fromToml-normalized raw base, so legacy aliases are honored.
* test(kap-server): retry temp-home cleanup in auth tests
auth.test.ts removed its temp home with a plain recursive rm, which races
late async writers in the server shutdown path and flakes with ENOTEMPTY
(seen on main CI and locally on main). Harden the cleanup with the same
maxRetries/retryDelay options sessions.test.ts already uses.
* chore(changeset): consolidate the env-override changesets into two entries
One feature entry (loop/background env overrides, both engines plus the
CLI) and one fix entry (env values persisting on writes and sticking
after degrade/unset, agent-core-v2 plus the CLI).
* fix(config): clear fully-stripped sections and refresh overlay domains on get()
Two follow-ups from review:
- stripEnvBoundFields returned an empty object when every written field
was env-owned, so a section with a non-empty default (e.g. subagent)
stored raw = {} and the default stopped applying until the next
reload. A fully stripped result now clears the raw section instead.
- get(domain) only recomputed env overlays for sections with env
bindings, so domains written solely by a ConfigEffectiveOverlay
(models / defaultModel from KIMI_MODEL_NAME) kept serving stale cached
values after the env changed. get() now derives every non-memory
domain from the fresh env-free base, matching getAll().
* test(config): drive the get() overlay freshness test with an inline overlay
The previous version imported the model env overlay to activate it, which
the CI runners failed to resolve from this test file (both tsgo and
vitest, while the same specifier resolves elsewhere — not reproducible
locally). An inline ConfigEffectiveOverlay double exercises the same
ConfigService contract with a tighter seam and no module dependency.
* fix(config): preserve unknown fields on full strip and clone nested env targets
Two follow-ups from review:
- stripEnvBoundFields cleared the raw section whenever the stripped
result was empty, so an env echo write could drop unknown
forward-compatible fields from the TOML table. An emptied section now
keeps its raw table while the env-free base still holds other fields,
and is cleared only when nothing remains (defaults keep applying).
- applyEnvBindings reused nested child objects in place, so a nested
env binding on a persistable key would mutate the env-free validated
base and serve stale values after the env var is removed. Children
are now cloned before descending.
* fix(config): keep the env-free base when a strip leaves nothing to persist
Returning {} for a fully stripped write stored the empty object into the
raw layer while the TOML table survived via rawSnake, so the bases
diverged; a second echo write then saw an empty base, cleared the
section, and deleted the table — dropping unknown forward-compatible
fields. When nothing persistable remains, the write is now a no-op for
the section (the env-free base is kept as-is), and the section is
cleared only when the base is empty.
* fix(config): revalidate stripped results so restored raw values never persist
A strip may restore on-disk values from the unvalidated raw base (e.g.
an env-masked invalid field), smuggling them past the merge-time
validation into the stored config: the section would then fail the next
buildValidated pass, dropping accompanying valid edits from the runtime
while the invalid value stayed persisted. set() now revalidates the
stripped result — discarding the parse output so unknown fields survive
— and rejects the write instead, matching replace().
|
||
|
|
e45832398d
|
fix(tui): keep long sessions responsive and bound resumed history (#1976)
* fix(pi-tui): reuse processed lines across frames in the renderer Each frame previously re-truncated, re-normalized, and re-compared every transcript line, so steady-state frames (spinner ticks, streaming flushes) cost O(total lines x chars) and pegged one core in long sessions. Keep the previous frame's raw lines, processed output, and per-line kitty image ids; a line whose raw string reference is unchanged reuses its processed output verbatim, and image-id consumers read the cache instead of re-scanning text. * fix(tui): stop tree-wide transcript invalidation on structural updates Grouping a second Read/Agent call, removing swarm progress, finalizing an MCP status row, replay tool-call removal, and /undo each invalidated the entire transcript tree, forcing every mounted message to re-render (markdown lexing + code highlighting) on the next frame. The container's render cache already validates per child by reference, so structural child-list changes are picked up without a tree-wide invalidate; reserve it for global style changes such as theme switches. * fix(tui): fold older assistant messages into the turn step summary Step merging only collapsed thinking/tool steps, so assistant text blocks accumulated without bound inside a turn (hundreds of markdown components in a single long turn, all re-processed every frame). A running turn now keeps its last 20 assistant messages (KIMI_CODE_TUI_KEEP_RECENT_ASSISTANT), and a finished turn folds down to its conclusion tail of 2 (KIMI_CODE_TUI_KEEP_RECENT_ASSISTANT_COMPLETED); older ones collapse into the step summary line with a message count. Entries are kept, so expand behavior is unchanged. * fix(session): bound resumed history to the most recent user turns Resume used to return every agent's full replay over the RPC boundary (a long session can reach ~100MB, serialized and parsed several times on an in-process call), while the TUI only renders the last 10 user turns. The resume payload now accepts an optional replayTurnLimit, the turn-boundary predicate moves into agent-core as the single source of truth (limitAgentReplayByTurns, re-exported through the SDK), and the CLI passes its existing 10-turn limit so resume transfers just the tail. * fix(session): count goal continuation rounds as replay turns Replay turn boundaries only matched real user input, so a 100-round goal (a handful of user prompts plus 100 system-trigger continuations) fell entirely inside the 10-turn replay window and resumed by rendering the whole run from the start. The goal driver already fires one synthetic prompt per goal turn and counts those as turns itself, so replay trimming now treats goal_continuation prompts as turn boundaries and keeps the most recent 10 rounds. The continuation prompt is model-facing and hidden live; replay no longer renders it as a user bubble either, while still advancing the replay turn so each round groups separately. * fix(tui): keep replay turn folding when goal continuations are hidden Suppressing goal continuation bubbles removed every turn-boundary component from goal-session replays, so mergeAllTurnSteps found no turn edges and skipped step/assistant folding entirely — a single oversized goal round still mounted all of its tool cards and assistant messages. Mount an invisible ReplayTurnBoundaryComponent for each hidden continuation: it renders zero lines but keeps the turn edges discoverable to folding and window trimming, matching replay trimming's per-round turn semantics. * chore: consolidate and simplify changesets |
||
|
|
ce0e3ceb04
|
feat: support custom agent files (#1735)
Some checks are pending
Nix Build / Check flake.nix workspace sync (push) Waiting to run
Nix Build / nix build .#kimi-code (push) Blocked by required conditions
Release / Release (push) Waiting to run
Release / Deploy docs (push) Blocked by required conditions
Release / Native release artifact (push) Blocked by required conditions
Release / Publish native release assets (push) Blocked by required conditions
CI / test (2) (push) Waiting to run
CI / test (3) (push) Waiting to run
CI / test (4) (push) Waiting to run
CI / test (5) (push) Waiting to run
CI / test-pi-tui (push) Waiting to run
CI / test-windows (push) Waiting to run
CI / lint (push) Waiting to run
CI / build (push) Waiting to run
CI / test (1) (push) Waiting to run
CI / typecheck (push) Waiting to run
* feat: support custom agent files
Discover Markdown+frontmatter agent definitions from user, project, and configured directories; merge them into a session-level profile catalog (priority: builtin < user < extra < project < explicit). Custom agents work as subagents via the Agent tool and as the main agent.
- agent-core-v2: new agentFileCatalog domain (discovery/parsing/profile factory, mirroring skillRoots/parseFrontmatter) and sessionAgentProfileCatalog merged view; AgentProfile.tools becomes optional (undefined = all tools) and gains disallowedTools deny list, evaluated in profileService.isToolActive and persisted in session wire records so resume keeps the gate even if the file is gone
- CLI: restore --agent/--agent-file for the v2 print runner (KIMI_CODE_EXPERIMENTAL_FLAG); the v1 TUI rejects them with a clear v2-only error
- kap-server/protocol: optional profile field on prompt submission with first-bind semantics (same-name no-op, different name rejected)
- docs: custom agents section (en/zh) + config/CLI references + changeset
* chore: shorten custom agents changeset
* fix(agent-core-v2): stringify caught errors in agent catalog log calls
* docs: drop v2 engine notes from custom agents docs
* fix: align custom agent binding semantics across engine and edges
- agent-core-v2: bind now owns the first-bind invariant — switching
profiles after bind throws profile.already_bound (checked again in the
synchronous segment before the first wire dispatch, so concurrent binds
cannot both pass); unknown names throw profile.unknown; same-name
rebinds keep the persisted thinking effort.
- kap-server / CLI: edges degrade to error mapping / same-name no-op
instead of their own divergent guards.
- agent files: reject non-string mode values, honor disallowedTools in
the append-mode Skill probe, pass --agent-file through unresolved so
the engine can expand ~, reject empty --agent-file values.
- session catalog: ready is recoverable via reload() after a fatal
source failure, and agent-file discovery is kicked at session
materialize so resumed sessions see file agents from the first turn.
- docs: first-bind semantics, name: agent override, tools: [] meaning,
--agent-file last-wins.
* fix: tighten custom agent behavior
* fix: address custom agent review findings
* fix(agent-core,agent-core-v2): keep active-tool wire records replayable by v1
Binding a profile without a tool allowlist (the default for file-defined
custom agents) persisted a tools.set_active_tools record with no `names`
key. v1 clients discover v2 sessions through the shared session index and
replay newer wire versions without migration, so the record crashed v1
resume with a TypeError and wedged the session permanently.
The v2 engine no longer writes the record when the base set is already
"every tool active" (its absence encodes the same state), and v1 replay
now skips records that lack `names` as defense in depth for wires already
written by preview builds. A same-name rebind that resets an allowlist to
"all tools" has no v1-safe encoding and is left as a documented gap (no
production caller today); a future tools.reset_active_tools Op is safe
because v1 silently no-ops unknown record types.
* fix(agent-core-v2): tolerate unreadable directories in agent-file discovery
A single unreadable subdirectory (EACCES anywhere in the scanned tree)
previously aborted the whole discovery pass, zeroing every agent of that
source on every session start with one path-less warning. The walker now
skips-and-warns per directory below the root (mirroring the skill
discovery it parallels), root-level failures are isolated per root so one
bad root no longer takes its siblings down, and only a genuinely transient
whole-fs outage (os.fs.unavailable) still propagates so the session
catalog keeps its previous contribution. Source-level warnings now name
the offending path, and repeated skip warnings are capped with a summary
that samples the suppressed paths.
Also consolidates the path primitives (~ expansion, base-relative
resolution, realpath type probes) shared by the root resolvers, the
walker, and the explicit-file source into agentFileCatalog/paths.ts, and
tightens parser diagnostics: frontmatter null is treated as absent, and a
present-but-wrong-typed name/description reports a type error instead of
"missing".
* fix(agent-core-v2): warn when a same-name builtin suppresses a file profile
A directory-discovered agent file colliding with a builtin profile
without override: true was silently dropped at merge time. The suppression
now logs a warning naming the profile and the opt-in.
* refactor(agent-core-v2): pass skillActive explicitly to renderSystemPrompt
The third parameter was a full tool list used only for includes('Skill'),
which forced the agent-file profile factory to answer a boolean question
with sentinel lists. The template now takes an explicit skillActive flag;
a skillActiveFor helper keeps builtin call sites derived from their tool
arrays.
* refactor(agent-core-v2): fall back to the configured default model in bind
BindAgentInput.model is now optional: the engine resolves a missing model
against the configured defaultModel and throws model.not_configured when
neither is set, so edges no longer each re-implement the fallback.
* fix(agent-core-v2,kap-server): reject unsupported thinking atomically at first bind
A REST prompt carrying profile + an unsupported thinking effort bound the
session first and failed setThinking after, wedging the session on an
identity the user never successfully used. The effort is now validated up
front when the caller marks it as an explicit request (strictThinking):
the bind rejects before any await or state mutation, and the requested
effort rides along in the bind instead of a separate setThinking. Internal
spawn/fork paths pass inherited thinking without the flag and keep the
previous clamp behavior — a persisted effort that drifted out of the
model's support list must not break subagent spawning. The route's
now-redundant model fallback is dropped in favor of the engine-side
default.
* fix(agent-core-v2): await the agent profile catalog at session materialize
The catalog's ready promise was only kicked, so a resumed session's first
turn could render the Agent tool description without the file-defined
agent types. Discovery is local-fs and cheap, so materialize now awaits
it; ready only rejects for a fatal explicit-source error, which is exactly
the case that should fail fast. A failure there now also removes and
disposes the half-materialized handle instead of leaving it registered in
the session cache.
* test: cover the --agent-file fatal path and tidy profile registration hygiene
The v2 print CLI now has a test asserting an invalid --agent-file fails
before any turn. The denylist profiles in binding.test.ts register in a
beforeAll (idempotent, scoped to the describe's run window) instead of at
module scope during collection.
* docs: align custom agent docs with v2-engine gating
--agent/--agent-file are rejected without the v2 engine, so restore the
requirement in the Agents and Command Reference pages (and use
KIMI_CODE_EXPERIMENTAL_FLAG=1 in the examples), note that tool lists only
shape model-visible disclosure (permission rules are the enforcement
layer), and remind authors of delegation-bound agents to state the
handoff contract in the prompt body.
* test(agent-core-v2): revert unrelated style churn in fs/workspace tests
Keep these files' diff limited to the realpath fakes the feature needs;
the lint-preference rewrites belong to a separate cleanup.
* feat(agent-core-v2): add permanent system prompt override via SYSTEM.md
Read $KIMI_CODE_HOME/SYSTEM.md on every startup and inject it as the default main-agent profile (name "agent", override: true), replacing the builtin default system prompt while inheriting builtin tools and description. Missing or empty files are ignored; unreadable files warn and fall back to the builtin profile.
The body supports variable substitution (${skills}, ${agents_md}, ${cwd}, ${cwd_listing}, ${os}, ${shell}, ${now}); unknown variables pass through verbatim.
Priority: --agent-file / --agent / project override > SYSTEM.md > same-name user-scope scan files.
* feat(agent-core-v2): gate tools globally and accept session disabledTools
Add a [tools] config section: "enabled" acts as a global allowlist (empty = unconstrained), "disabled" as a denylist applied on top, both intersected with the active profile's policy in isToolActive (mcp glob supported).
Plumb a session-persistent disabledTools parameter through the stack: v2 RPC PromptPayload, REST "disabled_tools" (protocol and kap-server parallel schemas), klient contract/facade, and node-sdk. The server applies it via profileService.setSessionDisabledTools, which replaces the client-owned denylist, keeps the profile's own deny, persists across resume, and rejects calls before a profile is bound with profile.not_bound (mapped to 40001). v1 core-api gains a type-only field and ignores it.
* fix: enforce session tool policy across agents
* fix(agent-core-v2): enforce tool policy at execution
* fix(agent-core-v2): align subagent tool descriptions with policy
* fix(agent-core-v2): harden custom agent policy state
* fix(agent-core-v2): harden custom agent lifecycle
* refactor(agent-core-v2): persist profile binding in a single profile.bind record
* fix(agent-core-v2): skip unreadable paths during agent file discovery
* fix(agent-core-v2): exempt select_tools from the executor policy guard
- share one composed profile/global/session tool-policy evaluation between
the executor gate and prompt rendering instead of two verbatim copies
- tolerate context-build failure in system prompt refresh instead of
rejecting callers (config watcher void-fire, session policy fan-out)
* test(agent-core-v2): resolve profile and tool-policy SUTs by interface
- drop the Object.assign patching of tool-policy methods onto the shared
profile service; rename describes so the SUT ownership is accurate
- classify profile.bind as v2-only with the accepted v1-replay tradeoff
documented, un-red the wire vocabulary guard test
- cover the select_tools guard exemption with an executor-level test
* fix(agent-core-v2): enforce explicit select_tools policy
* feat(agent-core-v2): accept Claude-style tool lists and rename agent-file mode to promptMode
* feat(agent-core-v2): unify prompt templating on ${var}
- Replace the nunjucks renderer with a single ${var} regex renderer
(unknown placeholders pass through verbatim) and drop nunjucks from
agent-core-v2.
- Merge the variable tables into one catalog shared by the builtin
system.md, SYSTEM.md, and agent file bodies; adds additional_dirs_info
plus code-composed blocks (windows_notes, additional_dirs_section,
skills_section).
- Replace the agent-file promptMode field with ${base_prompt}: bodies
are always rendered as templates, and ${base_prompt} expands to the
effective default profile prompt (honoring the SYSTEM.md override).
- Migrate the builtin system.md, goal reminders, compaction instruction,
and tool description templates to the same syntax.
* docs: complete agent priority chain and link SYSTEM.md precedence
* test: fix invalid custom agent fixture
* fix(cli): reject multiple agent selectors
* feat(agent-core-v2): add subagents allowlist to agent files
* fix(agent-core-v2): persist the subagent allowlist in the profile binding
The delegation allowlist now rides the profile.bind record like the tool
denylist, so a resumed session keeps enforcing it even when the source
agent file was deleted or changed. Agent/AgentSwarm resolve the caller's
allowlist from the persisted binding data instead of looking the profile
up in the live catalog.
* feat(agent-core-v2): warn on tool patterns that never match
Profile bind/apply and [tools] config changes now statically flag
entries that can never activate anything — wildcards without the mcp__
prefix (a bare * in an allowlist disables everything, in a denylist
nothing), incomplete mcp__ literals, and names no registered or
builtin-profile tool has — via a tool-pattern-no-match warning event,
once per pattern, instead of letting the tool set silently shrink. The
known-name vocabulary is the live registry plus literal names from the
builtin profiles, so flag-gated tools stay known and a typo in one agent
file cannot legitimize the same typo in another.
* docs: align --agent-file docs with the single-selector CLI
The flag accepts exactly one file and conflicts with --agent, but the
docs still described the earlier repeatable, composable design.
* docs: note the agent-file trust model and never-matching tool patterns
Spell out that project-scoped agent files can replace the default main
agent's whole system prompt (unlike AGENTS.md reference injection), and
list the three tool-pattern shapes that never match and now raise a
warning.
* chore: slim changeset wording to user-facing language
Drop wire record names, enforcement mechanics, and template syntax from
the entries; split the v1 resume fix into its own patch changeset; add a
patch entry for the tool-pattern warnings.
* chore: shorten changeset entries to one-line summaries
* Delete .changeset/v1-resume-v2-sessions.md
Signed-off-by: 7Sageer <12210216@mail.sustech.edu.cn>
* test: drop class-instance spread in sessionLifecycle test stub
* chore: clear the comments
---------
Signed-off-by: 7Sageer <12210216@mail.sustech.edu.cn>
|
||
|
|
e070a580f0
|
feat(telemetry): emit turn_id and agent_id on turn and tool events (#1675)
* fix: emit turn_id on turn_started/ended/interrupted telemetry
The turn lifecycle telemetry events (turn_started, turn_ended,
turn_interrupted) never carried the turn id, while tool_call and
tool_call_dedup_detected did. Any analysis correlating a turn's start,
end, or interruption back to its tool calls had nothing to join on.
Add turn_id (already in scope) to all three track() calls, matching the
existing key/value convention used by tool_call_dedup_detected.
Update the strict turn_started/turn_interrupted assertion to cover it.
* fix(agent-core-v2): emit turn_id on turn_started/ended/interrupted telemetry
Port the v1 fix to agent-core-v2: turn lifecycle telemetry events
(turn_started, turn_ended, turn_interrupted) carried no turn id while
tool_call did, leaving nothing to correlate a turn's start, end, or
interruption back to its tool calls.
Add turn_id to the three event interfaces, the telemetry registry
property docs, and the three track2() calls in AgentLoopService,
matching the existing ToolCallEvent key convention. Extend the
turn telemetry assertions in loop.test.ts to cover it.
* fix: emit turn_id on tool_call telemetry
* docs(agent-core-v2): clarify turn_id is a per-agent index in the telemetry registry
* feat(telemetry): add agent_id to turn and tool events
turn_id is a per-agent counter, so it collides across the main agent and subagents within a session. Emit each agent's scope id as agent_id on turn_*, tool_call, tool_call_dedup_detected, api_error and subagent_created (plus parent_agent_id) so events become attributable via (session_id, agent_id, turn_id).
* docs(changeset): cover turn_id emission alongside agent_id
* feat(telemetry): add agent_id to agent-level settings events
* feat(telemetry): link v2 events across agents, turns, and tool calls
- subagent_created: parent_tool_call_id, so a child run joins to the tool call that launched it
- permission_policy_decision / permission_approval_result: agent_id, turn_id, tool_call_id
- plan_submitted / plan_resolved / plan_enter_resolved, context_projection_repaired: agent_id
- compaction_finished / compaction_failed: agent_id + optional turn_id
- cron_scheduled / cron_deleted: optional agent_id of the scheduling agent
- api_error: turn_id + request_kind, so compaction request failures are distinguishable from turn requests
- tool_call_repeat: agent_id + optional turn_id; tool_call_dedup_detected stops fabricating turn_id: 0 outside a turn
- background_task_created/completed: task_id on both, unified kind vocabulary ('process' replaces the legacy 'bash' alias on created)
- agent lifecycle: auto-assigned agent-N ids now skip ids persisted by previous runs, so a resumed session cannot reissue agent-0 and collide with earlier telemetry
* fix(agent-core-v2): preserve telemetry turn attribution
* chore: split telemetry changeset into two logical changes
* refactor(agent-core-v2): bind agent telemetry context at scope
* test(agent-core-v2): fix telemetry assertions for scope-bound agent_id
- undoHistory/goal: include the injected agent_id in exact property assertions
- rpc-events: assert on the shared records array instead of a track2 spy,
which the scoped telemetry view bypasses
* refactor(agent-core-v2): bind agent identity ambiently in telemetry
- TelemetryService.withContext now returns a lightweight forwarding view:
transport state (appenders, enabled flag) stays with the App-scoped root,
so views created before an appender attaches no longer silently drop
events, and enablement changes apply to every live view.
- Agent-scope events register with defineAgentTelemetryEvent and compose the
centrally declared AgentTelemetryEventContext (agent_id) into their wire
schema; business payloads and call sites stay free of agent_id, enforced
at compile time and by the registry convention test.
- image_compress/image_crop stay plain events: the kap-server prompt routes
emit them through a session-scoped view without agent identity.
- v1 subagent_created gains parent_tool_call_id for parity with v2.
- Add a lifecycle test asserting each agent scope's telemetry view binds its
own agent id.
* Delete .changeset/subagent-id-reuse.md
Signed-off-by: 7Sageer <12210216@mail.sustech.edu.cn>
---------
Signed-off-by: 7Sageer <12210216@mail.sustech.edu.cn>
|
||
|
|
6dd4fd3368
|
refactor(agent-core-v2): rebuild the model wire layer on the kosong architecture (#1970)
* refactor(v2): land kosong contract layer (L0 wire contract) * refactor(v2): land kosong protocol layer (L1 traits and base registry) * refactor(v2): land kosong provider layer (bases, trait composers, kimi definition) * refactor(v2): land kosong model and catalog layers * refactor(v2): migrate engine callers to kosong, drop old llmProtocol layer * test(v2): migrate agent-core-v2 tests and harness to kosong * refactor: sync peripheral packages to kosong architecture * refactor(v2): replace provider dialects with per-protocol definitions * refactor(v2): merge the vertexai protocol into google-genai via providerOptions * refactor(v2): remove vendor-name gates outside the kosong layer * feat(v2): add model resolution inspection and connectivity ping * refactor(v2): remove the unused platform config layer * test(klient): pin invalid-input behavior across providers in e2e * refactor(agent-core-v2): merge kimi traits, split bases by protocol - merge the seven kimi trait modules into kimi.contrib.ts as two trait objects: kimiOpenAITrait (all native-transport hooks) and kimiAnthropicTrait (thinking only) - move bases/* implementations into per-protocol directories (openai/, anthropic/, google-genai/) and rename openai.contrib.ts to openai-legacy.contrib.ts - add per-directory index.ts registration barrels (import = registration) and exempt them in check-domain-layers.mjs alongside *.contrib.ts * refactor(agent-core-v2): reorganize model and request-layer types - consolidate shared model types (ModelOverrides, CompletionBudgetConfig/Params, ResolvedModelAuthMaterial, ThinkingDefaults/ModelThinkingMetadata) into kosong/model/model.types.ts and drop modelOverrides.ts - rename L2 request types to the ModelRequest* prefix: LLMEvent -> ModelRequestEvent, LLMRequestInput -> ModelRequestInput, LLMCallParams -> ModelRequestParams - extract ModelRequestTiming to replace three duplicated copies of the stream-timing shape - rename L3 llmRequester types with the Agent prefix to match IAgentLLMRequesterService (AgentLLMRequestOverrides/Finish/Task/Source/ PartHandler/LogFields) - delete the unused LLMRequestParams type * refactor(agent-core-v2): fold kosong/catalog into kosong/model - merge IModelCatalogService's enumeration surface (listModels / listProviders / getProvider / setDefaultModel and the wire shapes) into IModelCatalog; delete kosong/catalog/modelCatalog.ts - move the remote refresh path to the new IProviderDiscoveryService (discovery.ts + discoveryService.ts, renamed from catalog/modelCatalogService.ts) - relocate configSection / errors to discoveryConfigSection.ts / errors.ts; DEFAULT_MODEL_SECTION now lives in kosong/model/model.ts - drop the L3 catalog layer from check-domain-layers.mjs - kap-server routes / refresh scheduler / channelRegistry follow the split; klient renames the contract modelCatalogService to modelResolver and adds providerDiscovery; kimi-inspect reads via IModelCatalog - modelRequesterImpl: drop the streamedAnyPart backfill, onMessagePart already delivers every part - tests: remove the app/modelCatalog and kosong/catalog suites, add the kosong/model catalog and discovery suites * feat(klient): trim trailing undefined args and add boundary smoke probes - add trimTrailingUndefined helper so optional trailing args no longer cross the wire as null in http/ipc transports, which defeated server-side default parameters - add model-requester-boundary smoke probe for ChatProvider error wrapping behavior against real config and a local stub - add kimi-select-tools smoke probe verifying the kimi-only wire encoding of dynamic tool declarations - extend smoke.ts with a models set/get/delete round-trip, catalog list assertions, and update AGENTS.md with the new scripts * fix(agent-core-v2): declare openai chat hooks as function properties Method-shorthand members on OpenAIChatCompletionsHooks tripped typescript-eslint(unbound-method) at every extraction site (`const hook = this._hooks?.convertMessage` and friends), failing the repo-wide lint job. Every implementation is a plain closure composed by openaiHooks.ts, so declare the members as function-typed properties, which matches the actual semantics and clears the four errors. * fix: repair stale references surfaced by the origin/main rebase - agent-core-v2: point vacuousContent's ContentPart import at kosong/contract - kap-server: rewrite the transcript test seed as IModelCatalog (IModelResolver is gone) - klient: inline onceEvent/waitFor after the http transport helpers were dropped - kimi-inspect: remove useLiveEvent from ModelCatalogView; catalog polls on a slow interval * chore: downgrade the kosong architecture changeset to patch |
||
|
|
71bcfba54a
|
fix(agent-core): drop vacuous assistant messages that permanently wedge sessions (#1968)
* fix(agent-core): drop vacuous assistant messages that permanently wedge sessions A provider-filtered response can seal an assistant message holding only an empty thinking part into the recorded history. The projection's empty-message guard only counted parts, so it survived, and was serialized as an assistant message with no content and no tool calls — which the provider rejects with "the message at position N with role 'assistant' must not be empty" on every resend, permanently wedging the session. The projection now drops any message whose parts all serialize to nothing (empty/whitespace text, or empty unsigned thinking). Content-bearing messages keep every part verbatim, including empty thinking blocks that preserved-thinking providers require back. Drops surface through a new vacuous_message_dropped projection anomaly with log/telemetry counters, and the structural-error recovery now recognizes this provider rejection so the strict resend self-heals any residual path. Mirrored in agent-core-v2. * docs(agent-core-v2): move vacuous-message rationale into the file header The v2 comment convention keeps comments solely in the top-of-file block. Move the vacuous-message drop rationale there and drop the inline/JSDoc comments added beside statements in the projector service, the llmProtocol error patterns, and the projector tests. No behavior change. * fix(agent-core-v2): drop output-free steps at fold settle time The loop-event fold already dropped a settled assistant message when it was structurally empty (no content, no tool calls), but a step that recorded only an empty thinking part — e.g. a provider-filtered response carrying an empty reasoning field — survived into history, relying on the projection to keep it off the wire. Widen the settle condition to drop the step whenever nothing sendable was recorded (no tool calls; every content part vacuous), using a predicate now shared with the context projector. This heals stored v2 histories at restore time instead of only at request time. * fix(agent-core): drop duplicate-stripped assistants left with only vacuous content The strict resend's dedupe pass could re-create the exact empty-assistant shape the projection now guards against: when every tool call a message carried was removed as a later duplicate, a remainder holding only an empty thinking part was kept (its content was non-empty), serialized as an assistant message with no content and no tool calls, and rejected again with "the message at position N with role 'assistant' must not be empty" — the same failure the strict resend was recovering from. The keep condition now requires sendable content, so such a message is dropped wholesale with a vacuous_message_dropped anomaly alongside the duplicate-tool-call one. Mirrored in agent-core-v2. * fix(agent-core-v2): mirror output-free step drops in the transcript reducer The live loop-event fold drops a settled assistant step when nothing sendable was recorded (no tool calls; every content part vacuous), but the cold transcript reducer still pushed an assistant on every step.begin and never dropped it. Cold readers (snapshot / messages) therefore kept showing output-free phantom assistants — including one per failed retry attempt, which predates the vacuous-step case — and foldedLength overcounted the live folded history, which can misplace the unflushed-tail splice in the messages view while a turn is in flight. The reducer now settles steps the same way the fold does: at a step's end or at the next begin, an assistant with no tool calls and only vacuous content is removed and foldedLength is decremented. Steps carrying any sendable output — real text, real thinking, signed thinking, or tool calls — are kept verbatim. * fix(agent-core-v2): update vacuous-content predicate documentation for clarity * docs(changeset): simplify the vacuous-assistant fix changelog entry |
||
|
|
11c1683a1c
|
feat: scope thinking effort to the current session (#1933)
* feat: scope thinking effort to the current session Selecting a model or thinking mode in the TUI or web UI no longer persists the concrete effort to the config; only the boolean thinking toggle is saved. The web UI now restores and submits each session's own daemon-reported level instead of a browser-wide per-model pick, whose localStorage store is removed. * feat: persist bounded thinking efforts and migrate persisted max once Only low/medium/high/xhigh are written to the config on a pick; max and unrecognized levels stay session-only with just the boolean toggle saved. A one-shot migration rewrites a previously persisted thinking.effort max to high (recorded in migrations.json, never re-run). The web UI brings back the per-model localStorage pick as the seed for new sessions, while a session's own daemon-reported level keeps winning for existing ones. * feat: gate effort persistence on the model top declared level The session-only tier is now the last entry of the model's own support_efforts (ordered by strength) instead of a fixed name list, so custom provider-declared levels get the same treatment; anything below the top persists as before, and unknown model metadata keeps the concrete effort. * docs: align changeset wording with the top-tier persistence rule * refactor: rename the config migration marker file to migrations-effort.json * feat: drop the legacy global thinking pick fallback in the web UI A raw single-level string left in localStorage by older versions was migrated into a '*' fallback entry applied to every model; stale values (typically max) could silently steer any session. Non-map content is now discarded instead of inherited. * feat: drop the web per-model thinking pick store entirely Thinking in the web UI is now sourced only from the daemon: a session's own reported level wins when the model declares it, and everything else falls back to the model's catalog default. The stale localStorage key is cleaned up on startup. * feat: skip the effort write when the picker confirms its initial value Re-confirming the effort shown when the model/effort picker opened is not an explicit choice: the TUI persists only the model (no effort key, or no write at all when nothing changed), and the web skips the global config write the same way it already did for model switches. * revert: keep the web global thinking write on re-confirm for now Only the TUI skips the effort write when the picker confirms its initial value; the web behavior stays unchanged until the interaction is reconsidered. * fix: carry the draft thinking pick into a newly created session A level picked on the empty composer lived only in rawState.thinking, so selectSession's watcher overwrote it with the catalog default and the first prompt/skill submitted that instead of the pick. createDraftSession now captures the draft level and seeds the new session's own entry. * docs: simplify the changeset wording * fix: close the remaining top-tier persistence and draft-capture gaps The provider-add default-model flow now resolves support_efforts through effectiveModelForHost (overrides + protocol-profile inference), so a top-tier pick on catalog models without declared efforts no longer persists. The draft thinking level is captured before the session creation awaits, so a concurrent session switch can no longer seed the new session with another session's effort. * fix: wait for the session status fold before resolving prompt thinking In the cold window right after a reload or session switch, the session's own level has not landed from /status yet; resolving straight to the catalog default would carry the wrong level to the daemon, which writes it into the session profile. Send, steer, side-chat and skill-activation paths now await the fold when the session entry is missing. |