The stall-detection repeat key joined the tool name and canonical args
with a literal NUL (0x00) separator. The control byte caused git to
classify stall-hook.ts as binary, so diffs, blame, and code review on
the file were opaque — which prevented confirming the test history for
this feature. Replace the NUL with a normal space (tool names are
identifiers and never contain spaces, so keys stay collision-free) so
the file is plain UTF-8 text and remains reviewable.
Behavior is unchanged: the key still uniquely combines tool name and
canonical args. Verified by reverting the hook to a no-op stub to show
the three stall-detection test files go red (the discriminating block,
canonical-key, e2e turn-abort, and worker-stall-translation cases all
fail), then restoring the real implementation to confirm they pass —
the failing-first the prior atomic commit never recorded.
Full suite: 5049 passed / 25 skipped; make typecheck clean.
Introduce a central, env-driven flag registry in agent-core. Each flag is declared once with an id, full env var name, default, and surface. Within agent-core, flags are consulted through a process-global 'flags' constant that reads live process.env. Resolution precedence: master switch KIMI_CODE_EXPERIMENTAL_FLAG > per-feature KIMI_CODE_EXPERIMENTAL_<NAME> > registry default, with lenient boolean parsing via parseBooleanEnv. FlagId is a literal union derived from the registry for compile-time autocomplete and typo-checking.
SDK boundary: KimiHarness.getExperimentalFlags() returns the resolved values over RPC, and the SDK re-exports only the flag *types* — no runtime value crosses the boundary. The TUI caches that snapshot once at startup and reads it synchronously for command gating.
Gate the plugin system behind the 'plugins' flag, off by default: PluginManager.load() consults flags.enabled('plugins'), so when off no installed plugins are loaded or activated, and the TUI /plugins command is hidden from the palette and resolves as an unknown command.
Tests cover the resolver precedence matrix, registry invariants, the FlagId type guard, the live-env singleton, the plugin-load gate, the getExperimentalFlags RPC, and the TUI command gating.
* fix(update): don't report success when native update fails
The native auto-updater spawned `bash -c "curl -fsSL … | bash"`. A
pipeline's exit status is that of its last command, so when curl could
not connect (e.g. a dead proxy) it produced no output, the trailing bash
read empty stdin and exited 0, and the whole command looked successful —
printing "Updated … Restart the CLI" while nothing had been installed.
Run the spawned shell with `set -o pipefail` so curl's non-zero status
propagates. installUpdate() then rejects and runUpdatePreflight() warns
and continues on the current version instead of claiming success.
* chore: add changeset for native update fix
* fix(tui): show real terminal status for background agents
The Agent tool's run_in_background=true call returns a non-error
ToolResult whose body just says "status: running". The transcript
card derived its done/failed badge from that result, so every
terminated background agent — including ones reconcile reclassifies
as lost on resume — kept the green "✓ Completed" label even when
the actual task failed, was killed, or never came back.
Push the real BackgroundTaskInfo.status into the matching Agent
card so the badge reflects what happened. The card's resolver
prefers subagent agentId (live) and falls back to the description
on resume; on resume the apply step also runs after replay
finishes so the agent group can reach the borrowed components.
Also adds an agent-core regression test that pins live, busy,
group, race, and resume scenarios for the bg notification chain.
* fix(tui): also propagate bg agent terminal status to standalone cards
Standalone Agent cards (only one Agent tool call in a step, never
upgraded into an AgentGroupComponent) bypassed the previous
`setBackgroundTaskTerminalStatus` path: the standalone header reads
`getDerivedSubagentPhase`, which still derived `done` from the
non-error spawn-success ToolResult, and the method did not request
a header/content rebuild. Lost/failed/killed bg agents in this
shape still rendered as `✓ Completed`.
Thread the override through `getDerivedSubagentPhase`, populate
`subagentError` with the friendly failure message so both render
paths share one source of truth, and trigger the same header +
content rebuild that `onSubagentFailed` does. Also include the
override in `hasSubagentState` / the subagent-block early-return
so a replayed solo bg agent (no replayed subagent block, no
sub-tool activity) switches to the subagent-aware layout instead
of the generic `Used Agent` rendering.
Adds two standalone-render regression tests so the path no longer
relies on the grouped snapshot to stay correct.
* feat(agent-core): make resume actionable from the lost-task notification
A backgrounded subagent that ends as `lost`/`failed`/`killed` is
already a soft-recoverable thing — `subagentHost.resume` will
reanimate the persisted Agent instance — but the LLM had to dig
through the original spawn-success ToolResult to find the right id
and figure out the recovery shape on its own. The two look-alike
identifiers (the BackgroundManager `task_id` aka `source_id`, and
the `subagentHost` `agent_id`) regularly got confused in practice.
Surface what the model needs at the moment of decision:
- Add `agent_id` as a top-level `<notification>` attribute for
agent-* tasks, so the right id is structural, not buried in
prose. Render path keeps backward-compat by omitting the
attribute when no agent_id is known (bash tasks, old sessions).
- On non-success agent terminal states, append a recovery
paragraph to the body: the precise `Agent(resume=...)` call,
the disambiguation between `agent_id` and `source_id`, the
`run_in_background` option, and what state survives the
restart vs. what may need to be redone.
- Tighten the spawn-time `resume_hint` with the same
disambiguation and an explicit pointer at the
`task.lost`/`task.failed`/`task.killed` recovery trigger.
- Persist `agent_id` and `subagent_type` in PersistedTask so the
recovery body still works after a session restart, where
in-memory `BackgroundTaskInfo.agentId` would otherwise be
undefined. Optional fields keep the disk schema
forward/backward compatible — pre-PR records load without
them and silently fall back to the original short body.
* fix(tui): route bg-agent terminal events by stable agent_id, not description
`tc.subagentAgentId` is left undefined for every backgrounded agent.
`handleSubagentSpawned` early-returns for `runInBackground` before
calling `tc.onSubagentSpawned`, and the wire replay path drops the
`subagent` block entirely (`toolCallFromReplayMessage` returns only
id/name/args). So the `agentId` branch in
`applyBackgroundTaskTerminalStatus` never matched in practice, every
call fell through to the description-based fallback, and the
persisted `agent_id` we added in the previous commit was effectively
dead. That fallback also has a real failure mode: if a foreground
Agent and a backgrounded Agent share the same `args.description`,
the only candidate found is the live (unrelated) card, which gets
incorrectly relabeled as the lost task's terminal state.
Parse `agent_id: agent-N` out of the AgentTool spawn-success
ToolResult body inside `getSubagentAgentId` so the id is always
recoverable, regardless of whether the in-memory subagent metadata
was ever populated. Foreground and backgrounded Agent cards now
carry distinct ids and route correctly.
Also pipe the real `subagent.failed` error through to the parent
card. The background branch of `handleSubagentFailed` previously
only appended the dedicated transcript entry; the parent Agent
card was left with the generic "Background agent failed" written
by the later `background.task.terminated` event. Add an optional
`errorText` to `setBackgroundTaskTerminalStatus` /
`applyBackgroundTaskTerminalStatus` and pass `event.error` through
on the failed branch — the real reason now reaches both the card
and the entry.
* fix(tui): treat agent_id as authoritative when matching bg terminal events
Previously `applyBackgroundTaskTerminalStatus` always tried agent_id
first and then fell back to description match on miss. That fallback
caused two cross-card bugs:
1. On resume, `applyTerminalBackgroundAgentStatuses` iterates every
persisted terminal task, including ones whose tool calls fell
outside the `REPLAY_TURN_LIMIT` window and were never mounted.
Description fallback could route an old `lost` status onto an
unrelated recent Agent card sharing the same `args.description`.
2. During the live spawn → terminate window, the same card briefly
lives in both `_pendingToolComponents` and `transcriptContainer`.
A description-only walk visits the same component twice and flags
itself ambiguous, dropping the otherwise unambiguous update.
When `args.agentId` is provided we now match only by id and skip on
miss. With `getSubagentAgentId` already parsing `agent_id: agent-N`
out of the spawn-success ToolResult, the id path is reliable for
both live and resume even though `tc.subagentAgentId` is never
populated for backgrounded agents. Description fallback is preserved
solely for old pre-PR sessions whose persisted records lack
`agent_id` — same best-effort behavior as before.
A mid-stream SSE drop surfaces as a raw undici `TypeError: terminated`, which was classified as a non-retryable generic error and failed the turn on the first attempt. Route raw transport-layer errors through the connection-error heuristic so a dropped stream becomes a retryable APIConnectionError and is retried transparently. User aborts (ESC) are unaffected — the retry loop checks the abort signal before retrying.
Related to #149.
The `/status` and `/usage` "Plan usage" rows previously rendered the
progress bar from the used ratio while labelling it "X% left", so the
bar direction and the number disagreed.
Display "X% used" instead, aligning the number with the bar, and move
the reset hint to the right without parentheses to mirror the web
console layout.
Merge turnId/step into a single `turnStep` field ("0.1") and
attempt/maxAttempts into `attempt` ("2/3"), and drop the
messageCount/toolCallCount fields. The per-request `llm request`
line goes from up to 8 fields down to ~3; the `llm config` line
(including thinkingEffort, logged for all providers) is unchanged.
Nix's profile prepends its bin directory to PATH, shadowing the Node.js version installed by setup-node. Moving Nix installation before setup-node ensures the Node >=24 PATH entry remains first, fixing `ERR_PNPM_UNSUPPORTED_ENGINE` during release.
- Add version:release script combining changeset version with Nix hash update
- Update release workflow to install Nix and use version:release
- Document updated hash refresh workflow in for-agents/workflows.md
- Include changeset for the fix
* docs: simplify plugins documentation
* docs: restore plugin caveats lost during simplification
Re-add three behaviors that were dropped from the simplified plugins
docs: the stdio MCP `cwd` must start with `./`, local-path installs run
from the managed copy (so editing the source after install requires a
reinstall), and `/plugins remove` only deletes the install record while
leaving files on disk. Mirror the changes in both en and zh.