* fix(agent-core-v2): guard config persistence against lossy writes
- A failed load no longer clears the in-memory snapshot: the service keeps
the last-known-good config, reports an error diagnostic, and taints.
set/replace/replaceSections on the persisted layer then fail fast with
Error2(config.persist_blocked) instead of erasing the file; memory-layer
overrides stay available, and a successful reload clears the taint.
- persistDomains is now read-modify-write: the file is re-read and only the
domains being written are applied on top of current disk content, so
external edits are merged instead of clobbered, and an external delete is
honored instead of resurrected.
- External changes absorbed at persist time trigger a full reload so change
events fire for domains the writer did not touch.
* fix(agent-core-v2): rebase set() merges onto re-read config state
set(domain, patch) now merges the patch against the freshly re-read file
content and refreshes the in-memory snapshot from the same read, so external
edits to the same section survive a concurrent write instead of being
overwritten by the stale in-memory copy.
* fix(protocol): register config.persist_blocked in KimiErrorCode
Add the new code to the KimiErrorCode union and kimiErrorCodeSchema so the
persist-refusal error payload passes protocol validation across RPC
boundaries.
* fix(agent-core-v2): compute every config write against the re-read file
Move strip/merge/validate for set/replace/replaceSections into the persist
rebase callback so each write is derived from the file content re-read at
persist time. Overlay strip handlers (e.g. the KIMI_MODEL_* mask restoring
default_model) now read the fresh snapshot instead of the stale in-memory
one, and the unconditional snapshot sync makes the separate
absorbed-external reload redundant.
* fix(agent-core-v2): build defaults when the initial config load fails
A failed first load has no last-known-good state worth preserving, so fall
through with an empty document: registered section defaults are still
validated and applied (consumers of defaulted sections keep working), while
the taint keeps blocking persisted writes until a reload succeeds. Only
reload failures preserve the previous in-memory state.
* fix(agent-core-v2): stage re-read config snapshots until the write succeeds
Build the rebased raw/rawSnake snapshots in locals and publish them only
after the rebase and documentStore.set both succeed, so a validation error
or a storage failure cannot leave userValue and effective pointing at
different snapshots. stripEnv now takes the staged snapshots explicitly.
Split the dense prose walls into a minimal config, a one-line-per-field
table, constraint bullets, a numbered resolution order, and a separate
advanced subsection for per-entry thinking efforts. All behavioral facts
are preserved; zh and en stay mirrored.
The startup banner now comes from the backend client_configs endpoint
(config name client_banner) instead of the CDN-hosted tips.json, and is
fetched fresh on every startup with no caching. The payload keeps the
tips.json shape, plus two targeting additions:
- banner_platform (top level and fallback entries) limits display to
the given platform; missing, empty, or all means every platform, and
the CLI only shows entries targeting all or cli.
- banner_start_time/banner_end_time on fallback entries add scheduled
visibility windows with the same semantics as the active banner.
Finished `!` output collapses to the first 10 visual rows with a `... (N more lines, ctrl+o to expand)` marker, sharing the global ctrl+o toggle with agent tool output; ctrl+o also expands the live buffer while the command runs. Replayed output mounts the same card and behaves identically. The shared truncation component, the running card's default view, and the agent bash path are unchanged.
* feat(datasource): add NDA/NBS, standards, IGO, xhcj, and caixin sources
* fix(datasource): narrow real-time-news ban to coverage gaps, require PublishTime citation
* fix(datasource): scope the real-time-news limitation to coverage gaps only
* fix(datasource): trim redundant clause in the real-time-news limitation
* fix(datasource): stop on a result that covers the question, not the first success
* fix(datasource): front-load trigger terms in the skill listing description
* fix(datasource): exempt discovery calls from the one-call workflow
GET /api/v2/sessions gains view=by_workspace: one request returns every
workspace with a matching session, each carrying its first group.page_size
sessions under the requested sort plus the workspace's full matching total,
with group-level page_token pagination (40922 on condition drift). Groups
key on the alias-canonical workspace id, so legacy split buckets of one
physical directory merge into a single group, matching the v1 alias
semantics. meta.has_prompt filters sessions by prompt presence (the v1
exclude_empty equivalent) in both views. The flat view and v1 routes stay
byte-compatible.
The global WS stream now fans out event.session.archived (live and cold
paths; payload carries the session id and workspace_id) and
event.workspace.created/updated/deleted, published by the core
IWorkspaceService on every mutation path including the implicit
createOrTouch on session creation.
kimi-inspect consumes the grouped projection as a single-column
workspace/session tree in the chat view; the session pane merges into the
right dock as the Session tab. The server API reference (en + zh) documents
the new parameters, the grouped response, and the new events.
* feat(agent-core-v2): rework the title generation excerpts
- Rebalance the excerpt budgets toward the user's prompts (400 chars
each) and trim the assistant segments (300) so titles follow the
user's task instead of narrating the assistant's reply.
- Cap each prompt in the default user_prompts excerpt so one long
paste no longer starves the remaining prompts.
- Compose the digest excerpt from the full conversation arc: every
natural-language user prompt in the live window paired with its own
turn's final assistant text, interleaved chronologically, with
per-segment caps and a 3000-char total budget (middle turns elided).
* chore: scope the title changeset to agent-core-v2
* fix(agent-core-v2): dedupe digest prompts and elide whole turns
- Drop the redundant `| undefined` from the optional
TitleDigestTurn.assistant per the monorepo optional-property
convention.
- Deduplicate user messages by id when constructing digest turns, so a
prompt already in the context and still active in the queue does not
produce two turns.
- Elide the over-budget digest at whole-turn granularity, keeping each
assistant line paired with its own user line.
* docs(agent-core-v2): describe the full-arc digest in the SessionTitleSource contract
* fix: fail fast on provider-filtered empty responses
An APIEmptyResponseError carrying finishReason 'filtered' (OpenAI
content_filter, Anthropic refusal) is deterministic: replaying the same
request re-triggers the provider safety filter. Both isRetryableGenerateError
implementations (kosong, agent-core-v2) treated every empty response as
retryable, so step retry replayed the doomed request the full 10 attempts
before the filter notice surfaced. Return non-retryable for filtered empty
responses in both engines; the error already carries the provider.filtered
code, so the turn fails immediately with the existing filter notice.
* fix: skip the compaction shrink-retry for filtered empty responses
Both full-compaction loops routed every APIEmptyResponseError into the
shrink-and-continue branch before isRetryableGenerateError was consulted,
so a filtered response was retried with shrinking input instead of failing
fast. Exclude finishReason 'filtered' from the shrink branch in both
engines; it now falls through to the retryability check and throws
immediately. Add end-to-end tests (real kosong generate over a filtered
think-only stream) asserting a single attempt with the history untouched.
---------
Co-authored-by: kimi-agent-bot <kimi-agent-bot@users.noreply.github.com>
* feat(kimi-code): add China/International region selection for OAuth login
- Add region profiles (cn/overseas) and resolver in @moonshot-ai/kimi-code-oauth:
env override → persisted login host → install-channel marker → default cn
- /login now offers Kimi Code (China) / Kimi Code (International); the CLI
login entries (kimi login, kimi acp --login) accept --region cn|overseas
- Update/plugin/site/telemetry endpoints derive from the selected region;
plugin trust list covers both .com and .ai hosts
- kap-server: POST /oauth/login accepts an optional region; new GET /oauth/region
* fix(oauth): keep an explicit default-slot login ahead of the install marker
A China login persists no oauthHost (the default credential slot carries
no host trace), so after switching back from International the resolver
fell through to a stale overseas install marker. Treat a persisted
default-slot oauth ref (key === oauth/kimi-code) as an explicit-cn signal
that outranks the marker; getRegion() on the v2 side mirrors it.
* fix(agent-core-v2): thread the default-slot key through capability region resolution
Capability installs resolved the region from the persisted oauthHost only,
so an explicit China login (which persists no host) lost to a stale
overseas install marker. Pass the oauth ref key through as well, matching
getRegion(). Also move the region contract notes into the auth.ts file
header per the package comment convention.
* fix(agent-core-v2): honor the region-marker opt-out for the telemetry endpoint
Hosts that set KIMI_CODE_REGION_MARKER=off (the desktop embedded server)
skip the install marker in getRegion(), but the default telemetry endpoint
still consulted it, so a stale overseas marker could split the reported
region from the telemetry destination.
* feat(cli): show region site domains in login platform selector
* chore: reword oauth login changesets
* fix: honor the region marker opt-out in the CLI and capability resolvers
* refactor: rename login region values to mainland-cn and global
* fix: keep the --region help text in English
* fix: simplify the --region help text to site domains
* feat: drop the suggested login platform order
* feat: split a browser-safe region profile table out of the region resolver
* Revert "feat: split a browser-safe region profile table out of the region resolver"
This reverts commit a037b1143e.
* fix: read the install marker from the bootstrapped home directory
* fix: resolve the server plugin marketplace from the active login region
* feat: expose the login region option through the klient auth facade
* fix: drop a comment from the v2 auth region test
* fix: keep scoped base-only logins on their environment for a bare login
* fix: invalidate the region cache on the provider-manager logout path
* fix: route client-config fetches through the active region profile
* fix: resolve the telemetry endpoint per flush so a login region switch applies in-process
* test: expect the telemetry endpoint resolver in the CLI init assertions
* fix: resolve the default telemetry endpoint from the bootstrapped home
* chore: reword the oauth login changeset around the two login methods
* chore: trim the oauth login changeset to the headline
* feat: let hosts override the region marker env through the server bootstrap env bag
- give the builtin agent profile an explicit subagents allowlist (coder, explore, plan), restoring v1 semantics
- inherit the default profile's allowlist when a caller profile declares none, instead of leaving delegation unrestricted
- pass a lone "*" subagents field through as an explicit unrestricted marker
* docs: document KIMI_CODE_CUSTOM_HEADERS on the env vars page
* docs: address review on KIMI_CODE_CUSTOM_HEADERS entry
- use a neutral gateway header name in the example
- correct the release version to 0.20.2
- scope the override claim to exact-name matches and warn against
case-variant auth headers
* docs: describe protocol-dependent Authorization precedence
On the OpenAI-compatible protocols (kimi/openai/openai_responses) an
exact Authorization custom header is applied after the SDK-generated
bearer token and therefore replaces it; /models listing keeps its own
authentication.
* docs(zh): add the required space before the config-files link
Per the mixed-content spacing rule in docs/AGENTS.md.
---------
Co-authored-by: bj456736 <bj456736@users.noreply.github.com>
* fix(kap-server): serve real session usage in snapshot and persist per-turn context readings
* fix(kap-server): omit unknown session usage fields instead of reporting zero
* refactor(agent-core-v2): persist cron tasks as durable wire records
- write CronAdd/CronDelete/CronCursor as durable wire records and rebuild the cron task table from dispatcher replay
- migrate legacy per-workspace cron JSON files into the wire on first resume, then drop the file-based persistence service, its registrations, and the bootstrap cron scope
- derive the session cron view from the agent replayable cron state and remove the redundant session-level copy
- let session forks inherit cron tasks through the copied wire instead of duplicating task files
* fix(agent-core-v2): keep legacy cron tasks on cold forks and flush before cleanup
- inherit legacy cron task files into a full fork's wire so cold sessions forked before their first post-upgrade resume do not silently lose scheduled tasks
- flush the migrated wire records before deleting legacy files so a crash cannot lose both copies
* refactor(agent-core-v2): drop the legacy cron file migration
- stop reading legacy per-workspace cron JSON files entirely; pre-upgrade tasks simply stop applying instead of being migrated into the wire
- remove the legacy read path, the fork-time legacy inheritance, and the now-unused session context/document store injections
* refactor(kap-server): table-driven dispatch for multi-action routes
- add a shared action-dispatch helper; sessions, prompts, plugins, questions, and modelCatalog collection routes declare action tables with module-level handlers
- add ISessionManager.status returning the session summary; the archive action checks it instead of resuming the session
- archive cold sessions through the persisted-metadata path shared with batch archive
* fix(agent-core-v2): read-your-writes for session index point gets
Overlay pending mirror summaries in getFromReadModel so a freshly recorded
summary (e.g. cold-session archive) is visible to GET immediately, matching
the existing pending overlays in list and cursor resolution.
* feat(agent-core-v2): carry prompt attachments on turn.started so the live transcript projects them
* fix(transcript): clear the transcript goal when the goal is cleared
* fix(agent-core-v2): count a prompt media part as a transcript attachment only when its id matches its daemon file URL
* test(kap-server): expect the session-media file id on converted prompt parts
Move the external hook services out of app/externalHooksRunner,
session/externalHooks, and agent/externalHooks into
features/externalHooks, assembled as the ExternalHooksFeature unit:
- services live under per-scope subdirectories (app/, session/, agent/);
shared pure helpers (types, hook matching/dispatch, process spawn,
prompt result rendering) live under internal/
- the runner and the two observers are contributed through the Feature
seams (ScopeUnits materialization); the hooks config section stays on
the static import=register channel
- update the package entry leaf exports, the plugin domain imports, the
kap-server events-zod import, and the affected tests; regenerate the
state manifest
* fix(vscode): multi-select question jumps to next after only one answer selected
* chore: add changeset
---------
Co-authored-by: gaoyuan <gaoyuan@moonshot.ai>
* fix(agent-core-v2): emit subagent.spawned after task registration
The spawned signal previously fired at launch, before the run's task
registration, so clients learned the agent id with no task id to bind
cancel/status actions to; a failed registration also left a spawned row
behind for a run that never registered. Emit it only after registerTask
succeeds and carry the task id on the event.
* fix(agent-core-v2): keep spawned ahead of started for Agent-tool runs
The TUI drops subagent.started until spawned has established the row,
and a failed registration must not leave a started row behind with no
terminal event. Defer the mirrored started dispatch so the Agent tool
can emit it itself after registration and spawned.
* fix(agent-core-v2): void the deferred started dispatch
* fix(kap-server): key Agent-tool transcript rows by the registered task id
Transcript-protocol clients suppress the raw task.*/subagent.* session
events, so they only saw a subagent row keyed by agent id that cannot
address /tasks/{id}, plus a second row once task.started landed. Key the
spawned row by the task id it now carries, fold task.started and the
subagent lifecycle back into it, and keep the agent-id path for spawns
without a registration (swarm/session-init/tower). Statement-level
ordering notes move to the file headers per package convention.
* test(agent-core-v2): split the spawned/started ordering contract into its own test
* fix(kap-server): keep subagent result details across task termination and drop stale task mappings on taskless respawns
* fix(kap-server): recover the agent-to-task association from a backfilled task.started
* fix(kap-server): seed pre-attach Agent task mappings on the transcript binding
* fix(kap-server): seed the full in-flight task row on transcript bind, not only its id
* docs(agent-core-v2): name the state-domain event dispatcher in the Agent tool header
* style(kap-server): drop comments in transcript services per the no-comments lint rule
* feat(kimi-code): specialize the WaitFor tool's transcript display
* feat(agent-core-v2): emit status progress while WaitFor is pending
* fix(kimi-code): route WaitFor dimming through the TUI theme
* feat(kimi-code): support replaceable status updates in tool progress
* fix(kimi-code): forward status progress to subagent activity surfaces
* fix(agent-core-v2): drop the redundant undefined from ToolUpdate.replace
* fix(kimi-code): honor replace semantics in the subagent live status path
* test(agent-core-v2): drive the WaitFor progress test through a manual tick
* fix(kap-server): mirror ToolUpdate.replace in the ws event schema
* refactor(agent-core-v2): expose the WaitFor progress scheduler as a public seam
* fix(kimi-code): pass child wait statuses without the trailing newline
* feat(agent-core-v2): tick the WaitFor progress status every second
* feat(agent-core-v2): format WaitFor progress durations as 1m 15s
* feat(agent-core-v2): omit zero seconds and minutes in WaitFor durations
* refactor(agent-core-v2): unify the loop-event fold into one core with two materializations
The loop-event stream was reduced by two hand-mirrored state machines:
loopEventFold.ts for the live/replayed context and contextTranscript.ts
for the full transcript behind the messages endpoints, kept in sync by
comments alone and already drifted (transcript dropped tool-result note
metadata and never closed a dangling tool exchange at step.end).
createLoopEventFold now owns the shared state machine once (settle,
pending tool exchanges, deferred appends, vacuous tracking) and both
views plug in as LoopEventFoldSink materializations. New parity tests
pin the foldedLength === live length invariant the endpoints splice on.
* fix(agent-core-v2): drop every removed prompt's injections on multi-turn transcript undo
The transcript undo only walked prompt-owned injections off the oldest
counted anchor, so with count > 1 an injection owned by a newer removed
prompt (e.g. an image-compression caption) survived the display undo
while the live context removed it. Collect every counted anchor's id
during the walk and sweep their owned injections afterwards, keeping
the transcript's 'prompt-owned ones leave with their prompt' contract
for every count and matching the live view.
* refactor(agent-core-v2): drop module headers from the context fold modules
The comment-free zone lint only allows JSDoc on exported symbols.
* fix(agent-core-v2): recover fold state after rehydration
* fix(agent-core-v2): scope undo injections to their prompt
* fix(agent-core-v2): settle open transcript frames when compaction lands mid-fold
An overflow-triggered compaction arrives with the failed attempt's
frame still open. The transcript appended the summary marker and reset
the fold but left the frame, so a vacuous partial stayed in the entries
while the live context dropped it, and a pending tool exchange lost its
interrupted result. Settle through the shared fold core at the marker
instead: close pending tool calls, drop or seal the open frame, then
append the summary. recoverFoldedLength recomputes the absolute count
right after either way.
* fix(agent-core-v2): keep legacy compaction recovery on the pre-settlement count
A legacy context.apply_compaction record (compactedCount without
keptUserMessageCount) recovers foldedLength as 1 + (foldedLength -
compactedCount), and the live legacy tail shape keeps the unsettled
open frame inside history.slice(compactedCount). Settling the fold for
those records shifted foldedLength by the settlement delta before the
recovery read it, leaving the transcript count one off the live
context. Gate the settle to modern records; legacy records keep the
previous freeze-and-reset behavior.
* fix(agent-core-v2): stop advertising unavailable ReadMediaFile to non-multimodal models
* fix(agent-core-v2): honor the effective tool policy before advertising ReadMediaFile
* fix(agent-core-v2): keep media-unavailable guidance reason-neutral and within the active toolset
* fix(agent-core-v2): recommend MCP fallbacks only when an active MCP tool exists
* fix(agent-core-v2): stop naming other tools in Read descriptions and errors
* fix(agent-core-v2): drop tool roster from plan agent prompt
Assistant messages carrying only tool calls were serialized without a
content key (JSON.stringify drops undefined), which strict
chat-completions validators such as LiteLLM reject with a 422,
permanently poisoning the session. Emit content: null for such messages
in both the kosong and agent-core-v2 converters, matching the shape
OpenAI responses use alongside tool_calls. The think-only empty-string
behavior in agent-core-v2 and both Kimi providers' deliberate content
omission are unchanged.
* feat(agent-core-v2): add the WaitFor tool for waiting on background tasks
* fix(agent-core-v2): mark WaitFor deliveries only after formatting succeeds
* fix(agent-core-v2): cancel losing waits once the WaitFor race resolves
* test(node-sdk): project WaitFor out of the v1-v2 resume parity roster
* fix(agent-core-v2): gate WaitFor goal guidance behind the wait_for flag
* fix(agent-core-v2): gate WaitFor goal guidance on actual tool availability
* fix(agent-core-v2): enforce the wait_for flag at WaitFor execution time
* fix(agent-core-v2): consult the live tool policy in the WaitFor availability check
#2593 replaced AgentVideoResolverService with the image+video
AgentMediaResolverService and reduced videoResolverService.ts to a pure
deprecated alias with no DI registration. #2909's squash merge restored
the pre-#2593 file wholesale, bringing back the legacy class and its
registerScopedService call. Both classes then registered the same token
('agentVideoResolverService') at the Agent scope and the legacy
video-only resolver won on the production import order, so image
kimi-file:// references reached the provider unresolved. Gateways
reject the unknown scheme with a 400 ("unsupported image url"), the
media-strip fallback then hid the image from the model, and pasted
images only worked on undo-resend via the inline base64 fallback.
Delete the legacy alias files and their index exports, drop the stale
alias assertion, and pin the behavior with a klient e2e regression:
a kimi-file image prompt part must reach the provider as a data: URL,
never verbatim.
* fix(kimi-code): upload pasted videos to the daemon file store
Video paste staged a cache copy and submitted a bare file:// video_url,
which the v2 engine no longer resolves, so the submission failed and the
persisted history retried it on every turn. Mirror the image flow
instead: upload the paste to the daemon file store in the background and
submit a kimi-file:// reference that the engine's prompt intake
materializes. A video whose upload is still in flight, failed, or
expired now refuses the submission with an actionable error, since video
bytes have no inline fallback form.
* test(kimi-code): fix MessageDriver recallStashedMedia signature
* chore: sync web dist from code-app
code-app: 1fb57f0ee3c424675d8768dcc70e56f5c953cd9f
* chore: sync web dist from code-app
code-app: 93da508d079b0118cc0338da97dcb738f0a3d46f
Adds the web session admin page. Built from the code-app PR #261 branch tip before its merge; the squash-merged main tree is expected to be identical (will be checked at merge time and rebuilt here if not).
* chore: correct the code-app trailer of the previous sync commit
The previous commit's trailer had a mistyped code-app SHA. The dist content is byte-identical to a build of code-app main at the merge below (tree verified identical to the branch tip it was built from), so the watermark for the next sync is:
code-app: 025805b33f1e87a4bc479574971c6177ad1403a0
* feat(kimi-code): support automatic updates for native installations via staged swap
Native (SEA) installs previously could not self-update on Windows and
relied on 'curl | bash' re-install on Unix. Replace both with a staged
swap updater:
- startup swaps in a staged binary (verified against the release
manifest sha256, smoke-checked via --version) and re-execs it, so the
running process never replaces itself (Windows-safe)
- downloads run in a self-spawned hidden sub-command, in the background
from the update preflight or in the foreground from 'kimi upgrade'
- rollback from .bak on any swap failure; install failures keep the
existing retry/prompt thresholds
* fix(kimi-code): fully clean staged artifacts on swap discard paths
Real-binary smoke testing on macOS surfaced two cleanup gaps in the
discard path: the claimed metadata file was unlinked after the staging
dir rmdir (so the empty dir survived), and the staged exe was
rediscovered via the already-claimed staged.json (so it leaked on the
downgrade-guard path). Pass the known metadata through and order the
unlink before the rmdir.
* fix(kimi-code): restore staged metadata on swap failure and sweep update leftovers at startup
* fix(kimi-code): address codex review on lock contention and swap crash window
- The background native install no longer takes the outer install lock:
the self-spawned downloader holds it for the whole download, and the
parent's spawn-time lock raced the child into a false lastSuccess.
- Smoke-check the staged exe before moving anything, so a bad staged
binary is discarded with the install path never left empty; the
remaining crash window is two adjacent atomic renames (documented,
recoverable via the .bak or by re-running the install script).
* test(kimi-code): align swap test expectation with smoke-before-rename order
The restore-on-failure case now observes the early smoke check's
--version spawn; only the re-exec spawn must be absent.
* fix(kimi-code): stage the bare CDN binary instead of unzipping
The published per-release artifacts are the bare platform binaries
(kimi-code-<target>[.exe]), not zip archives — the staging flow now
streams the download straight to the staged exe after the manifest
sha256 check, and the zip reader is dropped. Verified end-to-end on
macOS against the live CDN: download -> sha256 match -> swap ->
re-exec into the real released binary.
* fix(kimi-code): address second codex review round
- re-exec: forward 128 + signo when the swapped-in child dies by signal
instead of reporting exit 0
- __update_download: only exit 0 without staging when the lock holder is
staging the SAME version; a different in-flight version (or a vanished
lock) no longer surfaces as a successful foreground upgrade
- staging: sweep orphaned .part downloads and unreferenced staged exes
before downloading, preserving live swap claims and their payloads
* feat(kimi-code): show download progress for native updates
The foreground 'kimi upgrade' path streamed 180 MB with a single static
'Downloading…' line. Render progress instead: a throttled in-place
percentage line on a TTY, one line per 32 MB when piped, and plain MB
counts when Content-Length is unknown.
* fix(kimi-code): bound native update downloads with an idle timeout
Codex review: the manifest fetch cleared its timer once headers arrived,
so a stalled response body hung the worker forever, and the binary
download had no abort at all. The manifest timeout now covers body
consumption, and the binary stream aborts after 30 s without a chunk
(total duration stays unbounded for slow networks). The idle timeout is
injectable for tests.
* fix(kimi-code): retry native updates blocked by an orphaned active record
Windows real-machine verification surfaced that a parent exiting before
the downloader's exit event leaves a fresh-looking 'active' record that
silently blocks every background retry for the 6 h TTL. For native
installs, lock liveness is the truth past a 60 s spawn grace window:
a held lock means a download is running, a free lock means the record
is an orphan and a new attempt may start. Package-manager sources keep
the TTL behavior (no lock to prove liveness).
* fix: skip staged swap while another instance holds a fresh claim
sweepStaleNativeUpdateArtifacts already detected an in-progress swap in
a concurrent instance, but the result stayed inside the cleanup helper:
startup still claimed a newly published staged.json and ran a second
swap, so the two launchers could rename the install path and delete each
other's rollback backup. Propagate the in-progress signal and skip
claiming until the existing claim is released or goes stale.
* fix: keep the install lock while its holder process is alive
The install lock went stale purely by age (30 min), but the native
downloader is idle-bounded, not duration-bounded: a slow link can
legitimately take longer. Another startup would then sweep the lock and
spawn a second downloader, and both would write and clean the same
.staging paths. Past the age threshold, fall back to a pid liveness
probe (signal 0) — the lock is stale only when the holder is gone.
* fix: keep recovery artifacts on rollback failure and wait out same-version downloads
Two robustness fixes from review:
- native-swap: when moving the staged exe into place fails AND the
rollback rename fails too (transient lock, AV), the install path is
left absent and no next launch can start. Discarding the staged
payload and claim on top of that removes the second recovery copy.
rollback() now reports its result; on a double failure the swap keeps
the .bak (which IS the old exe), the staged exe and the claim so
manual recovery or a re-install still works.
- update-download: a foreground `kimi upgrade` racing a background
downloader of the same version exited 0 immediately, so the CLI
printed a success message for a download that could still fail. The
worker now waits while the same-version holder is in flight, adopts
the verified staged result (staged.json lands before the lock is
released), and takes over the download when the holder finished
without staging.
* fix: stamp the swap claim with a fresh mtime when claiming
rename() preserves the staged metadata's mtime, which can be arbitrarily
old — the background download often finishes hours before the next
launch claims it. A concurrent launch's sweep would then classify the
live claim as crash residue (older than the 5-minute window) and delete
the claim, the staged exe, and eventually the first swap's rollback
backup. Stamp the claim file with the claim time so the staleness check
measures the swap's liveness, not the download's age.
* fix: stamp the claim before the rename so it is born fresh
Stamping after the rename left a window: a concurrent launch could
inspect the claim between the two syscalls, see the staged metadata's
old mtime, and delete the staged executable mid-swap. utimes the state
file first so the claim carries a fresh timestamp from the instant it is
published — no fresh-looking-later intermediate state exists.
* fix: chmod the staged download before publishing it at its final name
A swap claims only the staged METADATA; the staged exe stays in
.staging/. A concurrent same-version downloader (possible because swaps
do not hold the install lock) then re-downloads and renames its .part
over that path. If the swap moves the file into the install path between
the downloader's rename and its post-publish chmod, the chmod lands on a
path that is already gone and the installation is left non-executable —
every future launch fails. Apply the executable mode to the private
.part file before the publishing rename so the staged exe is executable
from the instant it appears.
* fix: publish the install lock atomically via hard link
The 'wx' open exposed a momentarily empty lock file before its contents
were written. A concurrent acquirer reading in that window got a
SyntaxError, treated the lock as stale, swept it and also won — two
"holders" then ran stageNativeUpdate against the same .staging paths.
Write the lock contents to a unique temp file and hard-link it into
place: link() fails when the destination exists (same exclusivity as
'wx') and the lock path only ever appears fully written.
* fix: serialize stale-lock takeover through a secondary lock
A pathname-level delete can never be conditioned on the file still being
the inspected stale instance, so a plain compare-and-delete still loses
exclusivity: two workers classifying the same stale lock could interleave
unlink and publish such that both won (proven by a 20-way contention
test). Takeovers now go through a secondary create-if-absent lock
(install.lock.takeover): the delete+publish section only ever runs in
one process, staleness is re-validated inside it, and a fast-path creator
that wins the briefly-free path simply beats the takeover. The takeover
lock itself is age-swept (a live section lasts microseconds), and handles
only release the lock instance they own.
* fix: verify lock ownership after publish and preserve freshly staged exes
Two more race fixes from review:
- install-lock: the stale-marker sweep repeats the inspect-then-delete
race one level up — two contenders sweeping the same aged takeover
marker could both win and enter the main-lock section together.
Pathname APIs offer no conditional delete, so both the takeover marker
and the main lock now verify ownership after publishing (unique marker
content, read-back compare): a racing sweep converts to a single
survivor instead of two holders. The irreducible residual (a delete
landing in the microsecond link-to-verify window) degrades to a wasted
download cycle, never a corrupt install — swap claims guard the exe
independently.
- native-swap: sweeping a stale swap claim deleted the exe it referenced
even when a FRESH staged.json referenced the same version-derived name
(a downloader re-staged the version after the swap crashed), throwing
away a verified ~180 MB stage. The sweep now preserves any exe the
current staged metadata still references.
* fix: reject mismatched manifests, take over from dead holders, unique .part names
Three robustness fixes from review:
- native-manifest: the per-release endpoint can answer with ANOTHER
release's manifest (stale cache, mispublish); its checksums would then
be applied to this version's binary and fail verification on every
attempt. Compare the parsed manifest version with the requested one.
- install-lock/update-download: a killed lock holder skips its finally
and never releases, stranding a waiting foreground `kimi upgrade`
forever. A lock whose recorded pid is dead is now stale at any age
(the atomic publish guarantees the pid was alive when written), and
the same-version wait loop polls the acquisition itself, so a dead
holder's lock is taken over within one poll instead of never.
Package-manager spawns are unaffected: they hold the lock only around
the spawn, and the active-record bookkeeping guards that layer.
- native-stage: the download intermediate is now unique per worker
(`.part` carries pid + counter), so overlapping same-version workers
can no longer interleave writes into the same file.
* fix: restrict staging cleanup to updater-owned names and retry short writes
- cleanupStagingOrphans recursively deleted anything it did not
recognize; the staging dir sits next to the exe and can contain files
belonging to the user or another tool. Deletion now requires a
positive match on updater-owned artifact names (staged exes and .part
intermediates) and only ever unlinks files.
- FileHandle.write may persist fewer bytes than requested (short write,
e.g. near disk exhaustion) while the running hash and size already
accounted for the whole chunk — publishing a truncated binary under a
valid checksum. The chunk write now loops until fully persisted.
* fix: scope failure cleanup, recognize all semvers, reverify staged checksums
Three fixes from review:
- native-stage failure cleanup deleted whatever staged update was
currently published — including a concurrent worker's valid result
that its caller had already reported as success. The catch path now
removes only this attempt's own artifacts: its unique .part file and
its staged exe name when the current metadata does not reference it.
- The orphan-cleanup ownership check only matched stable x.y.z names;
prerelease/build-metadata versions (1.2.3-rc.1, 1.2.3+build) would
never be cleaned and accumulate ~180 MB each. Ownership now derives
from the semver contract via the semver package's valid().
- The swap path trusted a staged exe whose size matched, though the
metadata records the release checksum; post-download on-disk damage
could pass the --version smoke check with corrupted bytes.
claimStagedUpdate now re-verifies the staged exe's sha256 before
claiming and discards the stage (for a later re-download) on mismatch
— paid only when an update is actually pending.
* fix: validate versions before path derivation and honor the update opt-out in the swap
- native-stage: stageNativeUpdate derived staging paths (including the
cleanup rm targets) from the version before fetchNativeReleaseManifest
rejected it; a traversal string like `x/../../kimi` would resolve the
staged-exe cleanup onto the running installation. The semver check now
happens before any path is derived, and the staged-metadata schema
constrains exeFileName to a plain file name.
- native-swap: the startup swap ran before the update preflight, so
KIMI_CODE_NO_AUTO_UPDATE / KIMI_CLI_NO_AUTO_UPDATE stopped gating
update behavior once a payload was pending. The swap now honors the
same opt-out: the staged payload stays in place for a later launch
without the variable, and the current exe starts.
* fix: restrict backup cleanup to updater-owned .bak names
cleanupBackups treated every <exe>.*.bak sibling as swap residue, so a
user's own backup like kimi.config.bak in a shared bin directory was
silently deleted on startup. Only the exact <exe>.bak and the numeric
PID fallback <exe>.<pid>.bak are updater-created — cleanup now
positively matches those two formats.
* fix: claim staged metadata before validating it and let manual upgrades bypass the opt-out
- native-swap: claimStagedUpdate validated the metadata and hashed the
staged exe BEFORE the atomic rename, so a concurrent downloader
superseding staged.json in between could get its fresh metadata
claimed under the older object — the smoke check then failed and
discard() deleted the newly published stage, recording a failure for
the wrong version. The claim (utimes + rename) now happens first and
validation acts on exactly the claimed file; discards use a new
discardClaimedUpdate that never removes anything a meanwhile-published
stage references.
- The auto-update env opt-out gated the startup swap unconditionally,
so an explicit `kimi upgrade` with the variable set staged the
version but no launch ever applied it. Stages now record
`manual: true` when they answer a user-initiated install
(`__update_download --manual`, threaded from installUpdate through
the hidden sub-command), and the swap applies manual stages even when
automatic updates are opted out.
* fix: promote adopted stages to manual and preserve claim-referenced payloads
Three follow-up fixes from review:
- An explicit `kimi upgrade` adopting an auto-staged payload (already
on disk, or still downloading via the wait path) returned before the
manual marker applied, so under the env opt-out the swap still skipped
it despite the success message. Both adoption paths now promote the
staged metadata to manual: true via a new promoteStagedUpdateToManual.
- The download-failure cleanup checked only the current staged metadata,
but a live swap holds the metadata renamed aside as its claim — a
failing same-version downloader could delete the exe an active swap
was about to move into place. The catch path now also preserves names
referenced by any live swap claim.
- Restoring a claimed stage after a failed exe move used rename, which
on POSIX replaces a newer staged.json a downloader published during
the smoke check. The restore is now a create-if-absent hard link: it
only lands when the state-file path is still free, and the older claim
is discarded when a newer stage has taken it.
* fix: drop exe deletion from stale-claim cleanup
The stale-claim sweep deleted the referenced exe based on a metadata
snapshot taken before the loop; a downloader republishing the same
version between the read and the unlink would have its fresh payload
deleted after reporting success. Publication can never be synchronized
with a pathname-level snapshot, so the sweep now removes only the claim
files themselves — genuinely unreferenced exes are reaped by the
downloader's own orphan cleanup (keep-set aware) before its next stage.
* fix: never delete the staged exe when discarding a claim
The same publication race existed one level down: a same-version
downloader can rename its fresh payload onto the shared exe path after
the discard's metadata snapshot but before the unlink (payloads publish
before their metadata), and the discard would delete a download whose
caller then reports success with nothing behind it. discardClaimedUpdate
now removes only the claimed metadata file; unreferenced exes are reaped
by the downloader's own orphan cleanup before its next stage.
* fix: only reap staging orphans old enough to be abandoned
The orphan sweep could delete a concurrent worker's freshly renamed
staged exe in the gap before its staged.json lands (payloads publish
before their metadata), turning the admitted duplicate-worker race into
a successful stage with no payload behind it. Unreferenced artifacts are
now only deleted once older than a one-hour grace period — publication
takes milliseconds, so unreferenced AND old means definitively
abandoned.
* fix: honor the persisted auto-update preference in the swap and drop claim-unsafe deletions
- The startup swap gated only on the env opt-out, so a payload staged
automatically still installed after the user disabled automatic
updates via [upgrade] auto_install = false. The swap now loads the
persisted preference (only when an automatic stage is actually
pending) and skips it, exactly like the env opt-out; manual stages
still always apply.
- Superseding a staged version deleted its exe through an uncoordinated
read-then-remove that could pull the payload from a live swap. The
supersede now removes only the old metadata record — the metadata
write atomically replaces it, and an unreferenced exe is reaped by a
later orphan cleanup. removeStagedNativeUpdate, left with no callers,
is removed.
- docs: the kimi upgrade reference (en + zh) no longer claims Windows
native installations cannot upgrade automatically; native installs
download and verify in the foreground and swap on the next start.
* fix: gate on claimed metadata, stop shared-path deletes on failure, exact smoke match
- The opt-out gate evaluated a pre-claim snapshot of the staged
metadata, but the claim could pick up a different (automatic) stage a
downloader published in between — smuggling it past the gate. The
env/preference check now runs on the CLAIMED metadata; when disabled,
the claim is restored via create-if-absent link so a newer stage is
never overwritten and a later launch can still apply it. The checksum
re-verify moves after the gates so opted-out launches stop paying for
the hash.
- The download-failure cleanup still deleted the shared staged-exe path
based on snapshot reference checks — the same publication race as the
paths already fixed. It now removes only the attempt's privately owned
.part file; the shared exe is left for the age-gated orphan cleanup.
- The smoke check accepted the staged version as a substring of the
--version output, so a mispublished 1.2.30 binary would satisfy a
1.2.3 target with a matching manifest checksum. It now requires the
trimmed output to equal the staged version exactly.
* fix: confirm the manual marker before reporting stage adoption
promoteStagedUpdateToManual silently no-oped when a startup swap had
claimed the state file, while the adoption paths still reported success
with manual: true synthesized — under the env opt-out the restored
automatic metadata would then be skipped on every later launch despite
the upgrade's success message. The helper now verifies the marker with a
confirming read (one retry) and returns whether it persisted; the
already-staged branch falls through to a fresh stage when it does not,
and the same-version wait loop only adopts after a confirmed promotion.
* fix(cli): verify the staged payload digest before adopting it as already-staged
readStagedNativeUpdate checks only the recorded size, so a same-size
corruption after the download was adopted and reported as success, only
for the startup swap's claim-time re-verify to reject and discard it.
Compare the actual sha256 before returning already-staged; a mismatch
falls through and re-stages from the CDN.
* fix(cli): keep staged metadata until its replacement is ready
Two related races around staged.json, both reported against the
duplicate-downloader residual:
- stageNativeUpdate deleted the previous record before downloading its
replacement; a pathname-only delete can remove a concurrent worker's
freshly published record, orphaning a payload whose worker already
reported success. The old record now stays until the final atomic
metadata write replaces it.
- promoteStagedUpdateToManual wrote the marker unconditionally onto
whichever generation owned staged.json. It now takes the adopted record
and promotes only while the on-disk metadata still matches it, and the
post-write confirmation requires the promoted candidate itself.
* fix(cli): preserve the exe referenced by the current staged record during orphan cleanup
Since the supersede path now keeps the previous staged.json until the
final atomic write replaces it, an aged staged exe is still the
applicable update while its replacement downloads — but
cleanupStagingOrphans only pinned exes referenced by swap claim files,
so a payload older than the grace period was unlinked out from under
its own record. Read staged.json itself in the pinning pass so the
current record's exe is preserved like any live claim's.
* chore(kimi-code): reword the native auto-update changeset
* chore(kimi-code): trim the native auto-update changeset
* fix(cli): support update locking on filesystems without hard links
link() fails with ENOTSUP/ENOSYS/EPERM on FAT/exFAT and some network
mounts, which aborted every native update before the download. Add a
shared createFileIfAbsent primitive (hard-link a fully written temp
file, falling back to an exclusive create + write) and use it for the
install lock, its takeover marker, and the swap's claim restore. The
fallback's create->write gap is observable, so the lock inspection now
grants young unparseable content a publish grace before sweeping it as
crash residue.
* fix(cli): publish staged exes under unique names and recover orphaned claims
Two related robustness fixes in the staged swap flow:
- A staged executable is now published under a unique per-worker name
(kimi-<version>.<pid>.<epoch-ms>.<n>[.exe]) and never replaced; the
atomic metadata write retargets the pointer. The pathname a swap
validates at claim time can no longer be exchanged by a concurrent
same-version publisher between validation and install.
- restoreClaimedUpdate only drops the claim when the restore landed or a
newer stage holds the state-file path; transient failures retain it.
The stale-claim sweep now restores aged claims (create-if-absent)
instead of deleting them, so a stage orphaned by a dead swap or a
transient restore failure is retried on a later launch.
* fix(cli): verify the staged payload digest in the lock-wait adoption path
waitForStagedUpdate relied on readStagedNativeUpdate, which checks only
the recorded size: while a holder re-stages a same-size-corrupted
payload (its metadata is replaced only when the repaired generation
publishes), a waiter could promote and report the corrupt stage as
downloaded, and startup would later reject its checksum. Apply the same
integrity bar as stageNativeUpdate's already-staged path — adopt only a
payload that hashes to its recorded checksum; a mismatch falls through
to the lock poll, which takes over once the holder finishes without
repairing it.
* fix(cli): serialize swap critical sections and preserve in-flight publishes
- The fresh-claim sweep is only a directory snapshot: two processes could
both pass it before either claimed, then rename the same installed exe
concurrently and delete each other's rollback backup. A create-if-absent
swap mutex (swap.lock, age-gated like the takeover marker) now serializes
the executable-renaming section; the loser restores its claim and defers.
The mutex is released as soon as the new exe is in place, before the
re-exec, so it is never held for the child session's lifetime.
- claimStagedUpdate no longer destroys a claimed record that is unparseable
but was young at claim time: on filesystems without hard links the
exclusive-create publish is observable mid-write, and discarding it would
orphan the staged exe while the writer reports success. Such a record is
put back with the same inode so the writer completes it; aged corrupt
residue and well-formed records with a missing/changed exe are still
discarded.
* fix(cli): keep backup cleanup inside the swap mutex
The early release let a subsequent swap rename the just-installed exe to
the shared .bak path while the previous swap's cleanup was still about to
unlink that same path, destroying the second swap's rollback source. The
mutex now covers the backup cleanup; the cosmetic staging-dir rmdir and
the re-exec stay outside it.
* test: replace prose-pinning system-prompt tests with a structural sharing check
The two removed tests pinned exact sentences of the default system prompt
('reversibility and blast radius', 'premature abstraction', optional-tool
phrasings that must not appear, ...). They broke on any intentional
wording change while only catching regressions that reused the same words.
The one real contract underneath — shared, ungated sections must render
byte-identically in the root agent and every subagent profile — is now
checked structurally by slicing the section out of the root prompt and
asserting the other profiles contain it, regardless of its wording.
* test: remove wording-pinning tests of model-facing prose across both suites
Sweep of the class identified in #3030: assertions pinning the exact
English wording of product model-facing text (system prompt, reminder
and injection .md files, tool descriptions, shipped profile/skill
bodies). They break on any intentional rewording yet only catch
regressions that reuse the same words.
Across 35 files (~60 test cases, net -1131 lines):
- deleted dedicated wording tests: 'exposes current metadata and
schema' description pins, goal/plan/todo reminder content tests,
tower skill-body prose pins, goal-outcome.test.ts;
- trimmed wording assertions from behavioral tests that otherwise
stand alone; kept identifiers (tool names, XML tags, section
markers), structural properties (wrapping/escaping/gating/cadence),
fixture data, tool outputs and error messages;
- re-anchored a few gating tests on exported constants
(WINDOWS_PATH_HINT, DEFAULT_REPLY_STYLE_GUIDE) instead of prose
literals.
Deferred for a follow-up decision: ~15 tests whose prose pin is the
only discriminator of which reminder/budget-band fired (constants not
exported). Wire baselines and snapshot machinery untouched.