mirror of
https://github.com/MoonshotAI/kimi-code.git
synced 2026-08-22 15:16:07 +00:00
* feat(kimi-code): support automatic updates for native installations via staged swap Native (SEA) installs previously could not self-update on Windows and relied on 'curl | bash' re-install on Unix. Replace both with a staged swap updater: - startup swaps in a staged binary (verified against the release manifest sha256, smoke-checked via --version) and re-execs it, so the running process never replaces itself (Windows-safe) - downloads run in a self-spawned hidden sub-command, in the background from the update preflight or in the foreground from 'kimi upgrade' - rollback from .bak on any swap failure; install failures keep the existing retry/prompt thresholds * fix(kimi-code): fully clean staged artifacts on swap discard paths Real-binary smoke testing on macOS surfaced two cleanup gaps in the discard path: the claimed metadata file was unlinked after the staging dir rmdir (so the empty dir survived), and the staged exe was rediscovered via the already-claimed staged.json (so it leaked on the downgrade-guard path). Pass the known metadata through and order the unlink before the rmdir. * fix(kimi-code): restore staged metadata on swap failure and sweep update leftovers at startup * fix(kimi-code): address codex review on lock contention and swap crash window - The background native install no longer takes the outer install lock: the self-spawned downloader holds it for the whole download, and the parent's spawn-time lock raced the child into a false lastSuccess. - Smoke-check the staged exe before moving anything, so a bad staged binary is discarded with the install path never left empty; the remaining crash window is two adjacent atomic renames (documented, recoverable via the .bak or by re-running the install script). * test(kimi-code): align swap test expectation with smoke-before-rename order The restore-on-failure case now observes the early smoke check's --version spawn; only the re-exec spawn must be absent. * fix(kimi-code): stage the bare CDN binary instead of unzipping The published per-release artifacts are the bare platform binaries (kimi-code-<target>[.exe]), not zip archives — the staging flow now streams the download straight to the staged exe after the manifest sha256 check, and the zip reader is dropped. Verified end-to-end on macOS against the live CDN: download -> sha256 match -> swap -> re-exec into the real released binary. * fix(kimi-code): address second codex review round - re-exec: forward 128 + signo when the swapped-in child dies by signal instead of reporting exit 0 - __update_download: only exit 0 without staging when the lock holder is staging the SAME version; a different in-flight version (or a vanished lock) no longer surfaces as a successful foreground upgrade - staging: sweep orphaned .part downloads and unreferenced staged exes before downloading, preserving live swap claims and their payloads * feat(kimi-code): show download progress for native updates The foreground 'kimi upgrade' path streamed 180 MB with a single static 'Downloading…' line. Render progress instead: a throttled in-place percentage line on a TTY, one line per 32 MB when piped, and plain MB counts when Content-Length is unknown. * fix(kimi-code): bound native update downloads with an idle timeout Codex review: the manifest fetch cleared its timer once headers arrived, so a stalled response body hung the worker forever, and the binary download had no abort at all. The manifest timeout now covers body consumption, and the binary stream aborts after 30 s without a chunk (total duration stays unbounded for slow networks). The idle timeout is injectable for tests. * fix(kimi-code): retry native updates blocked by an orphaned active record Windows real-machine verification surfaced that a parent exiting before the downloader's exit event leaves a fresh-looking 'active' record that silently blocks every background retry for the 6 h TTL. For native installs, lock liveness is the truth past a 60 s spawn grace window: a held lock means a download is running, a free lock means the record is an orphan and a new attempt may start. Package-manager sources keep the TTL behavior (no lock to prove liveness). * fix: skip staged swap while another instance holds a fresh claim sweepStaleNativeUpdateArtifacts already detected an in-progress swap in a concurrent instance, but the result stayed inside the cleanup helper: startup still claimed a newly published staged.json and ran a second swap, so the two launchers could rename the install path and delete each other's rollback backup. Propagate the in-progress signal and skip claiming until the existing claim is released or goes stale. * fix: keep the install lock while its holder process is alive The install lock went stale purely by age (30 min), but the native downloader is idle-bounded, not duration-bounded: a slow link can legitimately take longer. Another startup would then sweep the lock and spawn a second downloader, and both would write and clean the same .staging paths. Past the age threshold, fall back to a pid liveness probe (signal 0) — the lock is stale only when the holder is gone. * fix: keep recovery artifacts on rollback failure and wait out same-version downloads Two robustness fixes from review: - native-swap: when moving the staged exe into place fails AND the rollback rename fails too (transient lock, AV), the install path is left absent and no next launch can start. Discarding the staged payload and claim on top of that removes the second recovery copy. rollback() now reports its result; on a double failure the swap keeps the .bak (which IS the old exe), the staged exe and the claim so manual recovery or a re-install still works. - update-download: a foreground `kimi upgrade` racing a background downloader of the same version exited 0 immediately, so the CLI printed a success message for a download that could still fail. The worker now waits while the same-version holder is in flight, adopts the verified staged result (staged.json lands before the lock is released), and takes over the download when the holder finished without staging. * fix: stamp the swap claim with a fresh mtime when claiming rename() preserves the staged metadata's mtime, which can be arbitrarily old — the background download often finishes hours before the next launch claims it. A concurrent launch's sweep would then classify the live claim as crash residue (older than the 5-minute window) and delete the claim, the staged exe, and eventually the first swap's rollback backup. Stamp the claim file with the claim time so the staleness check measures the swap's liveness, not the download's age. * fix: stamp the claim before the rename so it is born fresh Stamping after the rename left a window: a concurrent launch could inspect the claim between the two syscalls, see the staged metadata's old mtime, and delete the staged executable mid-swap. utimes the state file first so the claim carries a fresh timestamp from the instant it is published — no fresh-looking-later intermediate state exists. * fix: chmod the staged download before publishing it at its final name A swap claims only the staged METADATA; the staged exe stays in .staging/. A concurrent same-version downloader (possible because swaps do not hold the install lock) then re-downloads and renames its .part over that path. If the swap moves the file into the install path between the downloader's rename and its post-publish chmod, the chmod lands on a path that is already gone and the installation is left non-executable — every future launch fails. Apply the executable mode to the private .part file before the publishing rename so the staged exe is executable from the instant it appears. * fix: publish the install lock atomically via hard link The 'wx' open exposed a momentarily empty lock file before its contents were written. A concurrent acquirer reading in that window got a SyntaxError, treated the lock as stale, swept it and also won — two "holders" then ran stageNativeUpdate against the same .staging paths. Write the lock contents to a unique temp file and hard-link it into place: link() fails when the destination exists (same exclusivity as 'wx') and the lock path only ever appears fully written. * fix: serialize stale-lock takeover through a secondary lock A pathname-level delete can never be conditioned on the file still being the inspected stale instance, so a plain compare-and-delete still loses exclusivity: two workers classifying the same stale lock could interleave unlink and publish such that both won (proven by a 20-way contention test). Takeovers now go through a secondary create-if-absent lock (install.lock.takeover): the delete+publish section only ever runs in one process, staleness is re-validated inside it, and a fast-path creator that wins the briefly-free path simply beats the takeover. The takeover lock itself is age-swept (a live section lasts microseconds), and handles only release the lock instance they own. * fix: verify lock ownership after publish and preserve freshly staged exes Two more race fixes from review: - install-lock: the stale-marker sweep repeats the inspect-then-delete race one level up — two contenders sweeping the same aged takeover marker could both win and enter the main-lock section together. Pathname APIs offer no conditional delete, so both the takeover marker and the main lock now verify ownership after publishing (unique marker content, read-back compare): a racing sweep converts to a single survivor instead of two holders. The irreducible residual (a delete landing in the microsecond link-to-verify window) degrades to a wasted download cycle, never a corrupt install — swap claims guard the exe independently. - native-swap: sweeping a stale swap claim deleted the exe it referenced even when a FRESH staged.json referenced the same version-derived name (a downloader re-staged the version after the swap crashed), throwing away a verified ~180 MB stage. The sweep now preserves any exe the current staged metadata still references. * fix: reject mismatched manifests, take over from dead holders, unique .part names Three robustness fixes from review: - native-manifest: the per-release endpoint can answer with ANOTHER release's manifest (stale cache, mispublish); its checksums would then be applied to this version's binary and fail verification on every attempt. Compare the parsed manifest version with the requested one. - install-lock/update-download: a killed lock holder skips its finally and never releases, stranding a waiting foreground `kimi upgrade` forever. A lock whose recorded pid is dead is now stale at any age (the atomic publish guarantees the pid was alive when written), and the same-version wait loop polls the acquisition itself, so a dead holder's lock is taken over within one poll instead of never. Package-manager spawns are unaffected: they hold the lock only around the spawn, and the active-record bookkeeping guards that layer. - native-stage: the download intermediate is now unique per worker (`.part` carries pid + counter), so overlapping same-version workers can no longer interleave writes into the same file. * fix: restrict staging cleanup to updater-owned names and retry short writes - cleanupStagingOrphans recursively deleted anything it did not recognize; the staging dir sits next to the exe and can contain files belonging to the user or another tool. Deletion now requires a positive match on updater-owned artifact names (staged exes and .part intermediates) and only ever unlinks files. - FileHandle.write may persist fewer bytes than requested (short write, e.g. near disk exhaustion) while the running hash and size already accounted for the whole chunk — publishing a truncated binary under a valid checksum. The chunk write now loops until fully persisted. * fix: scope failure cleanup, recognize all semvers, reverify staged checksums Three fixes from review: - native-stage failure cleanup deleted whatever staged update was currently published — including a concurrent worker's valid result that its caller had already reported as success. The catch path now removes only this attempt's own artifacts: its unique .part file and its staged exe name when the current metadata does not reference it. - The orphan-cleanup ownership check only matched stable x.y.z names; prerelease/build-metadata versions (1.2.3-rc.1, 1.2.3+build) would never be cleaned and accumulate ~180 MB each. Ownership now derives from the semver contract via the semver package's valid(). - The swap path trusted a staged exe whose size matched, though the metadata records the release checksum; post-download on-disk damage could pass the --version smoke check with corrupted bytes. claimStagedUpdate now re-verifies the staged exe's sha256 before claiming and discards the stage (for a later re-download) on mismatch — paid only when an update is actually pending. * fix: validate versions before path derivation and honor the update opt-out in the swap - native-stage: stageNativeUpdate derived staging paths (including the cleanup rm targets) from the version before fetchNativeReleaseManifest rejected it; a traversal string like `x/../../kimi` would resolve the staged-exe cleanup onto the running installation. The semver check now happens before any path is derived, and the staged-metadata schema constrains exeFileName to a plain file name. - native-swap: the startup swap ran before the update preflight, so KIMI_CODE_NO_AUTO_UPDATE / KIMI_CLI_NO_AUTO_UPDATE stopped gating update behavior once a payload was pending. The swap now honors the same opt-out: the staged payload stays in place for a later launch without the variable, and the current exe starts. * fix: restrict backup cleanup to updater-owned .bak names cleanupBackups treated every <exe>.*.bak sibling as swap residue, so a user's own backup like kimi.config.bak in a shared bin directory was silently deleted on startup. Only the exact <exe>.bak and the numeric PID fallback <exe>.<pid>.bak are updater-created — cleanup now positively matches those two formats. * fix: claim staged metadata before validating it and let manual upgrades bypass the opt-out - native-swap: claimStagedUpdate validated the metadata and hashed the staged exe BEFORE the atomic rename, so a concurrent downloader superseding staged.json in between could get its fresh metadata claimed under the older object — the smoke check then failed and discard() deleted the newly published stage, recording a failure for the wrong version. The claim (utimes + rename) now happens first and validation acts on exactly the claimed file; discards use a new discardClaimedUpdate that never removes anything a meanwhile-published stage references. - The auto-update env opt-out gated the startup swap unconditionally, so an explicit `kimi upgrade` with the variable set staged the version but no launch ever applied it. Stages now record `manual: true` when they answer a user-initiated install (`__update_download --manual`, threaded from installUpdate through the hidden sub-command), and the swap applies manual stages even when automatic updates are opted out. * fix: promote adopted stages to manual and preserve claim-referenced payloads Three follow-up fixes from review: - An explicit `kimi upgrade` adopting an auto-staged payload (already on disk, or still downloading via the wait path) returned before the manual marker applied, so under the env opt-out the swap still skipped it despite the success message. Both adoption paths now promote the staged metadata to manual: true via a new promoteStagedUpdateToManual. - The download-failure cleanup checked only the current staged metadata, but a live swap holds the metadata renamed aside as its claim — a failing same-version downloader could delete the exe an active swap was about to move into place. The catch path now also preserves names referenced by any live swap claim. - Restoring a claimed stage after a failed exe move used rename, which on POSIX replaces a newer staged.json a downloader published during the smoke check. The restore is now a create-if-absent hard link: it only lands when the state-file path is still free, and the older claim is discarded when a newer stage has taken it. * fix: drop exe deletion from stale-claim cleanup The stale-claim sweep deleted the referenced exe based on a metadata snapshot taken before the loop; a downloader republishing the same version between the read and the unlink would have its fresh payload deleted after reporting success. Publication can never be synchronized with a pathname-level snapshot, so the sweep now removes only the claim files themselves — genuinely unreferenced exes are reaped by the downloader's own orphan cleanup (keep-set aware) before its next stage. * fix: never delete the staged exe when discarding a claim The same publication race existed one level down: a same-version downloader can rename its fresh payload onto the shared exe path after the discard's metadata snapshot but before the unlink (payloads publish before their metadata), and the discard would delete a download whose caller then reports success with nothing behind it. discardClaimedUpdate now removes only the claimed metadata file; unreferenced exes are reaped by the downloader's own orphan cleanup before its next stage. * fix: only reap staging orphans old enough to be abandoned The orphan sweep could delete a concurrent worker's freshly renamed staged exe in the gap before its staged.json lands (payloads publish before their metadata), turning the admitted duplicate-worker race into a successful stage with no payload behind it. Unreferenced artifacts are now only deleted once older than a one-hour grace period — publication takes milliseconds, so unreferenced AND old means definitively abandoned. * fix: honor the persisted auto-update preference in the swap and drop claim-unsafe deletions - The startup swap gated only on the env opt-out, so a payload staged automatically still installed after the user disabled automatic updates via [upgrade] auto_install = false. The swap now loads the persisted preference (only when an automatic stage is actually pending) and skips it, exactly like the env opt-out; manual stages still always apply. - Superseding a staged version deleted its exe through an uncoordinated read-then-remove that could pull the payload from a live swap. The supersede now removes only the old metadata record — the metadata write atomically replaces it, and an unreferenced exe is reaped by a later orphan cleanup. removeStagedNativeUpdate, left with no callers, is removed. - docs: the kimi upgrade reference (en + zh) no longer claims Windows native installations cannot upgrade automatically; native installs download and verify in the foreground and swap on the next start. * fix: gate on claimed metadata, stop shared-path deletes on failure, exact smoke match - The opt-out gate evaluated a pre-claim snapshot of the staged metadata, but the claim could pick up a different (automatic) stage a downloader published in between — smuggling it past the gate. The env/preference check now runs on the CLAIMED metadata; when disabled, the claim is restored via create-if-absent link so a newer stage is never overwritten and a later launch can still apply it. The checksum re-verify moves after the gates so opted-out launches stop paying for the hash. - The download-failure cleanup still deleted the shared staged-exe path based on snapshot reference checks — the same publication race as the paths already fixed. It now removes only the attempt's privately owned .part file; the shared exe is left for the age-gated orphan cleanup. - The smoke check accepted the staged version as a substring of the --version output, so a mispublished 1.2.30 binary would satisfy a 1.2.3 target with a matching manifest checksum. It now requires the trimmed output to equal the staged version exactly. * fix: confirm the manual marker before reporting stage adoption promoteStagedUpdateToManual silently no-oped when a startup swap had claimed the state file, while the adoption paths still reported success with manual: true synthesized — under the env opt-out the restored automatic metadata would then be skipped on every later launch despite the upgrade's success message. The helper now verifies the marker with a confirming read (one retry) and returns whether it persisted; the already-staged branch falls through to a fresh stage when it does not, and the same-version wait loop only adopts after a confirmed promotion. * fix(cli): verify the staged payload digest before adopting it as already-staged readStagedNativeUpdate checks only the recorded size, so a same-size corruption after the download was adopted and reported as success, only for the startup swap's claim-time re-verify to reject and discard it. Compare the actual sha256 before returning already-staged; a mismatch falls through and re-stages from the CDN. * fix(cli): keep staged metadata until its replacement is ready Two related races around staged.json, both reported against the duplicate-downloader residual: - stageNativeUpdate deleted the previous record before downloading its replacement; a pathname-only delete can remove a concurrent worker's freshly published record, orphaning a payload whose worker already reported success. The old record now stays until the final atomic metadata write replaces it. - promoteStagedUpdateToManual wrote the marker unconditionally onto whichever generation owned staged.json. It now takes the adopted record and promotes only while the on-disk metadata still matches it, and the post-write confirmation requires the promoted candidate itself. * fix(cli): preserve the exe referenced by the current staged record during orphan cleanup Since the supersede path now keeps the previous staged.json until the final atomic write replaces it, an aged staged exe is still the applicable update while its replacement downloads — but cleanupStagingOrphans only pinned exes referenced by swap claim files, so a payload older than the grace period was unlinked out from under its own record. Read staged.json itself in the pinning pass so the current record's exe is preserved like any live claim's. * chore(kimi-code): reword the native auto-update changeset * chore(kimi-code): trim the native auto-update changeset * fix(cli): support update locking on filesystems without hard links link() fails with ENOTSUP/ENOSYS/EPERM on FAT/exFAT and some network mounts, which aborted every native update before the download. Add a shared createFileIfAbsent primitive (hard-link a fully written temp file, falling back to an exclusive create + write) and use it for the install lock, its takeover marker, and the swap's claim restore. The fallback's create->write gap is observable, so the lock inspection now grants young unparseable content a publish grace before sweeping it as crash residue. * fix(cli): publish staged exes under unique names and recover orphaned claims Two related robustness fixes in the staged swap flow: - A staged executable is now published under a unique per-worker name (kimi-<version>.<pid>.<epoch-ms>.<n>[.exe]) and never replaced; the atomic metadata write retargets the pointer. The pathname a swap validates at claim time can no longer be exchanged by a concurrent same-version publisher between validation and install. - restoreClaimedUpdate only drops the claim when the restore landed or a newer stage holds the state-file path; transient failures retain it. The stale-claim sweep now restores aged claims (create-if-absent) instead of deleting them, so a stage orphaned by a dead swap or a transient restore failure is retried on a later launch. * fix(cli): verify the staged payload digest in the lock-wait adoption path waitForStagedUpdate relied on readStagedNativeUpdate, which checks only the recorded size: while a holder re-stages a same-size-corrupted payload (its metadata is replaced only when the repaired generation publishes), a waiter could promote and report the corrupt stage as downloaded, and startup would later reject its checksum. Apply the same integrity bar as stageNativeUpdate's already-staged path — adopt only a payload that hashes to its recorded checksum; a mismatch falls through to the lock poll, which takes over once the holder finishes without repairing it. * fix(cli): serialize swap critical sections and preserve in-flight publishes - The fresh-claim sweep is only a directory snapshot: two processes could both pass it before either claimed, then rename the same installed exe concurrently and delete each other's rollback backup. A create-if-absent swap mutex (swap.lock, age-gated like the takeover marker) now serializes the executable-renaming section; the loser restores its claim and defers. The mutex is released as soon as the new exe is in place, before the re-exec, so it is never held for the child session's lifetime. - claimStagedUpdate no longer destroys a claimed record that is unparseable but was young at claim time: on filesystems without hard links the exclusive-create publish is observable mid-write, and discarding it would orphan the staged exe while the writer reports success. Such a record is put back with the same inode so the writer completes it; aged corrupt residue and well-formed records with a missing/changed exe are still discarded. * fix(cli): keep backup cleanup inside the swap mutex The early release let a subsequent swap rename the just-installed exe to the shared .bak path while the previous swap's cleanup was still about to unlink that same path, destroying the second swap's rollback source. The mutex now covers the backup cleanup; the cosmetic staging-dir rmdir and the re-exec stay outside it. |
||
|---|---|---|
| .. | ||
| .vitepress | ||
| en | ||
| media | ||
| public | ||
| zh | ||
| .gitignore | ||
| AGENTS.md | ||
| index.md | ||
| package.json | ||