The 0.9.20 menubar never completed a fetch on a large corpus: the cache
version bump forced a full rehydration, DataClient's fixed 45s kill ended
it mid-transaction, cache-refresh-lock had no signal cleanup so the dead
holder's lock survived, and DEFAULT_WAIT_MS (30s) < DEFAULT_STALE_MS (90s)
meant no waiter could ever recover it.
A. Swift port of #1096's watchdog. CLIWatchdog holds the constants and the
pure verdict; spawns and the resident serve child set CODEBURN_PROGRESS=1;
the window restarts on any stdout/stderr byte; 45s silence, 10min cold
floor until the first payload, 15min ceiling, SIGTERM then SIGKILL after
5s. ServeConnection's fixed 60s warm cap becomes the same silence window,
re-armed by each progress frame, and a spent death budget is now a
5-minute cooldown instead of disabling the resident for the app run.
B. cache-refresh-lock arms SIGINT/SIGTERM cleanup the way session-cache does
for hydrating.lock, so a SIGTERMed holder unlinks its own lock.
C. Staleness also opens on a dead holder pid, and the waiter budget derives
from staleMs so it can never expire before the gate it waits for. A live
holder - fresh heartbeat, pid answers signal 0 - is never taken from.
parser.ts heartbeats through the lock wait, the one silent stretch left.
D. The app closes its end of a retired child's stdin (dropping the handle
left the pipe alive inside the Process), reaps every serve child
synchronously at quit, and records pid+argv so a crash-orphaned child is
reaped next launch. serve's final exit no longer runs through a
monkeypatched process.exit, and its post-drain cleanup is bounded.
Closes#1117
A warm launch rewrote the entire session cache whenever any provider
appended a few KB: on a 6 GB corpus that is a 155 MB stringify + fsync
every run. The on-disk cache is now a version-suffixed directory holding
one shard per provider plus a small envelope, and a save rewrites only
the providers marked dirty.
- Dirtiness is tracked per provider (markCacheDirty) instead of one
global flag, so an appended Claude session no longer republishes
Codex, Copilot and the rest.
- Shards carry a nonce in their filename and the envelope is renamed
last, so a save is published at a single point: readers never see a
half-updated set, and a writer that loses the refresh ownership fence
leaves the canonical shards untouched.
- A shard that fails validation is treated as an absent provider rather
than rejecting the whole cache, so one malformed turn costs one
provider's re-parse instead of every provider's history.
- v7 migrates losslessly: the blob is re-laid-out into shards and
removed only once that save publishes. Nothing re-parses.
- Cold-parse progress saves now trigger every N files parsed rather than
every 5s, so a slow cold parse no longer rewrites the growing cache on
a wall clock.
parser.test.ts (a)/(f): createJsonlSession stamped events at a fixed 2026-05-01
that aged past the 90-day retention window, pruning to zero; date them relative
to now. cli-durable-totals: seedLiveTodaySession stamped noon, which is in the
future on a pre-noon run so the provider-scoped today slice (ends at now)
dropped it while the all path (ends at range end) kept it; seed a past-today
time. cache-refresh-lock and other integration tests starve under a saturated
parallel run and fail closed; add a small global retry and raise the two most
load-sensitive lock tests. Test-only; no production code changed.
observe() classified a stable unparseable session-refresh.lock body as the
terminal 'unavailable'. parseAllSessions routes that to a read-only parse, so a
zero-byte or truncated lock froze warm-cache ingestion permanently across every
later run while each command still exited successfully.
A corrupt body is now a recoverable observation carrying a real mtime, and is
recovered only through the UNMODIFIED staleness gate — tryTakeover and the age
check are byte-identical to main. sameObservation gains an explicit null/non-null
boundary and a sha1 of the raw bytes, because two corrupt bodies have no tokens
to compare and mtime granularity is coarse on some filesystems.
The heartbeat deliberately does NOT rewrite a body it cannot prove is its own.
An owner that cannot prove ownership ends its ownership: mtime stops advancing,
the publication fence refuses, and a successor recovers the lock one staleMs
later. Losing that parse is the price of never having two owners.
The 1ms-heartbeat stress test fails under saturated full-suite runs when
filesystem starvation makes guard operations report unavailable, which
the fence correctly treats as fail-closed. Isolation runs, the process
tests, and mutation checks all hold. Retries go 2 to 5 and iterations
300 to 120: a mutated build still fails every attempt (six percent per
verify compounds to a per-attempt pass chance under a tenth of a
percent), while transient starvation now has five chances to clear.
Co-authored-by: reviewer <review@local>
Under a saturated full-suite run, fd/CPU starvation can make the
takeover-guard fs ops report unavailable, which the fence correctly
treats as fail-closed; the 1ms-heartbeat stress test misread that as the
serialization race. Retries rescue environmental noise only: the real
race fails roughly 6 percent per verify across 300 verifies, so a
mutated build cannot pass any attempt.
* cache: strict cross-process gate for the warm session-cache refresh transaction (#645)
* lock: serialize the fence against the owner's own heartbeat
verifyStillOwner and the heartbeat tick both take the takeover guard;
without in-process serialization the fence could observe its own
heartbeat's guard file, read it as displacement, and abort a legitimate
publication. Fail-safe, but it discarded the parse the lock exists to
protect (~6% of verifies at a 1ms heartbeat in the repro). Owner-side
guard operations now run through one serializer; cross-process guard
semantics are unchanged. Regression test at heartbeatMs 1, mutation-
verified against the unserialized code.
---------
Co-authored-by: Aditya Vikram Singh <247195684+avs-io@users.noreply.github.com>
Co-authored-by: reviewer <review@local>