mirror of
https://github.com/openclaw/openclaw.git
synced 2026-10-03 01:29:56 +00:00
fix(gateway): derive the darwin stop budget from the launchd job (#157007)
* fix(gateway): derive the darwin stop budget from the launchd job resolveGatewayShutdownBudget derives the stop deadline from restart ownership, so a darwin host running with OPENCLAW_SUPERVISOR_MODE=external resolves drain=315000ms whatever launchd actually enforces. The linux-gated systemd probe becomes a platform dispatch, and a new readLaunchdStopTimeout mirrors readSystemdStopTimeout: it reads the running job's effective exit timeout and accepts it only when the printed pid is this process. run-loop.ts is untouched because the reader detects launchd from the environment itself. Closes #156968 * fix(gateway): accept the launchd launcher parent and cap its deadline The darwin reader accepted the printed job only when its pid was this process. The installed service can keep a launcher parent while the serving Gateway runs as its child, so the job prints the launcher's pid, the enforcing job was rejected, and the Gateway fell back to 20 seconds in the one layout where an operator's ExitTimeOut was meant to apply. Accepting that job at face value would be worse than rejecting it. node-runtime-recovery.mjs builds the launcher's own reap timer from the compile-time LAUNCH_AGENT_EXIT_TIMEOUT_SECONDS rather than from the job, so on a forwarded stop it re-sends SIGTERM to its child at 18000ms and SIGKILLs it at 19000ms. Adopting a 90 second job deadline in that layout would plan a 75000ms drain and then lose it to its own parent, which is the truncated drain this change exists to prevent. The reader now accepts process.pid or process.ppid and takes min(jobExitTimeout, 20000) in the launcher case, naming the cap in the source string only when it binds. A shorter job deadline still applies in both layouts, because launchd reaps the whole job regardless of the launcher. The docs claimed the deadline is read at startup and when accepting shutdown. run-loop.ts gates the budget refresh on linux at both consumption sites, so on darwin it is read once, at startup, and nativeStopBudget has no shutdown-time consumer there. The page now says startup-only and drops the sentence about a failed shutdown reread retaining the startup budget, which described linux behaviour only. * fix(gateway): scope the darwin stop budget to launchd-driven stops Address the two review findings on the previous revision. The launcher parent is now identified by an explicit OPENCLAW_LAUNCHER_STOP_TIMEOUT_MS declaration rather than a bare ppid match, so an unknown parent keeps the job's full deadline instead of being capped on a guess. The darwin reader inspects the launchd job only while a stop is under way, and adopts that job's deadline only when the job reports the SIGTERMed state. An external SIGTERM keeps the platform-neutral drain, because launchd's ExitTimeOut bounds only a stop launchd itself is running. The job state is read with an anchored single-tab pattern. launchctl print emits further "state = active" lines inside the resource and jetsam coalition blocks, and the shared key/value parser keeps the last occurrence, so the shared parser could never observe SIGTERMed. A test fixture reproduces the nested coalition shape. * fix(gateway): keep a failed launchd inspection off the native stop budget `readLaunchdStopTimeout` answered a failed job inspection with the Gateway's own 330000ms stop policy wrapped in a non-null result. `resolveGatewayShutdownBudget` classifies every non-null result as a native stop budget, so a probe that established nothing still set `nativeStopBudget`. On an externally supervised darwin Gateway that caps a longer requested restart drain and can arm a forced exit against a deadline no supervisor was confirmed to be enforcing. The reader now reports its two answers separately. `stop` carries a deadline only while launchd is confirmed to be stopping the job; `warning` reaches the operator either way. A failed inspection reports `stop: null` and leaves the caller on the platform-neutral policy it had already resolved, which is the same number as before, now classified honestly. The reader no longer imports GATEWAY_SERVICE_STOP_TIMEOUT_MS, because choosing the Gateway's policy number was never its job. Tests assert the flag and the downstream restart drain in the failure case, paired with a confirmed-deadline case so deleting the read outright cannot pass both. Also fixes TS2532 on the job-state regex: the named group is optional under `noUncheckedIndexedAccess`, so `.trim()` needed the optional chain. * test(infra): isolate the hoisted spawn mock and ratchet the OPENCLAW_* count `spawn` is hoisted once per file, so its call log survived across cases and `toHaveBeenCalledExactlyOnceWith` could only hold for the first one. Clear it in `beforeEach`. CI reports `OPENCLAW_* count 488 exceeds budget 487; update config/env-var-count-budget.txt` as a warning. The new name is the launcher's OPENCLAW_LAUNCHER_STOP_TIMEOUT_MS declaration, so the count is correct and the ratchet moves to 488 with the reason recorded in the file. * style(infra): apply oxfmt to the new launchd reader assertions Run by the repo's own `pnpm run format`; joins one over-wrapped assertion line. `pnpm run format:check` then reports "All matched files use the correct format." * test(cli): mock the launchd reader and split the darwin run-loop cases out Activating the darwin stop-budget refresh made `run-loop.test.ts` exercise the real `readLaunchdStopTimeout`, which spawns three `launchctl print` calls against live subprocess I/O while the suite holds `vi.useFakeTimers()`. One case ran 120039ms before failing and the abandoned loop leaked into later cases, for eight failures in total. All eight clear by mocking the reader; no existing expectation was wrong. The mock could not simply be added. `run-loop.test.ts` is 3813 lines against a 1000 line cap where `check-line-cap-ratchet` rejects any growth, and an earlier attempt failed with `3501 -> 3510 counted lines`. Taking the remedy that ratchet names, the darwin cases move to `run-loop.launchd.test.ts` and the original file shrinks by 82 lines. It keeps the mock too, because four of those cases are registered into it by shared `register*` helpers whose `it.each` arrays interleave systemd and launchd variants and cannot be split without touching those support files. The new file adds one case the default mock would otherwise hide: a job reporting a 30000ms deadline bounds the stop at 25000ms and arms the force exit there. `node-runtime-recovery.test.ts` asserts the replacement child's spawn env exactly, so it gains the launcher's own OPENCLAW_LAUNCHER_STOP_TIMEOUT_MS declaration rather than a loosened matcher. That argv is a non-foreground `doctor --fix`, so the value is the 1s signal exit grace plus the 1s force-kill grace. * fix(gateway): only retain a startup budget when the probe was inconclusive Extending the stop-budget refresh to darwin gave the retained-budget safety net a new and wrong trigger. On a launchd-owned Gateway the startup budget is already native, and an in-process restart signals the process without launchd running the stop, so the job still prints `running` and the reader reports `stop: null`. The old condition treated any null as unconfirmed, so every in-process restart of a default macOS install logged "Retaining the startup shutdown budget of 15000ms because the current supervisor stop timeout could not be confirmed." The timeout was confirmed; it was confirmed not to apply. `readNativeStopTimeout` now reports `inconclusive` alongside the deadline, and only an inconclusive probe retains. Linux keeps its exact previous condition, an absent unit or a warned read. The budget number was already unaffected, so this is the message and the wasted `launchctl print` calls, not the deadline. Two cases pin it: a confirmed not-stopping read warns nothing, and a failed inspection still retains and still says so. Also fixes two wording defects found in review. `unresolved()` no longer takes a label it only ever interpolated as "the launchd job the configured label"; each failure already names its own target. And the `readJobState` comment said `state` appears four times in `launchctl print` output when its own next clause counts three, which is what a live LaunchDaemon prints. * style(cli): apply oxfmt to the new retained-budget assertions Run by the repo's own `pnpm run format`; joins one over-wrapped assertion line. Whitespace only, in a test file, so it cannot change runtime behaviour. * fix(gateway): a defaulted deadline is not an inconclusive probe `inconclusive` keyed off the warning alone, so the defaulted-value case counted as a failed probe. There launchd is confirmed to be stopping the job and only its deadline had to be guessed, so a clock is genuinely running and there is nothing to retain. Under launchd ownership that combination logged both "using 20000ms default" and "could not be confirmed" for the same stop, the second being false. Now only a warning with no deadline at all counts as inconclusive. A third case pins it, asserting the defaulted read warns exactly once and about the deadline rather than about confirmation. Also corrects a test fixture comment that still said `launchctl print` emits `state` four times; the arrangement it documents, and a live LaunchDaemon, both show three. * test(cli): correct the regression guard's account of the reported drain The comment said active work "had time to finish" during the reported 315 second drain. The reporting host's log shows the opposite: that drain hit its own timeout with five tasks still active, and the Sep 20 stop with six. The drain was being used, which is a stronger reason for the guard, not a weaker one. * fix(infra): stop exporting a type nothing outside the module reads Knip's all-exports gate reads LaunchdStopTimeout as dead: the only public surface is readLaunchdStopTimeout's LaunchdStopRead return, and nothing imports the nested shape by name. Keep it module-local. * fix(gateway): keep a proportional drain and cap an undeclared launcher The fixed 10s reserve and 5s exit margin were sized against the 315s policy drain. Subtracting them outright from a launchd job's own ExitTimeOut spent a short deadline entirely on overhead: 5 seconds funded the margin alone and 15 funded margin plus reserve, so active work drained for zero milliseconds in both. Each allowance now takes at most a share of what it is carved from, so a deadline long enough to fund them is unchanged and a shorter one keeps a proportional drain. Drain is now positive for every positive deadline and never decreases as ExitTimeOut grows. The shutdown log also subtracted the unscaled reserve constant when reporting drain, which understated a short budget by the amount the scaled reserve gives back, so it now reports the margin and drain actually spent. An already-running launcher published before OPENCLAW_LAUNCHER_STOP_TIMEOUT_MS existed arms the same reap timer and declares nothing, which is the upgrade shape: replacing files cannot change a launcher that is already running. Reading that silence as "no deadline" let a long custom ExitTimeOut be budgeted past a force-kill the parent was already counting down. The launcher's arithmetic moves into the shared budget module so the serving Gateway reconstructs the same timer, used only where the printed job carries OpenClaw's own label and its pid is this process's immediate parent. In that position the parent is the process launchd started for OpenClaw's job and the recovery launcher is the only path that puts a Gateway underneath it, so an unrelated process manager holding the parent slot still caps nothing. * fix(infra): gate the launcher reconstruction on the recovery respawn marker Holding the launchd job's pid is not evidence of a reap timer. An operator wrapper can keep that pid and start the Gateway itself while running none, so reconstructing a cap from the parent relation alone would cut a drain nothing was going to interrupt. The recovery launcher has stamped OPENCLAW_NODE_UPDATE_RESPAWNED on every child it respawns since long before it declared a stop timer, so that marker is present in exactly the upgrade case and absent for any other parent. * fix(infra): declare the new shared budget helpers for the type checker The root module is typed by a hand-written declaration file rather than being compiled, so new exports are invisible to src without being declared there. * docs(gateway): scope the positive-drain claim to the allocation The elapsed cost of resolving the budget is debited after the allocation, so a deadline shorter than that cost can still leave nothing to spend. Claiming every positive deadline yields a positive drain overstated it. * test(cli): accept the retained-budget attribution on a synthetic launchd label The fixture declares a launchd label with no matching job, so the per-stop re-inspection cannot read one and the startup budget is retained. The deadline is unchanged at 15000ms; only the source attribution differs. The assertion still fails if the budget comes from the platform-neutral policy instead. * fix(gateway): derive the launcher's reap deadline instead of declaring it Measurement showed the declared value and the derived value are the same number: a published launcher and a candidate launcher both arm an 18000ms exit grace, and a candidate Gateway resolved an identical 19000ms cap and 14231ms budget under each. The declaration was therefore carrying no information the child could not compute, while adding an OPENCLAW_* name and leaving the upgrade path conditional on which build of the launcher happened to be running. Both sides now read one expression in gateway-shutdown-budget.mjs, the launcher to arm its escalation and the Gateway to bound its budget, so they cannot disagree and there is no published-versus-candidate launcher distinction left to prove. The cap stays gated on OPENCLAW_NODE_UPDATE_RESPAWNED so a parent that did not respawn this process still caps nothing. Drops the env-var count back to 486, so the name ratchet no longer needs an owner waiver, and takes config/env-var-count-budget.txt out of the diff entirely. Test deadlines move off 90 seconds onto 55, because launchd clamps ExitTimeOut at 60 and a 90-second case cannot occur on macOS 27. * test(infra): restore the recovery spawn-env cases to their base form The launcher no longer declares a deadline, so the assertions these cases grew have nothing left to check and their explanatory comment described behavior that is gone. * fix(gateway): cap every respawning launcher, not just the Node-recovery one All three runRespawnedChild call sites reach the same launcher through the same function and arm the same escalation, but only the Node-recovery one sets OPENCLAW_NODE_UPDATE_RESPAWNED. Gating on that marker alone left the two compile-cache respawns uncapped, and the packaged one can wrap a foreground gateway run on an installed service: the job's deadline would then be budgeted past a force-kill the parent had already armed, which is the hazard this cap exists to prevent. The marker set moves next to the deadline it authorises, so adding a respawn path cannot silently escape the cap. Keeping the literals in the root module also keeps them out of the env-count ratchet's scope: reading the marker from counted production source promoted a previously test-only name and held the count at 487 even after the declared-timer variable was removed. Also applies oxfmt's own wrap to the launcher-relation expression. * fix(infra): re-export the respawn marker set through the infra barrel The constant existed only on the root module, so the test importing it from the barrel got undefined and it.each(undefined) threw during collection: the suite reported zero tests rather than a failure, and typecheck flagged the missing member. Re-exporting the binding adds no OPENCLAW_* literal under src, so the env-var count stays at 486. * fix(gateway): address the pre-review findings on the launchd stop budget Docs were stale or wrong in four places. The systemd section still promised a fixed 10 second reserve and 5 second margin, but the share caps are platform-neutral and do change a Linux unit whose TimeoutStopSec is under 25 seconds; that radius is now stated with the 20 second and 15 second cases spelled out. The direct-SIGTERM paragraph claimed the platform-neutral drain "is the deadline that actually governs that stop", which the branch's own kill -TERM capture contradicts: no supervisor deadline governs that stop at all and the fallback is the template value. The updated-launcher requirement no longer applies to the stop budget and says so. The marker sentence named one marker when three are honoured, and one measured figure was stale. Two claims in the code were too strong. The marker list asserted every setter arms the shared escalation, which is false for the compile-cache name: entry.compile-cache sets it for a runner that reaps on a fixed short grace. That runner refuses a foreground Gateway run off Windows and this deadline is only read during a darwin stop, so it cannot be the parent here, and the comment now says that instead of implying exclusivity. The sync note next to the escalation now records that the other runner keeps its own copies of the graces. An absent job state was indistinguishable from a job launchd is not stopping, so a macOS that printed the block differently would silently revert to the platform-neutral policy. When a deadline parsed but the state did not, that now warns and marks the read inconclusive rather than passing as a positive answer. Tests: launchd cases move off 90 and 315 second deadlines, which launchd clamps to 60 and cannot produce; the marker list is restated locally and tied back to the implementation so dropping a marker fails; and the real-process case now asserts the darwin probe actually ran, which it could not before. Retires five exports with no production consumer that the dead-export scan flagged. * test(infra): move the empty job state onto the missing-state warning An empty value yields no state to recognise, so it reaches the same branch as a missing line and now carries the same warning. Keeping it in the unrecognised-state table asserted the absence of a field the reader deliberately reports. * fix(infra): treat a whitespace-only job state as missing, and sharpen two docs claims The state value is trimmed after matching, so a line carrying only whitespace yielded an empty string rather than undefined and skipped the warning while behaving exactly like a missing line. The systemd section now says the exit margin shrinks below a 20-second deadline, not just the reserve. The direct-signal paragraph distinguished only the launchd-supervised budget; external supervisor mode keeps the platform-neutral policy and arms no force-exit timer, and a signal delivered to the job pid does reach a resident recovery launcher, which arms its own reap timer even though launchd is not stopping the job. * docs(gateway): correct the loaded launchd job deadline guidance launchctl print reports the job launchd has loaded, not the plist on disk, so editing a loaded job's ExitTimeOut changes nothing until the job is reloaded. Measured on macOS 27 against a scratch job: the plist moves from 20 to 47 while the loaded job keeps reporting 20, and it still reports 20 after launchctl kickstart -k, which restarts the process without reloading the job. Only bootout then bootstrap makes it report 47, and that restarts the Gateway with it. State what reading per stop actually buys instead: the Gateway never plans against a deadline it cached at its own startup. * fix(gateway): keep the full cleanup reserve wherever a deadline funds it The reserve was capped at half the shutdown budget whenever the budget could not fund it outright. That kept a drain on short deadlines, but it also reallocated deadlines that already worked: a 20 second ExitTimeOut, which is both the shipped LaunchAgent template and launchd's own default, moved from a 10 second reserve and a 5 second drain to 7.5 seconds of each. Cleanup costing between 7.5 and 10 seconds finished before and would have been cut off. Bound the reserve by what keeping a drain actually requires instead. The floor is 5 seconds, which is the drain the 20 second template already yields once the fixed margin and reserve are subtracted, so funding the full reserve alongside it takes 15 seconds of budget: exactly what that template resolves to. Every deadline from 20 seconds up therefore keeps the allocation it had, and only a deadline the fixed subtraction had already driven under a 5 second drain gives any reserve up, never below half the budget. Both allocations stay monotone in the deadline, and the 5 second case is unchanged at 1875ms of each. * fix(gateway): correct shutdown-budget probe-debit overclaims - gateway-shutdown-budget.mjs: replace the floor-invariant claim with the real bound (reserve unchanged only at a >=15000ms post-margin budget, i.e. a >=20s deadline once the launchd probe's own cost, up to three launchctl print calls at 2000ms each, is subtracted); name the 8-20s sub-band reserve loss the old comment denied, including 16-19s where a flat subtraction already left a positive drain while still funding the reserve in full - docs/gateway/restart-recovery.md: drop the systemd section's "down to the millisecond" promise and the launchd section's "exact allocation ... before this change" claim; both now state the inspection cost that comes off instead - run-loop-shutdown-budget.test.ts: add a real 13ms elapsed-debit case at a 20s exit timeout (asserts reserve=9987, drain=5000, contrasting with the existing acceptedAtMs=MAX_SAFE_INTEGER zero-debit fixtures) and mirror darwin's 19/16/15/10s sub-20s it.each against a systemd stub, which previously had no sub-20s coverage * test(cli): wrap the flat-subtraction assertion to satisfy oxfmt The sub-20-second systemd cases added in 60e5eb509b6d left the flat-subtraction comparison on a single line past the formatter's width, which failed oxfmt --check and took check-lint and check-docs down with it. Lint itself reported zero warnings and zero errors. Formatting only, no assertion or value changed. * docs(gateway): bound the 20-second allocation claim by the inspection debit The previous wording paired 'gives up exactly what the probe cost' with a probe bounded near 6 seconds, and those two cannot both hold. Which allowance pays the debit turns on a 5 second threshold: at or under it the reserve absorbs the whole cost and the 5 second drain floor is untouched, so a 20 second deadline resolves 10000 minus the debit. Past it the floor is share-bounded as well and the two converge on half the remainder, so a 6 second debit splits 9000 into 4500/4500 rather than retaining the old reserve. Documents the threshold in the docs and the root-module comment, and pins the over-threshold case with a test so the boundary is measured rather than reasoned. * docs(gateway): qualify the full-reserve claim by the probe debit The drain-floor comment asserted that a whole 10 second reserve follows for every ExitTimeOut of 20 seconds or more once the probe that reads the deadline is subtracted. The probe cost is charged before the budget is split, so a 20 second ExitTimeOut clears the 15 second budget that a whole reserve needs only when that probe costs nothing. At the 13ms measured on this rig the reserve is 9987, and a whole reserve at that cost needs 20013. The same comment already states this correctly six lines down, where it resolves a 20 second deadline to 10000 - cost. This removes the contradiction by qualifying the claim at its first statement instead of restating the arithmetic twice. Comment-only: no non-comment line changes and the export list is unchanged. * fix(gateway): honor unlimited launchd stops Preserve the observed launcher cap only for launchd-driven stops, and shorten redundant budget documentation without changing the measured split. Co-authored-by: Patrick-Erichsen <20157849+Patrick-Erichsen@users.noreply.github.com> * test(gateway): use action-based drain budget Co-authored-by: Patrick-Erichsen <20157849+Patrick-Erichsen@users.noreply.github.com> * style(gateway): format launchd timeout assertions Co-authored-by: Patrick-Erichsen <20157849+Patrick-Erichsen@users.noreply.github.com> * fix(ci): complete frozen Docker planner import closures --------- Co-authored-by: roboclaw-bot <309084314+roboclaw-bot@users.noreply.github.com> Co-authored-by: Patrick-Erichsen <20157849+Patrick-Erichsen@users.noreply.github.com> Co-authored-by: Patrick Erichsen <patrick.a.erichsen@gmail.com>
This commit is contained in:
parent
d97d619dba
commit
6eeb120d69
17 changed files with 2027 additions and 144 deletions
|
|
@ -159,9 +159,14 @@ Node version and the service manager tracks a launcher parent. The launcher
|
|||
forwards the stop signal and waits for the serving Gateway to drain within the
|
||||
shared service budget. Managed restart intent targets the live serving owner,
|
||||
so unfinished work still follows restart recovery when its drain budget expires.
|
||||
The launchd stop budget remains 20 seconds; Linux units use the deadlines below.
|
||||
When launchd drives the stop, macOS uses the running job's own `ExitTimeOut`, capped
|
||||
by the launcher's own stop timer when a launcher is in the path, and Linux units use
|
||||
their own stop timeout; both are described in the deadline sections below.
|
||||
This requires a Gateway started with the updated launcher: replacing files cannot
|
||||
change a launcher that is already running.
|
||||
change a launcher that is already running. The stop deadlines below are the
|
||||
exception: the serving Gateway derives the launcher's reap timer rather than being
|
||||
told it, precisely so a Gateway started by an already-running older launcher still
|
||||
bounds itself correctly.
|
||||
|
||||
For these managed restarts, if the CLI cannot verify the service command, serving
|
||||
owner, or restart-intent recording, it refuses the restart before signaling with
|
||||
|
|
@ -214,11 +219,25 @@ unit's effective `TimeoutStopUSec`, including drop-ins. It logs the source and
|
|||
reconciled stop budget at both points, so a repaired unit takes effect without
|
||||
restarting first. Inspection and any wait for startup to finish consume the same
|
||||
shutdown deadline. Active-work drain uses at most
|
||||
315 seconds, with 10 seconds reserved for final chat writes and server cleanup
|
||||
and another 5 seconds before systemd's deadline. A unit with the default
|
||||
315 seconds, with up to 10 seconds reserved for final chat writes and server cleanup
|
||||
and up to another 5 seconds before systemd's deadline. A unit with the default
|
||||
90-second stop timeout therefore gets a 75-second drain and an 85-second Gateway
|
||||
shutdown deadline. A shorter supervisor timeout also caps requested restart waits.
|
||||
The drained work, ordering, and interruption behavior stay the same.
|
||||
|
||||
Both allowances are bounded when the deadline cannot fund them, by the same rule the
|
||||
launchd budget uses and for the same reason: subtracting two values sized for a
|
||||
315-second drain from a short stop timeout consumed it entirely and left active work
|
||||
nothing. A unit whose `TimeoutStopSec` is 20 seconds or longer keeps the allocation it
|
||||
already had, less whatever the stop inspection itself cost, rather than to the exact
|
||||
millisecond: 20 seconds is only the least that funds the full 10-second reserve
|
||||
alongside a 5-second drain before that cost comes off. Below that the deadline cannot fund
|
||||
both, and the reserve yields to keep a drain: a unit at or below 15 seconds now drains
|
||||
at all where it previously drained for zero milliseconds, and the four values between
|
||||
trade part of a reserve for a drain that was under 5 seconds. The reserve never falls
|
||||
below half the shutdown budget. The exit margin holds its full 5 seconds down to a
|
||||
20-second deadline and shrinks below that, to 3.75 seconds at 15 seconds and a quarter
|
||||
of anything shorter. The drained work, ordering, and interruption behavior stay the
|
||||
same, and systemd's own 90-second default is unaffected.
|
||||
|
||||
Service-child cleanup uses the remaining Gateway shutdown budget, leaving time
|
||||
for final exit bookkeeping. A forced restart drains admitted work within the same
|
||||
|
|
@ -309,6 +328,59 @@ before acknowledging, without overwriting a newer turn in that session.
|
|||
If another child cannot be stopped, the response still reports incomplete
|
||||
cancellation; the captured parent's cancellation is persisted before that error.
|
||||
|
||||
### Launchd stop deadlines
|
||||
|
||||
On a launchd-driven stop, the macOS Gateway reads the **loaded** job's effective
|
||||
`exit timeout` with `launchctl print`. It checks the system, GUI, and user
|
||||
domains, and accepts only a job whose PID is the Gateway or its launcher. Restart
|
||||
ownership can still be external: the supervisor that enforces the stop, not the
|
||||
owner of the next start, determines the deadline. A direct SIGTERM does not start
|
||||
launchd's stop clock, so it keeps the existing Gateway stop policy unless an
|
||||
OpenClaw launcher independently enforces its own child reap timer.
|
||||
|
||||
The **stop budget** is time available until the supervisor can kill the process.
|
||||
The Gateway sets aside an exit margin (up to 5 seconds) and plans its shutdown
|
||||
inside the remainder. **Drain** is the first part of that shutdown: stop admitting
|
||||
new work and wait for active turns and background tasks to settle. The rest is
|
||||
reserved for final writes and service cleanup (up to 10 seconds). The Gateway
|
||||
logs the chosen source, drain, shutdown deadline, reserve, and exit margin.
|
||||
|
||||
For the installed 20-second LaunchAgent job, the nominal split is 5 seconds of
|
||||
drain, 10 seconds for cleanup, and 5 seconds before launchd's deadline. Time
|
||||
spent inspecting the job is debited. For a shorter custom deadline the exit
|
||||
margin is at most a quarter, and the cleanup reserve gives way to leave active
|
||||
work up to 5 seconds of drain (at most half of the remaining shutdown budget
|
||||
when it is very short). For example, a 5-second job nominally leaves 1.875
|
||||
seconds each for drain and cleanup after its 1.25-second exit margin; a
|
||||
15-second job leaves 5 seconds of drain, 6.25 seconds for cleanup, and a
|
||||
3.75-second margin. **This reduces cleanup time for custom jobs below 20
|
||||
seconds.** The default systemd 90-second deadline is unaffected. The probe can
|
||||
consume up to three 2-second calls; a very slow inspection can leave no drain.
|
||||
|
||||
If a Node-recovery or compile-cache launcher is the job's PID, its own child
|
||||
reap timer may be shorter than launchd's deadline. The Gateway caps its budget
|
||||
at that timer only when its respawn markers establish that this launcher started
|
||||
it; an unrelated parent does not shorten the budget.
|
||||
|
||||
A direct signal to a marked launcher can start its own timer while launchd
|
||||
still reports `running`. The Gateway cannot distinguish that forwarded signal
|
||||
from a direct signal to its child, where the launcher has no timer, so it does
|
||||
not cap this case. Operators stopping a launcher directly should use the loaded
|
||||
job stop instead to get an enforceable, reported deadline.
|
||||
|
||||
`ExitTimeOut=0` means no launchd stop deadline, not a missing value. Without a
|
||||
launcher timer, the Gateway keeps its ordinary policy and does not classify the
|
||||
job as a native deadline. An unreadable job leaves
|
||||
the existing stop policy in place with a warning, while a confirmed stopping
|
||||
job whose timeout is missing uses launchd's 20-second default with a warning.
|
||||
|
||||
The value comes from the job launchd has **loaded**, not the plist on disk.
|
||||
Editing `ExitTimeOut` or running `launchctl kickstart -k` does not reload it;
|
||||
`launchctl bootout` followed by `launchctl bootstrap` does, restarting the
|
||||
Gateway. Reading per stop avoids a stale startup snapshot, not a stale loaded
|
||||
job. On macOS 27, launchd reports at most 60 seconds even if the plist asks
|
||||
for more; inspect the loaded job to confirm the effective value.
|
||||
|
||||
## Host sleep and process freezes
|
||||
|
||||
When a gateway host wakes from sleep, a virtual machine resumes, or the process
|
||||
|
|
|
|||
|
|
@ -4,3 +4,13 @@ export const GATEWAY_SHUTDOWN_TIMEOUT_MS: number;
|
|||
export const GATEWAY_SERVICE_STOP_TIMEOUT_MS: number;
|
||||
export const LAUNCH_AGENT_EXIT_TIMEOUT_SECONDS: 20;
|
||||
export const GATEWAY_RESTART_REPLACEMENT_TIMEOUT_MS: number;
|
||||
export const RESPAWN_SIGNAL_FORCE_KILL_GRACE_MS: number;
|
||||
export const RESPAWN_SIGNAL_HARD_EXIT_GRACE_MS: number;
|
||||
export function resolveSupervisorExitMarginMs(stopTimeoutMs: number): number;
|
||||
export function resolveShutdownReserveMs(shutdownTimeoutMs: number): number;
|
||||
export function isRespawnedByLauncher(env: NodeJS.ProcessEnv): boolean;
|
||||
export function resolveLauncherStopTimeoutMs(params: {
|
||||
env: NodeJS.ProcessEnv;
|
||||
platform: NodeJS.Platform;
|
||||
foreground: boolean;
|
||||
}): number;
|
||||
|
|
|
|||
|
|
@ -9,3 +9,72 @@ export const GATEWAY_SERVICE_STOP_TIMEOUT_MS =
|
|||
GATEWAY_SHUTDOWN_TIMEOUT_MS + GATEWAY_SUPERVISOR_EXIT_MARGIN_MS;
|
||||
|
||||
export const LAUNCH_AGENT_EXIT_TIMEOUT_SECONDS = 20;
|
||||
|
||||
// Keep a positive shutdown budget when a supervisor's deadline is under 20s.
|
||||
const GATEWAY_SUPERVISOR_EXIT_MARGIN_SHARE = 0.25;
|
||||
|
||||
// Preserve the template's 5s drain where possible. Under 20s, trade some cleanup
|
||||
// reserve for drain; on very short jobs, split the remaining budget in half.
|
||||
// Inspection time is debited before this allocation, so a slow probe can also
|
||||
// reduce the drain. The measured upgrade tradeoff is documented in restart-recovery.
|
||||
const GATEWAY_SHUTDOWN_DRAIN_FLOOR_MS = 5_000;
|
||||
const GATEWAY_SHUTDOWN_DRAIN_FLOOR_SHARE = 0.5;
|
||||
|
||||
/** The exit margin to hold back from a supervisor-enforced stop deadline. */
|
||||
export const resolveSupervisorExitMarginMs = (stopTimeoutMs) =>
|
||||
Math.min(
|
||||
GATEWAY_SUPERVISOR_EXIT_MARGIN_MS,
|
||||
Math.floor(Math.max(0, stopTimeoutMs) * GATEWAY_SUPERVISOR_EXIT_MARGIN_SHARE),
|
||||
);
|
||||
|
||||
/** The post-drain reserve to hold back from a resolved shutdown budget. */
|
||||
export const resolveShutdownReserveMs = (shutdownTimeoutMs) => {
|
||||
const budgetMs = Math.max(0, shutdownTimeoutMs);
|
||||
const drainFloorMs = Math.min(
|
||||
GATEWAY_SHUTDOWN_DRAIN_FLOOR_MS,
|
||||
Math.floor(budgetMs * GATEWAY_SHUTDOWN_DRAIN_FLOOR_SHARE),
|
||||
);
|
||||
return Math.min(GATEWAY_SHUTDOWN_RESERVE_MS, budgetMs - drainFloorMs);
|
||||
};
|
||||
|
||||
// Escalation graces the Node recovery launcher applies to a stopping child. Kept
|
||||
// here rather than in the launcher so the serving Gateway can derive the deadline
|
||||
// its parent enforces from the same numbers the parent armed it from.
|
||||
const RESPAWN_SIGNAL_EXIT_GRACE_MS = 1_000;
|
||||
export const RESPAWN_SIGNAL_FORCE_KILL_GRACE_MS = 1_000;
|
||||
export const RESPAWN_SIGNAL_HARD_EXIT_GRACE_MS = 1_000;
|
||||
|
||||
// A launcher inside OpenClaw's LaunchAgent gets its template deadline; otherwise
|
||||
// it uses the generic service stop policy. Its child inherits these markers.
|
||||
const resolveRespawnServiceStopTimeoutMs = (env, platform) => {
|
||||
const launchdService = env.OPENCLAW_LAUNCHD_LABEL?.trim();
|
||||
return platform === "darwin" && launchdService && env.XPC_SERVICE_NAME === launchdService
|
||||
? LAUNCH_AGENT_EXIT_TIMEOUT_SECONDS * 1_000
|
||||
: GATEWAY_SERVICE_STOP_TIMEOUT_MS;
|
||||
};
|
||||
|
||||
// Every runRespawnedChild call site stamps one of these. The compile-cache
|
||||
// marker is also used by a different respawner, but that one refuses foreground
|
||||
// Gateway runs on darwin. Keep this list with the launcher deadline it authorizes.
|
||||
const RESPAWN_LAUNCHER_MARKER_ENV_VARS = [
|
||||
"OPENCLAW_NODE_UPDATE_RESPAWNED",
|
||||
"OPENCLAW_COMPILE_CACHE_DISABLED_RESPAWNED",
|
||||
"OPENCLAW_PACKAGED_COMPILE_CACHE_RESPAWNED",
|
||||
];
|
||||
|
||||
/** Whether a `runRespawnedChild` parent started this process. */
|
||||
export const isRespawnedByLauncher = (env) =>
|
||||
RESPAWN_LAUNCHER_MARKER_ENV_VARS.some((name) => env[name] === "1");
|
||||
|
||||
// Shared with the serving Gateway: derive the parent's reap deadline from the
|
||||
// same graces, including on the first update while the old launcher is still live.
|
||||
export const resolveLauncherStopTimeoutMs = ({ env, platform, foreground }) => {
|
||||
const serviceStopTimeoutMs = resolveRespawnServiceStopTimeoutMs(env, platform);
|
||||
const signalExitGraceMs =
|
||||
platform !== "win32" && foreground
|
||||
? serviceStopTimeoutMs -
|
||||
RESPAWN_SIGNAL_FORCE_KILL_GRACE_MS -
|
||||
RESPAWN_SIGNAL_HARD_EXIT_GRACE_MS
|
||||
: RESPAWN_SIGNAL_EXIT_GRACE_MS;
|
||||
return signalExitGraceMs + RESPAWN_SIGNAL_FORCE_KILL_GRACE_MS;
|
||||
};
|
||||
|
|
|
|||
|
|
@ -6,8 +6,9 @@ import path from "node:path";
|
|||
import { consumeRootOptionToken as consumeLauncherRootOptionToken } from "./cli-root-options.mjs";
|
||||
import { isForegroundGatewayRunArgv } from "./gateway-run-argv.mjs";
|
||||
import {
|
||||
GATEWAY_SERVICE_STOP_TIMEOUT_MS,
|
||||
LAUNCH_AGENT_EXIT_TIMEOUT_SECONDS,
|
||||
RESPAWN_SIGNAL_FORCE_KILL_GRACE_MS,
|
||||
RESPAWN_SIGNAL_HARD_EXIT_GRACE_MS,
|
||||
resolveLauncherStopTimeoutMs,
|
||||
} from "./gateway-shutdown-budget.mjs";
|
||||
import {
|
||||
detectCurrentSqliteCapabilities,
|
||||
|
|
@ -39,22 +40,21 @@ const respawnSignals =
|
|||
process.platform === "win32"
|
||||
? ["SIGTERM", "SIGINT", "SIGBREAK"]
|
||||
: ["SIGTERM", "SIGINT", "SIGHUP", "SIGQUIT"];
|
||||
const respawnSignalExitGraceMs = 1_000;
|
||||
const respawnSignalForceKillGraceMs = 1_000;
|
||||
const respawnSignalHardExitGraceMs = 1_000;
|
||||
const respawnSignalForceKillGraceMs = RESPAWN_SIGNAL_FORCE_KILL_GRACE_MS;
|
||||
const respawnSignalHardExitGraceMs = RESPAWN_SIGNAL_HARD_EXIT_GRACE_MS;
|
||||
|
||||
export const runRespawnedChild = (command, args, env) => {
|
||||
const launchdService = env.OPENCLAW_LAUNCHD_LABEL?.trim();
|
||||
const serviceStopTimeoutMs =
|
||||
process.platform === "darwin" && launchdService && env.XPC_SERVICE_NAME === launchdService
|
||||
? LAUNCH_AGENT_EXIT_TIMEOUT_SECONDS * 1_000
|
||||
: GATEWAY_SERVICE_STOP_TIMEOUT_MS;
|
||||
// The serving Gateway owns drain and cleanup. Reap a stuck child only in the
|
||||
// supervisor's exit margin, after that owner has had its full shutdown budget.
|
||||
const signalExitGraceMs =
|
||||
process.platform !== "win32" && isForegroundGatewayRunArgv(process.argv)
|
||||
? serviceStopTimeoutMs - respawnSignalForceKillGraceMs - respawnSignalHardExitGraceMs
|
||||
: respawnSignalExitGraceMs;
|
||||
// The shared resolver owns this arithmetic so the serving Gateway derives the very
|
||||
// same deadline from the same expression, which is what lets a Gateway started by
|
||||
// any build of this launcher bound itself correctly without being told.
|
||||
const launcherStopTimeoutMs = resolveLauncherStopTimeoutMs({
|
||||
env,
|
||||
platform: process.platform,
|
||||
foreground: isForegroundGatewayRunArgv(process.argv),
|
||||
});
|
||||
const signalExitGraceMs = launcherStopTimeoutMs - respawnSignalForceKillGraceMs;
|
||||
const stdioIsTerminal = process.stdin.isTTY || process.stdout.isTTY;
|
||||
const child = spawn(command, args, {
|
||||
stdio: "inherit",
|
||||
|
|
@ -62,7 +62,10 @@ export const runRespawnedChild = (command, args, env) => {
|
|||
windowsHide: !stdioIsTerminal,
|
||||
});
|
||||
const listeners = new Map();
|
||||
// Keep signal forwarding and bounded shutdown in sync with src/entry.compile-cache.ts.
|
||||
// Keep signal forwarding and bounded shutdown in sync with src/entry.compile-cache.ts,
|
||||
// which drives src/process/respawn-child-runner.ts. That runner still holds its own
|
||||
// copies of the escalation graces and reaps on a fixed short one, so only this
|
||||
// launcher's deadline is the one the serving Gateway derives.
|
||||
let signalExitTimer = null;
|
||||
let signalForceKillTimer = null;
|
||||
let signalHardExitTimer = null;
|
||||
|
|
|
|||
|
|
@ -1180,11 +1180,6 @@ async function preflightFrozenTargetContracts(input, workflow = false, verifiedT
|
|||
for (const path of supportFiles[consumer] ?? []) {
|
||||
required(sources.tooling, `scripts/e2e/lib/${path}`);
|
||||
}
|
||||
if (
|
||||
["npm-onboard-channel-agent", "codex-on-demand", "update-corrupt-plugin"].includes(consumer)
|
||||
) {
|
||||
required(sources.tooling, "scripts/lib/record-shared.mjs");
|
||||
}
|
||||
if (consumer === "update-corrupt-plugin") {
|
||||
required(sources.tooling, "scripts/lib/update-compat-contract.mjs");
|
||||
required(sources.tooling, "scripts/lib/openclaw-e2e-instance.sh");
|
||||
|
|
|
|||
|
|
@ -1,16 +1,29 @@
|
|||
import { performance } from "node:perf_hooks";
|
||||
import { afterEach, beforeEach, describe, expect, it, vi } from "vitest";
|
||||
import { resolveGatewayShutdownBudget } from "./run-loop-shutdown-budget.js";
|
||||
import {
|
||||
GATEWAY_SHUTDOWN_RESERVE_MS,
|
||||
GATEWAY_SUPERVISOR_EXIT_MARGIN_MS,
|
||||
} from "../../infra/gateway-shutdown-budget.js";
|
||||
import {
|
||||
resolveGatewayShutdownBudget,
|
||||
resolveGatewayShutdownDrainBudget,
|
||||
} from "./run-loop-shutdown-budget.js";
|
||||
|
||||
const { readFile, execUser, execSystem } = vi.hoisted(() => ({
|
||||
const { readFile, execUser, execSystem, execLaunchctl } = vi.hoisted(() => ({
|
||||
readFile: vi.fn(),
|
||||
execUser: vi.fn(),
|
||||
execSystem: vi.fn(),
|
||||
execLaunchctl: vi.fn(),
|
||||
}));
|
||||
vi.mock("node:fs/promises", () => ({ default: { readFile } }));
|
||||
vi.mock("../../daemon/systemd-exec.js", () => ({
|
||||
execSystemctlUser: execUser,
|
||||
execSystemctl: execSystem,
|
||||
}));
|
||||
vi.mock("../../daemon/launchd-exec.js", async (importOriginal) => ({
|
||||
...(await importOriginal<typeof import("../../daemon/launchd-exec.js")>()),
|
||||
execLaunchctl,
|
||||
}));
|
||||
|
||||
beforeEach(() => {
|
||||
vi.stubGlobal("process", { ...process, platform: "linux", getuid: () => 1000, env: {} });
|
||||
|
|
@ -23,6 +36,7 @@ beforeEach(() => {
|
|||
"User=openclaw\nType=simple\nNotifyAccess=none\nKillMode=control-group",
|
||||
stderr: "",
|
||||
});
|
||||
execLaunchctl.mockReset();
|
||||
});
|
||||
afterEach(() => vi.unstubAllGlobals());
|
||||
|
||||
|
|
@ -93,4 +107,472 @@ describe("Gateway stop deadline independent of restart ownership", () => {
|
|||
expect(execSystem).not.toHaveBeenCalled();
|
||||
expect(execUser).not.toHaveBeenCalled();
|
||||
});
|
||||
|
||||
// The systemd coverage above only pins TimeoutStopUSec at 90 seconds and 10 minutes,
|
||||
// both comfortably above the 20 second threshold. This mirrors the darwin exit-timeout
|
||||
// table's sub-20-second cases against a systemd stub: a unit deadline in this band
|
||||
// gives up part of a reserve a flat subtraction would have funded in full, rather than
|
||||
// draining for zero milliseconds as that flat subtraction did.
|
||||
it.each([
|
||||
{ seconds: 19, timeoutMs: 14_250, reserveMs: 9_250, drainMs: 5_000, fixedDrainMs: 4_000 },
|
||||
{ seconds: 16, timeoutMs: 12_000, reserveMs: 7_000, drainMs: 5_000, fixedDrainMs: 1_000 },
|
||||
{ seconds: 15, timeoutMs: 11_250, reserveMs: 6_250, drainMs: 5_000, fixedDrainMs: 0 },
|
||||
{ seconds: 10, timeoutMs: 7_500, reserveMs: 3_750, drainMs: 3_750, fixedDrainMs: 0 },
|
||||
])(
|
||||
"keeps a drain a $seconds second systemd TimeoutStopUSec previously spent on overhead",
|
||||
async ({ seconds, timeoutMs, reserveMs, drainMs, fixedDrainMs }) => {
|
||||
process.env.OPENCLAW_SUPERVISOR_MODE = "external";
|
||||
execSystem.mockResolvedValue({
|
||||
code: 0,
|
||||
stdout: `LoadState=loaded\nTimeoutStopUSec=${seconds}s\nInvocationID=own`,
|
||||
stderr: "",
|
||||
});
|
||||
const budget = await resolveGatewayShutdownBudget("external", {
|
||||
info: vi.fn(),
|
||||
warn: vi.fn(),
|
||||
});
|
||||
expect(budget.timeoutMs).toBe(timeoutMs);
|
||||
expect(budget.reserveMs).toBe(reserveMs);
|
||||
expect(budget.timeoutMs - budget.reserveMs).toBe(drainMs);
|
||||
// What a flat, unshared subtraction would have left active work at the same deadline.
|
||||
expect(
|
||||
Math.max(
|
||||
0,
|
||||
seconds * 1_000 - GATEWAY_SUPERVISOR_EXIT_MARGIN_MS - GATEWAY_SHUTDOWN_RESERVE_MS,
|
||||
),
|
||||
).toBe(fixedDrainMs);
|
||||
expect(budget.reserveMs).toBeGreaterThanOrEqual(Math.floor(timeoutMs / 2));
|
||||
},
|
||||
);
|
||||
});
|
||||
|
||||
describe("Gateway stop deadline follows the launchd stop that is actually running", () => {
|
||||
// A stop that is under way. `previous` is the startup budget a darwin Gateway
|
||||
// resolves before any stop exists, which is the platform-neutral policy.
|
||||
// The budget subtracts `performance.now() - acceptedAtMs` and floors it at
|
||||
// zero, so an acceptance stamped ahead of the clock records exactly no elapsed
|
||||
// time and keeps the asserted numbers exact instead of off by a stray
|
||||
// millisecond.
|
||||
const stoppingNow = {
|
||||
previous: { timeoutMs: 325_000, nativeStopBudget: false },
|
||||
acceptedAtMs: Number.MAX_SAFE_INTEGER,
|
||||
};
|
||||
const printed = (state: string, fields: string) => ({
|
||||
code: 0,
|
||||
stdout: `system/ai.openclaw.gateway = {\n\tstate = ${state}\n\n${fields}\tresource coalition = {\n\t\tstate = active\n\t}\n}\n`,
|
||||
stderr: "",
|
||||
termination: "exit",
|
||||
});
|
||||
|
||||
beforeEach(() => {
|
||||
vi.stubGlobal("process", {
|
||||
...process,
|
||||
platform: "darwin",
|
||||
pid: 4242,
|
||||
getuid: () => 501,
|
||||
env: {},
|
||||
});
|
||||
process.env.XPC_SERVICE_NAME = "ai.openclaw.gateway";
|
||||
process.env.OPENCLAW_SUPERVISOR_MODE = "external";
|
||||
});
|
||||
|
||||
// THE REGRESSION GUARD. The linked report is an externally delivered SIGTERM
|
||||
// under a five second job, where launchd never starts its clock and the drain ran
|
||||
// its full 315 seconds. Measured on the reporting host, it then hit its own
|
||||
// timeout with work still active rather than finishing early, so the drain was
|
||||
// being used. Adopting the job deadline there would hand that same supported
|
||||
// setup a zero drain and cut that work off at once.
|
||||
it("keeps the full drain when launchd did not initiate the stop", async () => {
|
||||
execLaunchctl.mockResolvedValue(printed("running", "\texit timeout = 5\n\tpid = 4242\n"));
|
||||
const info = vi.fn();
|
||||
const budget = await resolveGatewayShutdownBudget(
|
||||
"external",
|
||||
{ info, warn: vi.fn() },
|
||||
stoppingNow,
|
||||
);
|
||||
budget.log("shutdown");
|
||||
expect(info).toHaveBeenCalledWith(
|
||||
"shutdown budget at shutdown: drain=315000ms shutdown=325000ms reserve=10000ms exitMargin=5000ms; source=Gateway stop policy=330000ms",
|
||||
);
|
||||
expect(budget.timeoutMs).toBe(325_000);
|
||||
expect(budget.nativeStopBudget).toBe(false);
|
||||
});
|
||||
|
||||
it.each(["external", "launchd"])(
|
||||
"does not treat a launchd ExitTimeOut of zero as a native deadline under %s ownership",
|
||||
async (supervisor) => {
|
||||
execLaunchctl.mockResolvedValue(printed("SIGTERMed", "\texit timeout = 0\n\tpid = 4242\n"));
|
||||
const budget = await resolveGatewayShutdownBudget(
|
||||
supervisor,
|
||||
{ info: vi.fn(), warn: vi.fn() },
|
||||
stoppingNow,
|
||||
);
|
||||
expect(budget.timeoutMs).toBe(325_000);
|
||||
expect(budget.nativeStopBudget).toBe(false);
|
||||
const drain = resolveGatewayShutdownDrainBudget({
|
||||
budget,
|
||||
action: "restart",
|
||||
forceRestart: false,
|
||||
restartWithoutSupervisor: false,
|
||||
acceptedAtMs: performance.now(),
|
||||
requestedRestartDrainTimeoutMs: 600_000,
|
||||
});
|
||||
expect(drain.drainTimeoutMs).toBeGreaterThan(590_000);
|
||||
},
|
||||
);
|
||||
|
||||
// A deadline this short cannot fund the fixed 5s margin and 10s reserve, and
|
||||
// subtracting them outright left the job's whole 5 seconds spent on overhead with
|
||||
// nothing to drain. Each allowance is capped at a share of what it is carved from,
|
||||
// so the short job keeps a proportional drain that still fits inside the deadline.
|
||||
it("keeps a proportional drain when the job's exit timeout cannot fund the fixed allowances", async () => {
|
||||
execLaunchctl.mockResolvedValue(
|
||||
printed("SIGTERMed", "\tminimum runtime = 10\n\texit timeout = 5\n\tpid = 4242\n"),
|
||||
);
|
||||
const info = vi.fn();
|
||||
const budget = await resolveGatewayShutdownBudget(
|
||||
"external",
|
||||
{ info, warn: vi.fn() },
|
||||
stoppingNow,
|
||||
);
|
||||
budget.log("shutdown");
|
||||
expect(info).toHaveBeenCalledWith(
|
||||
"shutdown budget at shutdown: drain=1875ms shutdown=3750ms reserve=1875ms exitMargin=1250ms; source=launchd system/ai.openclaw.gateway exit timeout=5000ms",
|
||||
);
|
||||
expect(budget.timeoutMs).toBe(3_750);
|
||||
expect(budget.reserveMs).toBe(1_875);
|
||||
expect(budget.nativeStopBudget).toBe(true);
|
||||
});
|
||||
|
||||
// Drain must never be starved to zero by the allowances: every positive deadline
|
||||
// leaves active work some time, and a longer deadline never yields less of it.
|
||||
it.each([1, 2, 5, 10, 15, 20, 25, 47, 60])(
|
||||
"leaves a positive drain for a %s second exit timeout",
|
||||
async (seconds) => {
|
||||
execLaunchctl.mockResolvedValue(
|
||||
printed("SIGTERMed", `\texit timeout = ${seconds}\n\tpid = 4242\n`),
|
||||
);
|
||||
const budget = await resolveGatewayShutdownBudget(
|
||||
"external",
|
||||
{ info: vi.fn(), warn: vi.fn() },
|
||||
stoppingNow,
|
||||
);
|
||||
expect(budget.timeoutMs - budget.reserveMs).toBeGreaterThan(0);
|
||||
expect(budget.timeoutMs).toBeLessThanOrEqual(seconds * 1_000);
|
||||
},
|
||||
);
|
||||
|
||||
it.each([
|
||||
{ seconds: 20, timeoutMs: 15_000, reserveMs: 10_000, drainMs: 5_000, exitMarginMs: 5_000 },
|
||||
{ seconds: 47, timeoutMs: 42_000, reserveMs: 10_000, drainMs: 32_000, exitMarginMs: 5_000 },
|
||||
{ seconds: 55, timeoutMs: 50_000, reserveMs: 10_000, drainMs: 40_000, exitMarginMs: 5_000 },
|
||||
])(
|
||||
"derives the budget from a $seconds second exit timeout",
|
||||
async ({ seconds, timeoutMs, reserveMs, drainMs, exitMarginMs }) => {
|
||||
execLaunchctl.mockResolvedValue(
|
||||
printed("SIGTERMed", `\texit timeout = ${seconds}\n\tpid = 4242\n`),
|
||||
);
|
||||
const info = vi.fn();
|
||||
const budget = await resolveGatewayShutdownBudget(
|
||||
"external",
|
||||
{ info, warn: vi.fn() },
|
||||
stoppingNow,
|
||||
);
|
||||
budget.log("shutdown");
|
||||
expect(budget.timeoutMs).toBe(timeoutMs);
|
||||
expect(budget.reserveMs).toBe(reserveMs);
|
||||
expect(info).toHaveBeenCalledWith(
|
||||
`shutdown budget at shutdown: drain=${drainMs}ms shutdown=${timeoutMs}ms reserve=${reserveMs}ms exitMargin=${exitMarginMs}ms; source=launchd system/ai.openclaw.gateway exit timeout=${seconds * 1_000}ms`,
|
||||
);
|
||||
},
|
||||
);
|
||||
|
||||
// Every case above pins `acceptedAtMs` to `Number.MAX_SAFE_INTEGER`, which floors
|
||||
// elapsed at zero and proves nothing about a real, nonzero debit. `performance.now`
|
||||
// is not faked anywhere in this file, so this pins it directly rather than trusting
|
||||
// the wall clock to land on a specific millisecond by chance: the reserve gives up
|
||||
// exactly that observed 13ms debit and the 5 second drain floor still funds in full.
|
||||
it("debits the reserve by a real elapsed delay, leaving the drain floor untouched", async () => {
|
||||
execLaunchctl.mockResolvedValue(printed("SIGTERMed", "\texit timeout = 20\n\tpid = 4242\n"));
|
||||
const nowMs = performance.now();
|
||||
const clock = vi.spyOn(performance, "now").mockReturnValue(nowMs);
|
||||
try {
|
||||
const budget = await resolveGatewayShutdownBudget(
|
||||
"external",
|
||||
{ info: vi.fn(), warn: vi.fn() },
|
||||
{ previous: stoppingNow.previous, acceptedAtMs: nowMs - 13 },
|
||||
);
|
||||
expect(budget.timeoutMs).toBe(14_987);
|
||||
expect(budget.reserveMs).toBe(9_987);
|
||||
expect(budget.timeoutMs - budget.reserveMs).toBe(5_000);
|
||||
} finally {
|
||||
clock.mockRestore();
|
||||
}
|
||||
});
|
||||
|
||||
// The case above only covers a debit small enough that the reserve pays it alone. The
|
||||
// probe can cost far more than 13ms: it is up to three `launchctl print` calls at a
|
||||
// 2 second timeout each. Past a 5 second debit the drain floor is share-bounded too,
|
||||
// so the claim that a 20 second deadline keeps its old allocation less the debit stops
|
||||
// holding and both allowances converge on half the remainder. Pinning that boundary
|
||||
// keeps the documented threshold honest instead of reasoned.
|
||||
it("converges the reserve and the drain once the elapsed debit passes the floor", async () => {
|
||||
execLaunchctl.mockResolvedValue(printed("SIGTERMed", "\texit timeout = 20\n\tpid = 4242\n"));
|
||||
const nowMs = performance.now();
|
||||
const clock = vi.spyOn(performance, "now").mockReturnValue(nowMs);
|
||||
try {
|
||||
const budget = await resolveGatewayShutdownBudget(
|
||||
"external",
|
||||
{ info: vi.fn(), warn: vi.fn() },
|
||||
{ previous: stoppingNow.previous, acceptedAtMs: nowMs - 6_000 },
|
||||
);
|
||||
expect(budget.timeoutMs).toBe(9_000);
|
||||
expect(budget.reserveMs).toBe(4_500);
|
||||
expect(budget.timeoutMs - budget.reserveMs).toBe(4_500);
|
||||
// Not the old allocation less the debit: that would have left the reserve at 4000.
|
||||
expect(budget.reserveMs).not.toBe(GATEWAY_SHUTDOWN_RESERVE_MS - 6_000);
|
||||
} finally {
|
||||
clock.mockRestore();
|
||||
}
|
||||
});
|
||||
|
||||
// The allowances are only ever capped to keep a drain, never to reallocate a deadline
|
||||
// that already worked. A deadline able to fund the 10s reserve alongside the 5s drain
|
||||
// the 20s template yields needs 15s of shutdown budget, which every deadline from 20s
|
||||
// up has, so all of them must resolve exactly what subtracting the fixed allowances
|
||||
// outright resolved. Asserting against that arithmetic rather than against literals is
|
||||
// what makes this a regression test: capping the reserve at a share of the budget, as
|
||||
// an earlier revision did, drops a 20s job's reserve to 7500ms and fails here.
|
||||
it.each([20, 21, 25, 30, 47, 55, 60])(
|
||||
"allocates a %s second exit timeout exactly as the fixed allowances did",
|
||||
async (seconds) => {
|
||||
execLaunchctl.mockResolvedValue(
|
||||
printed("SIGTERMed", `\texit timeout = ${seconds}\n\tpid = 4242\n`),
|
||||
);
|
||||
const budget = await resolveGatewayShutdownBudget(
|
||||
"external",
|
||||
{ info: vi.fn(), warn: vi.fn() },
|
||||
stoppingNow,
|
||||
);
|
||||
const fixedTimeoutMs = seconds * 1_000 - GATEWAY_SUPERVISOR_EXIT_MARGIN_MS;
|
||||
expect(budget.timeoutMs).toBe(fixedTimeoutMs);
|
||||
expect(budget.reserveMs).toBe(GATEWAY_SHUTDOWN_RESERVE_MS);
|
||||
expect(budget.timeoutMs - budget.reserveMs).toBe(
|
||||
fixedTimeoutMs - GATEWAY_SHUTDOWN_RESERVE_MS,
|
||||
);
|
||||
},
|
||||
);
|
||||
|
||||
// Under 20 seconds the budget cannot fund both allowances, so one has to give. The
|
||||
// fixed subtraction gave up the drain: 15 seconds and below drained for 0ms, and the
|
||||
// four deadlines between left active work under 5 seconds. These pin what is given up
|
||||
// instead, and that the reserve never falls below half the budget doing it. No shipped
|
||||
// template or platform default lands here: the LaunchAgent template and launchd's own
|
||||
// default are both 20 seconds, and systemd's default stop timeout is 90.
|
||||
it.each([
|
||||
{ seconds: 19, timeoutMs: 14_250, reserveMs: 9_250, drainMs: 5_000, fixedDrainMs: 4_000 },
|
||||
{ seconds: 16, timeoutMs: 12_000, reserveMs: 7_000, drainMs: 5_000, fixedDrainMs: 1_000 },
|
||||
{ seconds: 15, timeoutMs: 11_250, reserveMs: 6_250, drainMs: 5_000, fixedDrainMs: 0 },
|
||||
{ seconds: 10, timeoutMs: 7_500, reserveMs: 3_750, drainMs: 3_750, fixedDrainMs: 0 },
|
||||
])(
|
||||
"keeps a drain a $seconds second exit timeout previously spent on overhead",
|
||||
async ({ seconds, timeoutMs, reserveMs, drainMs, fixedDrainMs }) => {
|
||||
execLaunchctl.mockResolvedValue(
|
||||
printed("SIGTERMed", `\texit timeout = ${seconds}\n\tpid = 4242\n`),
|
||||
);
|
||||
const budget = await resolveGatewayShutdownBudget(
|
||||
"external",
|
||||
{ info: vi.fn(), warn: vi.fn() },
|
||||
stoppingNow,
|
||||
);
|
||||
expect(budget.timeoutMs).toBe(timeoutMs);
|
||||
expect(budget.reserveMs).toBe(reserveMs);
|
||||
expect(budget.timeoutMs - budget.reserveMs).toBe(drainMs);
|
||||
// What the fixed subtraction left active work at the same deadline.
|
||||
expect(
|
||||
Math.max(
|
||||
0,
|
||||
seconds * 1_000 - GATEWAY_SUPERVISOR_EXIT_MARGIN_MS - GATEWAY_SHUTDOWN_RESERVE_MS,
|
||||
),
|
||||
).toBe(fixedDrainMs);
|
||||
expect(budget.reserveMs).toBeGreaterThanOrEqual(Math.floor(timeoutMs / 2));
|
||||
},
|
||||
);
|
||||
|
||||
// Failing to inspect the job establishes nothing, so shortening the drain here
|
||||
// would cut work that no launchd deadline was bounding.
|
||||
it("warns and keeps the platform-neutral policy when the job cannot be inspected", async () => {
|
||||
execLaunchctl.mockResolvedValue({
|
||||
code: 1,
|
||||
stdout: "",
|
||||
stderr: "permission denied",
|
||||
termination: "exit",
|
||||
});
|
||||
const warn = vi.fn();
|
||||
const budget = await resolveGatewayShutdownBudget(
|
||||
"external",
|
||||
{ info: vi.fn(), warn },
|
||||
stoppingNow,
|
||||
);
|
||||
expect(budget.timeoutMs).toBe(325_000);
|
||||
// The number alone is not the contract. A failed probe confirmed no launchd
|
||||
// deadline, so this must not be classified as a native stop budget either:
|
||||
// that flag is what caps a restart drain and arms a forced exit.
|
||||
expect(budget.nativeStopBudget).toBe(false);
|
||||
expect(warn).toHaveBeenCalledExactlyOnceWith(
|
||||
expect.stringContaining("Unable to inspect the launchd job"),
|
||||
);
|
||||
});
|
||||
|
||||
// The flag is only worth asserting because of what it does downstream, so drive
|
||||
// the real consumer. An operator restart that asked to drain for ten minutes
|
||||
// keeps that request when no launchd deadline was ever confirmed.
|
||||
it("leaves a longer requested restart drain uncapped when the job cannot be inspected", async () => {
|
||||
execLaunchctl.mockResolvedValue({
|
||||
code: 1,
|
||||
stdout: "",
|
||||
stderr: "permission denied",
|
||||
termination: "exit",
|
||||
});
|
||||
const budget = await resolveGatewayShutdownBudget(
|
||||
"external",
|
||||
{ info: vi.fn(), warn: vi.fn() },
|
||||
stoppingNow,
|
||||
);
|
||||
expect(budget.nativeStopBudget).toBe(false);
|
||||
const drain = resolveGatewayShutdownDrainBudget({
|
||||
budget,
|
||||
action: "restart",
|
||||
forceRestart: false,
|
||||
restartWithoutSupervisor: false,
|
||||
acceptedAtMs: performance.now(),
|
||||
requestedRestartDrainTimeoutMs: 600_000,
|
||||
});
|
||||
// Only elapsed time comes off the request; no supervisor ceiling applies.
|
||||
expect(drain.drainTimeoutMs).toBeGreaterThan(590_000);
|
||||
expect(drain.restartTimeoutMs()).toBe(325_000);
|
||||
});
|
||||
|
||||
// The same consumer, with a deadline that WAS confirmed, still gets capped.
|
||||
// Without this pair the test above would also pass if the launchd read were
|
||||
// deleted outright.
|
||||
it("caps that same restart drain at a confirmed job deadline", async () => {
|
||||
execLaunchctl.mockResolvedValue(printed("SIGTERMed", "\texit timeout = 47\n\tpid = 4242\n"));
|
||||
const budget = await resolveGatewayShutdownBudget(
|
||||
"external",
|
||||
{ info: vi.fn(), warn: vi.fn() },
|
||||
stoppingNow,
|
||||
);
|
||||
expect(budget.nativeStopBudget).toBe(true);
|
||||
const drain = resolveGatewayShutdownDrainBudget({
|
||||
budget,
|
||||
action: "restart",
|
||||
forceRestart: false,
|
||||
restartWithoutSupervisor: false,
|
||||
acceptedAtMs: performance.now(),
|
||||
requestedRestartDrainTimeoutMs: 600_000,
|
||||
});
|
||||
// 47s job - 5s exit margin = 42000ms shutdown, less the 10000ms reserve.
|
||||
expect(drain.drainTimeoutMs).toBe(32_000);
|
||||
});
|
||||
|
||||
// A launchd-OWNED Gateway, rather than an externally supervised one. Its startup
|
||||
// budget is already native, so this is the configuration where the retained-budget
|
||||
// safety net can fire. `previous` is what a 20 second template job resolves.
|
||||
const launchdOwnedStop = {
|
||||
previous: { timeoutMs: 15_000, nativeStopBudget: true },
|
||||
acceptedAtMs: Number.MAX_SAFE_INTEGER,
|
||||
};
|
||||
|
||||
// An in-process restart signals the Gateway without launchd running the stop, so
|
||||
// the job still prints `running`. That is a confirmed answer, not a failed probe,
|
||||
// and claiming the deadline "could not be confirmed" there would be false on every
|
||||
// in-process restart of a default macOS install.
|
||||
it("does not claim an unconfirmed timeout when launchd is confirmed not to be stopping", async () => {
|
||||
delete process.env.OPENCLAW_SUPERVISOR_MODE;
|
||||
execLaunchctl.mockResolvedValue(printed("running", "\texit timeout = 20\n\tpid = 4242\n"));
|
||||
const warn = vi.fn();
|
||||
const budget = await resolveGatewayShutdownBudget(
|
||||
"launchd",
|
||||
{ info: vi.fn(), warn },
|
||||
launchdOwnedStop,
|
||||
);
|
||||
expect(warn).not.toHaveBeenCalled();
|
||||
expect(budget.timeoutMs).toBe(15_000);
|
||||
expect(budget.nativeStopBudget).toBe(true);
|
||||
});
|
||||
|
||||
// A probe that established nothing is the case the safety net exists for, so the
|
||||
// startup budget is held rather than widened.
|
||||
it("retains the startup budget when the job could not be inspected at all", async () => {
|
||||
delete process.env.OPENCLAW_SUPERVISOR_MODE;
|
||||
execLaunchctl.mockResolvedValue({
|
||||
code: 1,
|
||||
stdout: "",
|
||||
stderr: "permission denied",
|
||||
termination: "exit",
|
||||
});
|
||||
const warn = vi.fn();
|
||||
const budget = await resolveGatewayShutdownBudget(
|
||||
"launchd",
|
||||
{ info: vi.fn(), warn },
|
||||
launchdOwnedStop,
|
||||
);
|
||||
expect(warn).toHaveBeenCalledWith(expect.stringContaining("Unable to inspect the launchd job"));
|
||||
expect(warn).toHaveBeenCalledWith(
|
||||
"Retaining the startup shutdown budget of 15000ms because the current supervisor stop timeout could not be confirmed.",
|
||||
);
|
||||
expect(budget.timeoutMs).toBe(15_000);
|
||||
});
|
||||
|
||||
// Warned but non-null is the defaulted-value case, not a failed probe: launchd is
|
||||
// confirmed to be stopping the job, so a clock is running and there is nothing to
|
||||
// retain. Warning about a timeout that "could not be confirmed" here would be the
|
||||
// same false statement in a different place.
|
||||
it("does not retain when only the deadline's value had to be defaulted", async () => {
|
||||
delete process.env.OPENCLAW_SUPERVISOR_MODE;
|
||||
execLaunchctl.mockResolvedValue(printed("SIGTERMed", "\tpid = 4242\n"));
|
||||
const warn = vi.fn();
|
||||
const budget = await resolveGatewayShutdownBudget(
|
||||
"launchd",
|
||||
{ info: vi.fn(), warn },
|
||||
launchdOwnedStop,
|
||||
);
|
||||
expect(warn).toHaveBeenCalledExactlyOnceWith(
|
||||
expect.stringContaining("its exit timeout is missing or invalid"),
|
||||
);
|
||||
expect(budget.timeoutMs).toBe(15_000);
|
||||
});
|
||||
|
||||
// No stop is running at startup, so there is no enforcing deadline to read and
|
||||
// no reason to spend a launchctl print discovering that.
|
||||
it("does not inspect the job at startup", async () => {
|
||||
const info = vi.fn();
|
||||
const budget = await resolveGatewayShutdownBudget("external", { info, warn: vi.fn() });
|
||||
budget.log("startup");
|
||||
expect(execLaunchctl).not.toHaveBeenCalled();
|
||||
expect(budget.timeoutMs).toBe(325_000);
|
||||
expect(budget.nativeStopBudget).toBe(false);
|
||||
expect(info).toHaveBeenCalledWith(
|
||||
"shutdown budget at startup: drain=315000ms shutdown=325000ms reserve=10000ms exitMargin=5000ms; source=Gateway stop policy=330000ms",
|
||||
);
|
||||
});
|
||||
|
||||
it("keeps the platform-neutral policy when darwin is not running a launchd job", async () => {
|
||||
process.env = {};
|
||||
const warn = vi.fn();
|
||||
const budget = await resolveGatewayShutdownBudget(null, { info: vi.fn(), warn }, stoppingNow);
|
||||
expect(budget.timeoutMs).toBe(325_000);
|
||||
expect(budget.nativeStopBudget).toBe(false);
|
||||
expect(warn).not.toHaveBeenCalled();
|
||||
expect(execLaunchctl).not.toHaveBeenCalled();
|
||||
});
|
||||
|
||||
it("never reads systemd on darwin", async () => {
|
||||
execLaunchctl.mockResolvedValue(printed("SIGTERMed", "\texit timeout = 20\n\tpid = 4242\n"));
|
||||
await resolveGatewayShutdownBudget("external", { info: vi.fn(), warn: vi.fn() }, stoppingNow);
|
||||
expect(execSystem).not.toHaveBeenCalled();
|
||||
expect(execUser).not.toHaveBeenCalled();
|
||||
expect(readFile).not.toHaveBeenCalled();
|
||||
});
|
||||
});
|
||||
|
|
|
|||
|
|
@ -2,13 +2,58 @@ import { performance } from "node:perf_hooks";
|
|||
import { LAUNCH_AGENT_EXIT_TIMEOUT_SECONDS } from "../../daemon/launchd-plist.js";
|
||||
import {
|
||||
GATEWAY_SERVICE_STOP_TIMEOUT_MS,
|
||||
GATEWAY_SHUTDOWN_RESERVE_MS,
|
||||
GATEWAY_SHUTDOWN_TIMEOUT_MS,
|
||||
GATEWAY_SUPERVISOR_EXIT_MARGIN_MS,
|
||||
resolveShutdownReserveMs,
|
||||
resolveSupervisorExitMarginMs,
|
||||
} from "../../infra/gateway-shutdown-budget.js";
|
||||
import { readLaunchdStopTimeout } from "../../infra/launchd-stop-timeout.js";
|
||||
import { readSystemdStopTimeout } from "../../infra/systemd-stop-timeout.js";
|
||||
import type { GatewayRunSignalAction } from "./run-loop-request.js";
|
||||
|
||||
type NativeStopTimeout = { timeoutMs: number; source: string };
|
||||
|
||||
/**
|
||||
* Ask whichever supervisor actually enforces the deadline on this platform.
|
||||
*
|
||||
* Three independent answers. `stop` is a deadline that may be spent as a native
|
||||
* stop budget. `warning` is what the operator needs to hear. `inconclusive` says
|
||||
* the probe could not establish an answer at all, which is the only case the
|
||||
* retained-budget safety net below is for: a read that positively determined no
|
||||
* launchd deadline governs this stop is an answer, not a failure, so retaining a
|
||||
* startup budget and warning that the timeout "could not be confirmed" would be
|
||||
* false on every in-process restart of a launchd-owned Gateway.
|
||||
*/
|
||||
async function readNativeStopTimeout(stopping: boolean): Promise<{
|
||||
stop: NativeStopTimeout | null;
|
||||
warning?: string;
|
||||
inconclusive: boolean;
|
||||
}> {
|
||||
if (process.platform === "linux") {
|
||||
const systemd = await readSystemdStopTimeout();
|
||||
// Unchanged from the linux-only original: absent unit or warned read both
|
||||
// count as unconfirmed there.
|
||||
return {
|
||||
stop: systemd,
|
||||
warning: systemd?.warning,
|
||||
inconclusive: !systemd || Boolean(systemd.warning),
|
||||
};
|
||||
}
|
||||
// launchd's ExitTimeOut bounds a stop that launchd is running and nothing else:
|
||||
// an externally delivered SIGTERM never starts that clock, and the job outlives
|
||||
// the deadline untouched. There is no enforcing deadline to read before a stop
|
||||
// is under way, and reading one at startup would spend a launchctl print only
|
||||
// to adopt a deadline that does not govern the stop the Gateway will get.
|
||||
if (process.platform === "darwin" && stopping) {
|
||||
const read = await readLaunchdStopTimeout();
|
||||
// Warned but non-null is the defaulted-value case: launchd is confirmed to be
|
||||
// stopping the job and only its deadline had to be guessed, so a clock is
|
||||
// genuinely running and nothing needs retaining. Only a warning with no
|
||||
// deadline at all means the probe established nothing.
|
||||
return { ...read, inconclusive: read.stop === null && read.warning !== undefined };
|
||||
}
|
||||
return { stop: null, inconclusive: false };
|
||||
}
|
||||
|
||||
export async function resolveGatewayShutdownBudget(
|
||||
supervisor: string | null,
|
||||
logger: { info(message: string): void; warn(message: string): void },
|
||||
|
|
@ -17,37 +62,45 @@ export async function resolveGatewayShutdownBudget(
|
|||
acceptedAtMs: number;
|
||||
},
|
||||
) {
|
||||
// Restart ownership may be external while systemd still enforces the stop deadline.
|
||||
const systemdStop = process.platform === "linux" ? await readSystemdStopTimeout() : null;
|
||||
// Restart ownership may be external while the platform supervisor still
|
||||
// enforces the stop deadline. That holds on darwin exactly as it does on linux.
|
||||
const native = await readNativeStopTimeout(refresh !== undefined);
|
||||
const nativeStop = native.stop;
|
||||
const retained =
|
||||
refresh?.previous.nativeStopBudget && (!systemdStop || systemdStop.warning)
|
||||
? refresh.previous
|
||||
: undefined;
|
||||
if (systemdStop?.warning) {
|
||||
logger.warn(systemdStop.warning);
|
||||
refresh?.previous.nativeStopBudget && native.inconclusive ? refresh.previous : undefined;
|
||||
if (native.warning) {
|
||||
logger.warn(native.warning);
|
||||
}
|
||||
if (retained) {
|
||||
logger.warn(
|
||||
`Retaining the startup shutdown budget of ${retained.timeoutMs}ms because the current systemd stop timeout could not be confirmed.`,
|
||||
`Retaining the startup shutdown budget of ${retained.timeoutMs}ms because the current supervisor stop timeout could not be confirmed.`,
|
||||
);
|
||||
}
|
||||
const stop = systemdStop ?? {
|
||||
const stop = nativeStop ?? {
|
||||
timeoutMs:
|
||||
supervisor === "launchd"
|
||||
? LAUNCH_AGENT_EXIT_TIMEOUT_SECONDS * 1_000
|
||||
: GATEWAY_SERVICE_STOP_TIMEOUT_MS,
|
||||
source: supervisor === "launchd" ? "launchd ExitTimeOut" : "Gateway stop policy",
|
||||
};
|
||||
const nativeStopBudget = systemdStop !== null || supervisor === "launchd" || Boolean(retained);
|
||||
// ExitTimeOut=0 is unlimited. It is an observed job value, not a native
|
||||
// deadline that may cap a requested restart or arm a forced exit.
|
||||
const nativeStopBudget = nativeStop
|
||||
? Number.isFinite(nativeStop.timeoutMs)
|
||||
: supervisor === "launchd" || Boolean(retained);
|
||||
// An operator job may enforce a deadline far shorter than the policy these fixed
|
||||
// allowances were sized against, so each is capped at a share of what it is carved
|
||||
// from. A deadline long enough to fund them is unaffected; a short one keeps a
|
||||
// proportional drain instead of surrendering all of it to margin and reserve.
|
||||
const exitMarginMs = resolveSupervisorExitMarginMs(stop.timeoutMs);
|
||||
const limitMs =
|
||||
retained?.timeoutMs ??
|
||||
Math.min(GATEWAY_SHUTDOWN_TIMEOUT_MS, stop.timeoutMs - GATEWAY_SUPERVISOR_EXIT_MARGIN_MS);
|
||||
retained?.timeoutMs ?? Math.min(GATEWAY_SHUTDOWN_TIMEOUT_MS, stop.timeoutMs - exitMarginMs);
|
||||
const elapsedMs =
|
||||
refresh && nativeStopBudget
|
||||
? Math.max(0, Math.ceil(performance.now() - refresh.acceptedAtMs))
|
||||
: 0;
|
||||
const timeoutMs = Math.max(0, limitMs - elapsedMs);
|
||||
const reserveMs = Math.min(GATEWAY_SHUTDOWN_RESERVE_MS, timeoutMs);
|
||||
const reserveMs = resolveShutdownReserveMs(timeoutMs);
|
||||
return {
|
||||
nativeStopBudget,
|
||||
timeoutMs,
|
||||
|
|
@ -64,7 +117,10 @@ export async function resolveGatewayShutdownBudget(
|
|||
},
|
||||
log: (phase: "startup" | "shutdown") => {
|
||||
logger.info(
|
||||
`shutdown budget at ${phase}: drain=${Math.max(0, timeoutMs - GATEWAY_SHUTDOWN_RESERVE_MS)}ms shutdown=${timeoutMs}ms reserve=${reserveMs}ms exitMargin=${GATEWAY_SUPERVISOR_EXIT_MARGIN_MS}ms; source=${retained ? `startup shutdown budget=${retained.timeoutMs}ms` : `${stop.source}=${stop.timeoutMs}ms`}`,
|
||||
// Report the drain and margin actually spent. Subtracting the unscaled
|
||||
// reserve constant here understated a short budget's drain by the amount the
|
||||
// scaled reserve gave back.
|
||||
`shutdown budget at ${phase}: drain=${Math.max(0, timeoutMs - reserveMs)}ms shutdown=${timeoutMs}ms reserve=${reserveMs}ms exitMargin=${exitMarginMs}ms; source=${retained ? `startup shutdown budget=${retained.timeoutMs}ms` : `${stop.source}=${stop.timeoutMs}ms`}`,
|
||||
);
|
||||
},
|
||||
};
|
||||
|
|
|
|||
678
src/cli/gateway-cli/run-loop.launchd.test.ts
Normal file
678
src/cli/gateway-cli/run-loop.launchd.test.ts
Normal file
|
|
@ -0,0 +1,678 @@
|
|||
// darwin launchd-supervised run-loop cases. These live beside run-loop.test.ts because
|
||||
// the darwin stop budget reads the launchd job on every stop/restart request, so each
|
||||
// case here has to state the deadline launchd is enforcing instead of letting the real
|
||||
// reader spawn launchctl print, and run-loop.test.ts is at its line cap.
|
||||
import { performance } from "node:perf_hooks";
|
||||
import { afterEach, beforeEach, describe, expect, it, vi } from "vitest";
|
||||
import { useAutoCleanupTempDirTracker } from "../../../test/helpers/temp-dir.js";
|
||||
import type { HostedGatewayStop } from "../../daemon/hosted-stop.js";
|
||||
import type { GatewayActiveWorkSnapshot } from "../../infra/gateway-active-work.js";
|
||||
import type { GatewayBootLifecycleCompletion } from "../../infra/gateway-boot-lifecycle.js";
|
||||
import type { GatewayRestartIntent } from "../../infra/restart-intent.js";
|
||||
import { SUPERVISOR_HINT_ENV_VARS } from "../../infra/supervisor-markers.js";
|
||||
import type { RuntimeEnv } from "../../runtime.js";
|
||||
import { captureEnv, deleteTestEnvValue } from "../../test-utils/env.js";
|
||||
import type { GatewayRestartSnapshot } from "../daemon-cli/restart-health.js";
|
||||
import {
|
||||
createActiveWorkSnapshot,
|
||||
createCloseMock,
|
||||
createRuntimeWithExitSignal,
|
||||
createSignaledStart,
|
||||
originalPlatformDescriptor,
|
||||
type UpdateRespawnResultFixture,
|
||||
setPlatform,
|
||||
waitForStart,
|
||||
withIsolatedSignals,
|
||||
} from "./run-loop.test-support.js";
|
||||
|
||||
useAutoCleanupTempDirTracker(afterEach);
|
||||
const { readCgroup } = vi.hoisted(() => ({ readCgroup: vi.fn() }));
|
||||
|
||||
const spawnProcess = vi.hoisted(() => vi.fn<typeof import("node:child_process").spawn>());
|
||||
vi.mock("node:child_process", async () => {
|
||||
const actual = await vi.importActual<typeof import("node:child_process")>("node:child_process");
|
||||
const { mockNodeChildProcessModule } =
|
||||
await import("../../gateway/server-methods/node-child-process.test-support.js");
|
||||
if (!spawnProcess.getMockImplementation()) {
|
||||
spawnProcess.mockImplementation(actual.spawn);
|
||||
}
|
||||
const mocked = await mockNodeChildProcessModule({});
|
||||
vi.spyOn(mocked, "spawn").mockImplementation(spawnProcess);
|
||||
return mocked;
|
||||
});
|
||||
|
||||
vi.mock("node:fs/promises", async (original) => {
|
||||
const actual = await original<typeof import("node:fs/promises")>();
|
||||
// Foreground fixtures must not inherit the CI runner's systemd service or filesystem timing.
|
||||
const readFile = (...args: Parameters<typeof actual.readFile>) =>
|
||||
args[0] === "/proc/self/cgroup" ? readCgroup() : actual.readFile(...args);
|
||||
return { ...actual, readFile, default: { ...actual, readFile } };
|
||||
});
|
||||
|
||||
const systemctl = vi.fn(async () => ({
|
||||
code: 0,
|
||||
stdout: "LoadState=loaded\nTimeoutStopUSec=5min 30s",
|
||||
stderr: "",
|
||||
}));
|
||||
vi.mock("../../daemon/systemd-exec.js", () => ({
|
||||
execSystemctl: () => systemctl(),
|
||||
execSystemctlUser: () => systemctl(),
|
||||
}));
|
||||
|
||||
// darwin now refreshes the stop budget on every stop/restart request, and the real
|
||||
// reader spawns launchctl print against three domains through the mocked
|
||||
// node:child_process module. Keep launchctl out of the loop and let each case state
|
||||
// the deadline launchd is enforcing. The reader reports stop: null when launchd is not
|
||||
// running the stop, which leaves the platform-neutral policy in force.
|
||||
const readLaunchdStopTimeout = vi.fn<
|
||||
typeof import("../../infra/launchd-stop-timeout.js").readLaunchdStopTimeout
|
||||
>(async () => ({ stop: null }));
|
||||
vi.mock("../../infra/launchd-stop-timeout.js", () => ({
|
||||
readLaunchdStopTimeout: (...args: Parameters<typeof readLaunchdStopTimeout>) =>
|
||||
readLaunchdStopTimeout(...args),
|
||||
}));
|
||||
|
||||
const acquireGatewayLock = vi.fn(async (_opts?: { port?: number }) => ({
|
||||
release: vi.fn(async () => {}),
|
||||
}));
|
||||
const hostedStopExecute = vi.fn<HostedGatewayStop["execute"]>();
|
||||
const hostedStopDispose = vi.fn<HostedGatewayStop["dispose"]>();
|
||||
const hostedStopPrepare =
|
||||
vi.fn<typeof import("../../daemon/hosted-stop.js").prepareHostedGatewayStop>();
|
||||
vi.mock("../../daemon/hosted-stop.js", () => ({
|
||||
prepareHostedGatewayStop: (...args: Parameters<typeof hostedStopPrepare>) =>
|
||||
hostedStopPrepare(...args),
|
||||
}));
|
||||
const consumeGatewayRestartIntentPayloadSync = vi.fn<
|
||||
() => { reason?: string; force?: boolean; waitMs?: number } | null
|
||||
>(() => null);
|
||||
const consumeGatewayRestartIntent = vi.fn<() => GatewayRestartIntent | null>(() => null);
|
||||
type ManagedUpdateOwner = NonNullable<GatewayRestartIntent["successorOwner"]>;
|
||||
const cancelManagedServiceUpdateHandoff = vi.fn<
|
||||
(_identity: ManagedUpdateOwner) => Promise<false | "restored-in-process" | "restart-after-exit">
|
||||
>(async () => "restored-in-process");
|
||||
const claimManagedServiceUpdateHandoff = vi.fn((_identity: ManagedUpdateOwner) => true);
|
||||
const isForegroundUpdateHandoff = vi.fn((_identity: ManagedUpdateOwner) => false);
|
||||
const completeForegroundUpdateHandoffAfterClose =
|
||||
vi.fn<
|
||||
typeof import("../../infra/update-managed-service-handoff.js").completeForegroundUpdateHandoffAfterClose
|
||||
>();
|
||||
const captureForegroundUpdateHandoffStop =
|
||||
vi.fn<
|
||||
typeof import("../../infra/update-managed-service-handoff.js").captureForegroundUpdateHandoffStop
|
||||
>();
|
||||
const requestManagedServiceUpdateHandoffPark = vi.fn(async (_identity: ManagedUpdateOwner) => true);
|
||||
const waitForSystemServiceUpdateHandoffs = vi.fn<() => Promise<void> | undefined>();
|
||||
const commitManagedServiceUpdateHandoff = vi.fn(
|
||||
async (_identity: ManagedUpdateOwner, _outcome?: "update" | "restore") => true,
|
||||
);
|
||||
const consumeGatewayRestartAuthorization = vi.fn(() => true);
|
||||
const consumeGatewayRestartIntentSync = vi.fn(() => false);
|
||||
const isGatewayRestartExternallyAllowed = vi.fn(() => false);
|
||||
const markGatewayRestartHandled = vi.fn();
|
||||
const peekGatewayRestartReason = vi.fn<() => string | undefined>(() => undefined);
|
||||
const resetGatewayRestartStateForInProcessRestart = vi.fn();
|
||||
const resetGatewaySuspendCoordinatorForLifecycleRestart = vi.fn();
|
||||
const consumeGatewaySuspendHandoff =
|
||||
vi.fn<typeof import("../../infra/gateway-suspend-coordinator.js").consumeGatewaySuspendHandoff>();
|
||||
const disarmGatewaySuspendHandoff = vi.fn();
|
||||
const rollbackGatewayRestartSignalAdmission = vi.fn();
|
||||
const requestGatewayRestartWithSignalAdmission = vi.fn(() => ({ status: "emitted" as const }));
|
||||
const writeGatewayRestartHandoffSync = vi.fn(
|
||||
(
|
||||
_opts: unknown,
|
||||
): {
|
||||
kind: "gateway-supervisor-restart-handoff";
|
||||
version: 1;
|
||||
intentId: string;
|
||||
pid: number;
|
||||
createdAt: number;
|
||||
expiresAt: number;
|
||||
source: "unknown";
|
||||
restartKind: "full-process";
|
||||
supervisorMode: "external";
|
||||
} | null => ({
|
||||
kind: "gateway-supervisor-restart-handoff",
|
||||
version: 1,
|
||||
intentId: "test-intent",
|
||||
pid: process.pid,
|
||||
createdAt: Date.now(),
|
||||
expiresAt: Date.now() + 60_000,
|
||||
source: "unknown",
|
||||
restartKind: "full-process",
|
||||
supervisorMode: "external",
|
||||
}),
|
||||
);
|
||||
const scheduleGatewayRestart = vi.fn((_opts?: { delayMs?: number; reason?: string }) => ({
|
||||
ok: true,
|
||||
pid: process.pid,
|
||||
signal: "SIGUSR2" as const,
|
||||
delayMs: 0,
|
||||
mode: "emit" as const,
|
||||
coalesced: false,
|
||||
cooldownMsApplied: 0,
|
||||
}));
|
||||
const idleActiveWorkSnapshot = createActiveWorkSnapshot();
|
||||
const createGatewayActiveWorkSnapshot = vi.fn(() => idleActiveWorkSnapshot);
|
||||
const waitForGatewayActiveWork = vi.fn(
|
||||
async (
|
||||
_timeoutMs?: number,
|
||||
options?: { onSnapshot?: (snapshot: GatewayActiveWorkSnapshot) => void },
|
||||
) => {
|
||||
const snapshot = createGatewayActiveWorkSnapshot();
|
||||
options?.onSnapshot?.(snapshot);
|
||||
return { drained: snapshot.idle, snapshot };
|
||||
},
|
||||
);
|
||||
const advanceCronActiveJobGeneration = vi.fn();
|
||||
const resetCronActiveJobs = vi.fn();
|
||||
const abortActiveCronTaskRuns = vi.fn((_reason?: string) => 0);
|
||||
const retireActiveCronTaskRunTracking = vi.fn();
|
||||
const waitForActiveCronTaskRuns = vi.fn(async (_timeoutMs?: number) => ({
|
||||
drained: true,
|
||||
active: 0,
|
||||
}));
|
||||
const waitForActiveCronJobs = vi.fn(async (_timeoutMs?: number) => ({
|
||||
drained: true,
|
||||
active: 0,
|
||||
}));
|
||||
const reloadTaskRuntimeStateFromStore = vi.fn();
|
||||
const clearRuntimeConfigSnapshot = vi.fn();
|
||||
const restartGatewayProcessWithFreshPid = vi.fn<
|
||||
(_opts?: { env?: NodeJS.ProcessEnv }) => {
|
||||
mode: "supervised" | "disabled" | "failed";
|
||||
detail?: string;
|
||||
exitCode?: number;
|
||||
handoffSpawned?: Promise<boolean>;
|
||||
}
|
||||
>(() => ({ mode: "disabled" }));
|
||||
const respawnGatewayProcessForUpdate = vi.fn<
|
||||
(_opts?: { env?: NodeJS.ProcessEnv }) => UpdateRespawnResultFixture
|
||||
>(() => ({ mode: "disabled", detail: "OPENCLAW_NO_RESPAWN" }));
|
||||
const { killProcessTree } = vi.hoisted(() => ({ killProcessTree: vi.fn() }));
|
||||
vi.mock("../../process/kill-tree.js", async (importOriginal) => ({
|
||||
...(await importOriginal<typeof import("../../process/kill-tree.js")>()),
|
||||
killProcessTree,
|
||||
}));
|
||||
const markUpdateRestartSentinelFailure = vi.fn<(reason: string) => Promise<null>>(async () => null);
|
||||
const writeRestartSentinelIfUnchanged = vi.fn<
|
||||
typeof import("../../infra/restart-sentinel.js").writeRestartSentinelIfUnchanged
|
||||
>(async () => null);
|
||||
const readRestartSentinelReadOnly =
|
||||
vi.fn<typeof import("../../infra/restart-sentinel.js").readRestartSentinelReadOnly>();
|
||||
const waitForGatewayHealthyRestart =
|
||||
vi.fn<typeof import("../daemon-cli/restart-health.js").waitForGatewayHealthyRestart>();
|
||||
vi.mock("../daemon-cli/restart-health.js", async (importOriginal) => ({
|
||||
...(await importOriginal<typeof import("../daemon-cli/restart-health.js")>()),
|
||||
waitForGatewayHealthyRestart: (...args: Parameters<typeof waitForGatewayHealthyRestart>) =>
|
||||
waitForGatewayHealthyRestart(...args),
|
||||
}));
|
||||
const respawnHealth = (
|
||||
overrides: Partial<GatewayRestartSnapshot> = {},
|
||||
): GatewayRestartSnapshot => ({
|
||||
runtime: { status: "running", pid: 7777 },
|
||||
portUsage: { port: 18789, status: "busy", listeners: [{ pid: 7777 }], hints: [] },
|
||||
healthy: true,
|
||||
waitOutcome: "healthy",
|
||||
staleGatewayPids: [],
|
||||
...overrides,
|
||||
});
|
||||
const abortPendingChannelReloads = vi.fn();
|
||||
const abortEmbeddedAgentRun = vi.fn(
|
||||
(_sessionId?: string, _opts?: { mode?: "all" | "compacting"; reason?: "restart" }) => false,
|
||||
);
|
||||
const gatewayLog = {
|
||||
debug: vi.fn(),
|
||||
info: vi.fn(),
|
||||
warn: vi.fn(),
|
||||
error: vi.fn(),
|
||||
};
|
||||
const flushLogger = vi.fn(async () => {});
|
||||
const writeDiagnosticStabilityBundleForFailureSync = vi.fn(() => ({
|
||||
message: "stability bundle recorded",
|
||||
}));
|
||||
const hasManagedProviderLocalServices = vi.fn(() => false);
|
||||
const stopManagedProviderLocalServices = vi.fn(async () => {});
|
||||
const cancelShutdownHardExitWatchdog = vi.fn();
|
||||
const armShutdownHardExitWatchdog = vi.fn(
|
||||
(_params: { delayMs: number; onError: (error: unknown) => void }) => ({
|
||||
cancel: cancelShutdownHardExitWatchdog,
|
||||
}),
|
||||
);
|
||||
|
||||
vi.mock("../../infra/gateway-lock.js", async (original) => ({
|
||||
...(await original<typeof import("../../infra/gateway-lock.js")>()),
|
||||
acquireGatewayLock: (opts?: { port?: number }) => acquireGatewayLock(opts),
|
||||
}));
|
||||
|
||||
vi.mock("../../infra/restart.js", async (importOriginal) => {
|
||||
const actual = await importOriginal<typeof import("../../infra/restart.js")>();
|
||||
return {
|
||||
...actual,
|
||||
consumeGatewayRestartIntent: () => consumeGatewayRestartIntent(),
|
||||
consumeGatewayRestartAuthorization: () => consumeGatewayRestartAuthorization(),
|
||||
isGatewayRestartExternallyAllowed: () => isGatewayRestartExternallyAllowed(),
|
||||
markGatewayRestartHandled: () => markGatewayRestartHandled(),
|
||||
peekGatewayRestartReason: () => peekGatewayRestartReason(),
|
||||
resetGatewayRestartStateForInProcessRestart: () =>
|
||||
resetGatewayRestartStateForInProcessRestart(),
|
||||
rollbackGatewayRestartSignalAdmission: () => rollbackGatewayRestartSignalAdmission(),
|
||||
requestGatewayRestartWithSignalAdmission,
|
||||
scheduleGatewayRestart: (opts?: { delayMs?: number; reason?: string }) =>
|
||||
scheduleGatewayRestart(opts),
|
||||
};
|
||||
});
|
||||
|
||||
vi.mock("../../infra/restart-intent.js", () => ({
|
||||
consumeGatewayRestartIntentPayloadSync: () => consumeGatewayRestartIntentPayloadSync(),
|
||||
consumeGatewayRestartIntentSync: () => consumeGatewayRestartIntentSync(),
|
||||
}));
|
||||
|
||||
vi.mock("../../infra/update-managed-service-handoff.js", () => ({
|
||||
waitForSystemServiceUpdateHandoffs: () => waitForSystemServiceUpdateHandoffs(),
|
||||
captureForegroundUpdateHandoffStop: (
|
||||
params: Parameters<typeof captureForegroundUpdateHandoffStop>[0],
|
||||
) => captureForegroundUpdateHandoffStop(params),
|
||||
isForegroundUpdateHandoff: (identity: ManagedUpdateOwner) => isForegroundUpdateHandoff(identity),
|
||||
completeForegroundUpdateHandoffAfterClose: (identity: ManagedUpdateOwner) =>
|
||||
completeForegroundUpdateHandoffAfterClose(identity),
|
||||
cancelManagedServiceUpdateHandoff: (identity: ManagedUpdateOwner) =>
|
||||
cancelManagedServiceUpdateHandoff(identity),
|
||||
claimManagedServiceUpdateHandoff: (identity: ManagedUpdateOwner) =>
|
||||
claimManagedServiceUpdateHandoff(identity),
|
||||
requestManagedServiceUpdateHandoffPark: (identity: ManagedUpdateOwner) =>
|
||||
requestManagedServiceUpdateHandoffPark(identity),
|
||||
commitManagedServiceUpdateHandoff: (
|
||||
identity: ManagedUpdateOwner,
|
||||
outcome?: "update" | "restore",
|
||||
) => commitManagedServiceUpdateHandoff(identity, outcome),
|
||||
}));
|
||||
|
||||
vi.mock("../../infra/gateway-suspend-coordinator.js", () => ({
|
||||
consumeGatewaySuspendHandoff: (...args: Parameters<typeof consumeGatewaySuspendHandoff>) =>
|
||||
consumeGatewaySuspendHandoff(...args),
|
||||
disarmGatewaySuspendHandoff: (...args: unknown[]) => disarmGatewaySuspendHandoff(...args),
|
||||
resetGatewaySuspendCoordinatorForLifecycleRestart: () =>
|
||||
resetGatewaySuspendCoordinatorForLifecycleRestart(),
|
||||
}));
|
||||
|
||||
vi.mock("../../infra/process-respawn.js", async (importOriginal) => ({
|
||||
...(await importOriginal<typeof import("../../infra/process-respawn.js")>()),
|
||||
respawnGatewayProcessForUpdate: (opts?: { env?: NodeJS.ProcessEnv }) =>
|
||||
respawnGatewayProcessForUpdate(opts),
|
||||
restartGatewayProcessWithFreshPid: (opts?: { env?: NodeJS.ProcessEnv }) =>
|
||||
restartGatewayProcessWithFreshPid(opts),
|
||||
}));
|
||||
|
||||
vi.mock("../../infra/restart-sentinel.js", () => ({
|
||||
readRestartSentinelReadOnly: () => readRestartSentinelReadOnly(),
|
||||
markUpdateRestartSentinelFailure: (reason: string) => markUpdateRestartSentinelFailure(reason),
|
||||
writeRestartSentinelIfUnchanged: (...args: Parameters<typeof writeRestartSentinelIfUnchanged>) =>
|
||||
writeRestartSentinelIfUnchanged(...args),
|
||||
}));
|
||||
|
||||
vi.mock("../../infra/restart-handoff.js", () => ({
|
||||
writeGatewayRestartHandoffSync: (opts: unknown) => writeGatewayRestartHandoffSync(opts),
|
||||
}));
|
||||
|
||||
vi.mock("../../infra/gateway-active-work.js", () => ({
|
||||
createGatewayActiveWorkSnapshot: () => createGatewayActiveWorkSnapshot(),
|
||||
waitForGatewayActiveWork: (
|
||||
timeoutMs?: number,
|
||||
options?: { onSnapshot?: (snapshot: GatewayActiveWorkSnapshot) => void },
|
||||
) => waitForGatewayActiveWork(timeoutMs, options),
|
||||
}));
|
||||
|
||||
vi.mock("../../cron/active-jobs.js", () => ({
|
||||
advanceCronActiveJobGeneration: () => advanceCronActiveJobGeneration(),
|
||||
resetCronActiveJobs: () => resetCronActiveJobs(),
|
||||
waitForActiveCronJobs: (timeoutMs: number) => waitForActiveCronJobs(timeoutMs),
|
||||
}));
|
||||
|
||||
vi.mock("../../cron/service/active-run-cancellation.js", () => ({
|
||||
abortActiveCronTaskRuns: (reason?: string) => abortActiveCronTaskRuns(reason),
|
||||
retireActiveCronTaskRunTracking: () => retireActiveCronTaskRunTracking(),
|
||||
waitForActiveCronTaskRuns: (timeoutMs: number) => waitForActiveCronTaskRuns(timeoutMs),
|
||||
}));
|
||||
|
||||
vi.mock("../../tasks/runtime-internal.js", () => ({
|
||||
reloadTaskRuntimeStateFromStore: () => reloadTaskRuntimeStateFromStore(),
|
||||
}));
|
||||
|
||||
vi.mock("../../config/runtime-snapshot.js", () => ({
|
||||
clearRuntimeConfigSnapshot: () => clearRuntimeConfigSnapshot(),
|
||||
getRuntimeConfigSourceSnapshot: () => null,
|
||||
registerRuntimeConfigSnapshotPreparer: vi.fn(),
|
||||
}));
|
||||
|
||||
vi.mock("../../agents/embedded-agent-runner/runs.js", () => ({
|
||||
abortEmbeddedAgentRun: (
|
||||
sessionId?: string,
|
||||
opts?: { mode?: "all" | "compacting"; reason?: "restart" },
|
||||
) => abortEmbeddedAgentRun(sessionId, opts),
|
||||
}));
|
||||
|
||||
vi.mock("../../logging/subsystem.js", () => ({
|
||||
createSubsystemLogger: () => gatewayLog,
|
||||
}));
|
||||
|
||||
vi.mock("../../logging/logger.js", () => ({
|
||||
flushLogger: () => flushLogger(),
|
||||
}));
|
||||
|
||||
vi.mock("../../logging/diagnostic-stability-bundle.js", () => ({
|
||||
writeDiagnosticStabilityBundleForFailureSync,
|
||||
}));
|
||||
|
||||
vi.mock("../../agents/provider-runtime-lifecycle.js", () => ({
|
||||
hasManagedProviderLocalServices: () => hasManagedProviderLocalServices(),
|
||||
}));
|
||||
|
||||
vi.mock("../../agents/provider-local-service.js", () => ({
|
||||
stopManagedProviderLocalServices: () => stopManagedProviderLocalServices(),
|
||||
}));
|
||||
|
||||
vi.mock("../../gateway/server-reload-generation.js", () => ({
|
||||
abortPendingChannelReloads: () => abortPendingChannelReloads(),
|
||||
}));
|
||||
|
||||
vi.mock("./shutdown-hard-exit.js", () => ({
|
||||
armShutdownHardExitWatchdog: (params: { delayMs: number; onError: (error: unknown) => void }) =>
|
||||
armShutdownHardExitWatchdog(params),
|
||||
}));
|
||||
|
||||
async function runLoopWithStart(params: {
|
||||
start: ReturnType<typeof vi.fn>;
|
||||
runtime: RuntimeEnv;
|
||||
ownsProcessLifecycle?: boolean;
|
||||
lockPort?: number;
|
||||
healthHost?: string;
|
||||
beginBoot?: (startedAtMs: number) => void | Promise<void>;
|
||||
completeBoot?: (completion: GatewayBootLifecycleCompletion) => void;
|
||||
}) {
|
||||
vi.resetModules();
|
||||
const { runGatewayLoop } = await import("./run-loop.js");
|
||||
const loopPromise = runGatewayLoop({
|
||||
start: params.start as unknown as Parameters<typeof runGatewayLoop>[0]["start"],
|
||||
runtime: params.runtime,
|
||||
ownsProcessLifecycle: params.ownsProcessLifecycle,
|
||||
lockPort: params.lockPort,
|
||||
healthHost: params.healthHost,
|
||||
beginBoot: params.beginBoot,
|
||||
completeBoot: params.completeBoot,
|
||||
});
|
||||
return { loopPromise };
|
||||
}
|
||||
|
||||
async function createSignaledLoopHarness(exitCallOrder?: string[], ownsProcessLifecycle = false) {
|
||||
const close = createCloseMock();
|
||||
const { start, started } = createSignaledStart(close);
|
||||
const { runtime, exited } = createRuntimeWithExitSignal(exitCallOrder);
|
||||
const { loopPromise } = await runLoopWithStart({ start, runtime, ownsProcessLifecycle });
|
||||
await waitForStart(started);
|
||||
return { close, start, runtime, exited, loopPromise };
|
||||
}
|
||||
|
||||
function expectRestartHandoffCall(expected: {
|
||||
restartKind: "full-process" | "update-process";
|
||||
reason: string | undefined;
|
||||
supervisorMode: "external" | "launchd";
|
||||
}) {
|
||||
expect(writeGatewayRestartHandoffSync).toHaveBeenCalledTimes(1);
|
||||
const [handoff] = writeGatewayRestartHandoffSync.mock.calls[0] ?? [];
|
||||
if (!handoff || typeof handoff !== "object" || Array.isArray(handoff)) {
|
||||
throw new Error("expected restart handoff options object");
|
||||
}
|
||||
const processInstanceId = (handoff as { processInstanceId?: unknown }).processInstanceId;
|
||||
expect(typeof processInstanceId).toBe("string");
|
||||
if (typeof processInstanceId !== "string") {
|
||||
throw new Error("expected restart handoff processInstanceId string");
|
||||
}
|
||||
expect(processInstanceId).toMatch(
|
||||
/^[0-9a-f]{8}-[0-9a-f]{4}-[1-5][0-9a-f]{3}-[89ab][0-9a-f]{3}-[0-9a-f]{12}$/,
|
||||
);
|
||||
expect(handoff).toEqual({
|
||||
...expected,
|
||||
processInstanceId,
|
||||
});
|
||||
}
|
||||
|
||||
let gatewayWorkAdmissionActual: typeof import("../../process/gateway-work-admission.js");
|
||||
let supervisorEnvSnapshot: ReturnType<typeof captureEnv> | undefined;
|
||||
|
||||
beforeEach(async () => {
|
||||
vi.useRealTimers();
|
||||
vi.clearAllMocks();
|
||||
spawnProcess
|
||||
.mockReset()
|
||||
.mockImplementation(
|
||||
(await vi.importActual<typeof import("node:child_process")>("node:child_process")).spawn,
|
||||
);
|
||||
acquireGatewayLock.mockReset().mockImplementation(async () => ({
|
||||
release: vi.fn(async () => {}),
|
||||
}));
|
||||
setPlatform("linux");
|
||||
readCgroup.mockReset().mockResolvedValue("0::/\n");
|
||||
systemctl.mockReset().mockResolvedValue({
|
||||
code: 0,
|
||||
stdout: "LoadState=loaded\nTimeoutStopUSec=5min 30s",
|
||||
stderr: "",
|
||||
});
|
||||
// mockReset also drops any one-shot launchd deadline a previous case left queued.
|
||||
readLaunchdStopTimeout.mockReset().mockResolvedValue({ stop: null });
|
||||
hostedStopExecute.mockReset().mockResolvedValue({ outcome: "accepted" });
|
||||
hostedStopDispose.mockReset().mockResolvedValue(undefined);
|
||||
hostedStopPrepare.mockReset().mockImplementation(async (_owner, assertCurrent) => {
|
||||
assertCurrent();
|
||||
return { execute: hostedStopExecute, dispose: hostedStopDispose };
|
||||
});
|
||||
supervisorEnvSnapshot = captureEnv([...SUPERVISOR_HINT_ENV_VARS, "OPENCLAW_NO_RESPAWN"]);
|
||||
for (const key of [...SUPERVISOR_HINT_ENV_VARS, "OPENCLAW_NO_RESPAWN"]) {
|
||||
deleteTestEnvValue(key);
|
||||
}
|
||||
|
||||
// clearAllMocks preserves queued one-shot results. A skipped lifecycle branch
|
||||
// must not shift a stale supervisor or respawn decision into the next case.
|
||||
consumeGatewayRestartIntent.mockReset();
|
||||
consumeGatewayRestartIntentPayloadSync.mockReset().mockReturnValue(null);
|
||||
consumeGatewaySuspendHandoff.mockReset().mockReturnValue({ ok: true, value: false });
|
||||
disarmGatewaySuspendHandoff.mockClear();
|
||||
consumeGatewayRestartIntent.mockReturnValue(null);
|
||||
peekGatewayRestartReason.mockReset();
|
||||
peekGatewayRestartReason.mockReturnValue(undefined);
|
||||
restartGatewayProcessWithFreshPid.mockReset();
|
||||
restartGatewayProcessWithFreshPid.mockReturnValue({ mode: "disabled" });
|
||||
respawnGatewayProcessForUpdate.mockReset();
|
||||
waitForGatewayHealthyRestart.mockReset().mockResolvedValue(respawnHealth());
|
||||
writeRestartSentinelIfUnchanged.mockReset().mockResolvedValue(null);
|
||||
readRestartSentinelReadOnly.mockReset().mockResolvedValue(null);
|
||||
respawnGatewayProcessForUpdate.mockReturnValue({
|
||||
mode: "disabled",
|
||||
detail: "OPENCLAW_NO_RESPAWN",
|
||||
});
|
||||
hasManagedProviderLocalServices.mockReset();
|
||||
hasManagedProviderLocalServices.mockReturnValue(false);
|
||||
stopManagedProviderLocalServices.mockReset();
|
||||
stopManagedProviderLocalServices.mockResolvedValue(undefined);
|
||||
|
||||
gatewayWorkAdmissionActual = await vi.importActual("../../process/gateway-work-admission.js");
|
||||
gatewayWorkAdmissionActual.resetGatewayWorkAdmission();
|
||||
createGatewayActiveWorkSnapshot.mockReset();
|
||||
createGatewayActiveWorkSnapshot.mockReturnValue(idleActiveWorkSnapshot);
|
||||
waitForGatewayActiveWork.mockReset();
|
||||
waitForGatewayActiveWork.mockImplementation(async (_timeoutMs, options) => {
|
||||
const snapshot = createGatewayActiveWorkSnapshot();
|
||||
options?.onSnapshot?.(snapshot);
|
||||
return { drained: snapshot.idle, snapshot };
|
||||
});
|
||||
cancelManagedServiceUpdateHandoff.mockReset();
|
||||
cancelManagedServiceUpdateHandoff.mockResolvedValue("restored-in-process");
|
||||
claimManagedServiceUpdateHandoff.mockReset();
|
||||
claimManagedServiceUpdateHandoff.mockReturnValue(true);
|
||||
isForegroundUpdateHandoff.mockReset().mockReturnValue(false);
|
||||
completeForegroundUpdateHandoffAfterClose.mockReset().mockResolvedValue({ respawn: true });
|
||||
captureForegroundUpdateHandoffStop.mockReset().mockReturnValue(undefined);
|
||||
requestManagedServiceUpdateHandoffPark.mockReset();
|
||||
requestManagedServiceUpdateHandoffPark.mockResolvedValue(true);
|
||||
waitForSystemServiceUpdateHandoffs.mockReset().mockReturnValue(undefined);
|
||||
commitManagedServiceUpdateHandoff.mockReset();
|
||||
commitManagedServiceUpdateHandoff.mockResolvedValue(true);
|
||||
});
|
||||
|
||||
afterEach(() => {
|
||||
supervisorEnvSnapshot?.restore();
|
||||
supervisorEnvSnapshot = undefined;
|
||||
vi.useRealTimers();
|
||||
if (originalPlatformDescriptor) {
|
||||
Object.defineProperty(process, "platform", originalPlatformDescriptor);
|
||||
}
|
||||
});
|
||||
|
||||
describe("runGatewayLoop darwin launchd supervision", () => {
|
||||
it("waits briefly before exiting on launchd supervised restart", async () => {
|
||||
vi.clearAllMocks();
|
||||
peekGatewayRestartReason.mockReturnValue(undefined);
|
||||
try {
|
||||
setPlatform("darwin");
|
||||
process.env.OPENCLAW_LAUNCHD_LABEL = "ai.openclaw.gateway";
|
||||
restartGatewayProcessWithFreshPid.mockReturnValueOnce({
|
||||
mode: "supervised",
|
||||
handoffSpawned: Promise.resolve(true),
|
||||
});
|
||||
|
||||
await withIsolatedSignals(async ({ captureSignal }) => {
|
||||
const { runtime, exited } = await createSignaledLoopHarness();
|
||||
const restartSignal = captureSignal("SIGUSR2");
|
||||
|
||||
vi.useFakeTimers();
|
||||
restartSignal();
|
||||
await vi.advanceTimersByTimeAsync(1499);
|
||||
expect(runtime.exit).not.toHaveBeenCalled();
|
||||
await vi.advanceTimersByTimeAsync(1);
|
||||
|
||||
await expect(exited).resolves.toBe(0);
|
||||
expect(runtime.exit).toHaveBeenCalledWith(0);
|
||||
expectRestartHandoffCall({
|
||||
restartKind: "full-process",
|
||||
reason: undefined,
|
||||
supervisorMode: "launchd",
|
||||
});
|
||||
});
|
||||
} finally {
|
||||
vi.useRealTimers();
|
||||
delete process.env.OPENCLAW_LAUNCHD_LABEL;
|
||||
if (originalPlatformDescriptor) {
|
||||
Object.defineProperty(process, "platform", originalPlatformDescriptor);
|
||||
}
|
||||
}
|
||||
});
|
||||
|
||||
it("falls back in-process when the launchd restart handoff fails to spawn", async () => {
|
||||
vi.clearAllMocks();
|
||||
peekGatewayRestartReason.mockReturnValue(undefined);
|
||||
try {
|
||||
setPlatform("darwin");
|
||||
process.env.OPENCLAW_LAUNCHD_LABEL = "ai.openclaw.gateway";
|
||||
restartGatewayProcessWithFreshPid.mockReturnValueOnce({
|
||||
mode: "supervised",
|
||||
handoffSpawned: Promise.resolve(false),
|
||||
});
|
||||
|
||||
await withIsolatedSignals(async ({ captureSignal }) => {
|
||||
const { start, runtime, exited } = await createSignaledLoopHarness();
|
||||
const restartSignal = captureSignal("SIGUSR2");
|
||||
const sigint = captureSignal("SIGINT");
|
||||
|
||||
vi.useFakeTimers();
|
||||
restartSignal();
|
||||
await vi.advanceTimersByTimeAsync(1500);
|
||||
|
||||
expect(start).toHaveBeenCalledTimes(2);
|
||||
expect(runtime.exit).not.toHaveBeenCalled();
|
||||
expect(acquireGatewayLock).toHaveBeenCalledTimes(2);
|
||||
expect(gatewayLog.warn).toHaveBeenCalledWith(
|
||||
"launchd restart handoff failed to spawn; falling back to in-process restart",
|
||||
);
|
||||
|
||||
sigint();
|
||||
await expect(exited).resolves.toBe(0);
|
||||
});
|
||||
} finally {
|
||||
vi.useRealTimers();
|
||||
delete process.env.OPENCLAW_LAUNCHD_LABEL;
|
||||
if (originalPlatformDescriptor) {
|
||||
Object.defineProperty(process, "platform", originalPlatformDescriptor);
|
||||
}
|
||||
}
|
||||
});
|
||||
|
||||
it("leaves the successor to launchd after a SIGTERM restart intent", async () => {
|
||||
vi.clearAllMocks();
|
||||
consumeGatewayRestartIntentPayloadSync.mockReturnValueOnce({ reason: "gateway.restart" });
|
||||
setPlatform("darwin");
|
||||
process.env.OPENCLAW_LAUNCHD_LABEL = "ai.openclaw.gateway";
|
||||
restartGatewayProcessWithFreshPid.mockReturnValueOnce({
|
||||
mode: "supervised",
|
||||
handoffSpawned: Promise.resolve(true),
|
||||
});
|
||||
|
||||
await withIsolatedSignals(async ({ captureSignal }) => {
|
||||
const { start, exited } = await createSignaledLoopHarness();
|
||||
captureSignal("SIGTERM")();
|
||||
await expect(exited).resolves.toBe(0);
|
||||
expect(start).toHaveBeenCalledOnce();
|
||||
expect(restartGatewayProcessWithFreshPid).not.toHaveBeenCalled();
|
||||
expect(respawnGatewayProcessForUpdate).not.toHaveBeenCalled();
|
||||
});
|
||||
});
|
||||
|
||||
// The stop budget refresh is the reason darwin reads the launchd job at all, so prove
|
||||
// the job deadline it reports is what bounds the stop and arms the force exit.
|
||||
it("bounds a launchd-supervised stop on the deadline the printed job reports", async () => {
|
||||
vi.clearAllMocks();
|
||||
try {
|
||||
setPlatform("darwin");
|
||||
process.env.OPENCLAW_LAUNCHD_LABEL = "ai.openclaw.gateway";
|
||||
readLaunchdStopTimeout.mockResolvedValue({
|
||||
stop: { timeoutMs: 30_000, source: "launchd system/ai.openclaw.gateway exit timeout" },
|
||||
});
|
||||
hasManagedProviderLocalServices.mockReturnValue(true);
|
||||
stopManagedProviderLocalServices.mockReturnValue(new Promise<void>(() => {}));
|
||||
|
||||
await withIsolatedSignals(async ({ captureSignal }) => {
|
||||
const { close, runtime } = await createSignaledLoopHarness();
|
||||
|
||||
vi.useFakeTimers();
|
||||
const clock = vi.spyOn(performance, "now").mockImplementation(() => Date.now());
|
||||
try {
|
||||
captureSignal("SIGTERM")();
|
||||
await vi.advanceTimersByTimeAsync(24_999);
|
||||
|
||||
expect(close).toHaveBeenCalledOnce();
|
||||
expect(stopManagedProviderLocalServices).toHaveBeenCalledOnce();
|
||||
expect(runtime.exit).not.toHaveBeenCalled();
|
||||
await vi.advanceTimersByTimeAsync(1);
|
||||
|
||||
// launchd owns the successor, so an abandoned drain still exits 0 for the job.
|
||||
expect(runtime.exit).toHaveBeenCalledExactlyOnceWith(0);
|
||||
expect(writeDiagnosticStabilityBundleForFailureSync).toHaveBeenCalledWith(
|
||||
"gateway.stop_shutdown_timeout",
|
||||
undefined,
|
||||
);
|
||||
expect(gatewayLog.info).toHaveBeenCalledWith(
|
||||
"shutdown budget at shutdown: drain=15000ms shutdown=25000ms reserve=10000ms exitMargin=5000ms; source=launchd system/ai.openclaw.gateway exit timeout=30000ms",
|
||||
);
|
||||
} finally {
|
||||
clock.mockRestore();
|
||||
vi.clearAllTimers();
|
||||
vi.useRealTimers();
|
||||
}
|
||||
});
|
||||
} finally {
|
||||
delete process.env.OPENCLAW_LAUNCHD_LABEL;
|
||||
if (originalPlatformDescriptor) {
|
||||
Object.defineProperty(process, "platform", originalPlatformDescriptor);
|
||||
}
|
||||
}
|
||||
});
|
||||
});
|
||||
|
|
@ -79,9 +79,23 @@ it
|
|||
expect(exit, output).toEqual([0, null]);
|
||||
expect(elapsed, output).toBeLessThan(stopTimeoutMs);
|
||||
expect(output).toContain("process proof: acquisition-cancelled");
|
||||
// The fixture's launchd label is synthetic, so the per-stop re-inspection cannot
|
||||
// find a job to read and the startup budget is retained instead. Both
|
||||
// attributions prove the same thing this case is guarding: the shutdown budget
|
||||
// came from the 20 second native stop timeout and not from the platform-neutral
|
||||
// policy, which would report a 330000ms source and a far longer deadline.
|
||||
expect(output).toMatch(
|
||||
new RegExp(`shutdown budget at shutdown:.*source=.*=${stopTimeoutMs}ms`),
|
||||
new RegExp(
|
||||
`shutdown budget at shutdown:.*source=(?:.*=${stopTimeoutMs}ms|startup shutdown budget=${shutdownTimeoutMs}ms)`,
|
||||
),
|
||||
);
|
||||
if (process.platform === "darwin") {
|
||||
// The label resolves to no real job, so this cannot assert a successful read.
|
||||
// What it does assert is that the darwin probe ran inside a real spawned
|
||||
// Gateway on a real stop: only the launchd reader emits this, and reverting the
|
||||
// darwin dispatch removes it.
|
||||
expect(output).toContain("Unable to inspect the launchd job");
|
||||
}
|
||||
if (mode === "cooperative") {
|
||||
expect(output).toContain("process proof: acquisition-joined");
|
||||
expect(output).not.toContain("shutdown deadline reached");
|
||||
|
|
|
|||
|
|
@ -78,6 +78,19 @@ vi.mock("../../daemon/systemd-exec.js", () => ({
|
|||
execSystemctlUser: () => systemctl(),
|
||||
}));
|
||||
|
||||
// darwin now refreshes the stop budget on every stop/restart request, and the real
|
||||
// reader spawns launchctl print against three domains through the mocked
|
||||
// node:child_process module. Keep launchctl out of the loop and let each case state
|
||||
// the deadline launchd is enforcing. The reader reports stop: null when launchd is not
|
||||
// running the stop, which leaves the platform-neutral policy in force.
|
||||
const readLaunchdStopTimeout = vi.fn<
|
||||
typeof import("../../infra/launchd-stop-timeout.js").readLaunchdStopTimeout
|
||||
>(async () => ({ stop: null }));
|
||||
vi.mock("../../infra/launchd-stop-timeout.js", () => ({
|
||||
readLaunchdStopTimeout: (...args: Parameters<typeof readLaunchdStopTimeout>) =>
|
||||
readLaunchdStopTimeout(...args),
|
||||
}));
|
||||
|
||||
const acquireGatewayLock = vi.fn(async (_opts?: { port?: number }) => ({
|
||||
release: vi.fn(async () => {}),
|
||||
}));
|
||||
|
|
@ -468,6 +481,8 @@ beforeEach(async () => {
|
|||
stdout: "LoadState=loaded\nTimeoutStopUSec=5min 30s",
|
||||
stderr: "",
|
||||
});
|
||||
// mockReset also drops any one-shot launchd deadline a previous case left queued.
|
||||
readLaunchdStopTimeout.mockReset().mockResolvedValue({ stop: null });
|
||||
hostedStopExecute.mockReset().mockResolvedValue({ outcome: "accepted" });
|
||||
hostedStopDispose.mockReset().mockResolvedValue(undefined);
|
||||
hostedStopPrepare.mockReset().mockImplementation(async (_owner, assertCurrent) => {
|
||||
|
|
@ -2794,103 +2809,6 @@ describe("runGatewayLoop", () => {
|
|||
}
|
||||
});
|
||||
|
||||
it("waits briefly before exiting on launchd supervised restart", async () => {
|
||||
vi.clearAllMocks();
|
||||
peekGatewayRestartReason.mockReturnValue(undefined);
|
||||
try {
|
||||
setPlatform("darwin");
|
||||
process.env.OPENCLAW_LAUNCHD_LABEL = "ai.openclaw.gateway";
|
||||
restartGatewayProcessWithFreshPid.mockReturnValueOnce({
|
||||
mode: "supervised",
|
||||
handoffSpawned: Promise.resolve(true),
|
||||
});
|
||||
|
||||
await withIsolatedSignals(async ({ captureSignal }) => {
|
||||
const { runtime, exited } = await createSignaledLoopHarness();
|
||||
const restartSignal = captureSignal("SIGUSR2");
|
||||
|
||||
vi.useFakeTimers();
|
||||
restartSignal();
|
||||
await vi.advanceTimersByTimeAsync(1499);
|
||||
expect(runtime.exit).not.toHaveBeenCalled();
|
||||
await vi.advanceTimersByTimeAsync(1);
|
||||
|
||||
await expect(exited).resolves.toBe(0);
|
||||
expect(runtime.exit).toHaveBeenCalledWith(0);
|
||||
expectRestartHandoffCall({
|
||||
restartKind: "full-process",
|
||||
reason: undefined,
|
||||
supervisorMode: "launchd",
|
||||
});
|
||||
});
|
||||
} finally {
|
||||
vi.useRealTimers();
|
||||
delete process.env.OPENCLAW_LAUNCHD_LABEL;
|
||||
if (originalPlatformDescriptor) {
|
||||
Object.defineProperty(process, "platform", originalPlatformDescriptor);
|
||||
}
|
||||
}
|
||||
});
|
||||
|
||||
it("falls back in-process when the launchd restart handoff fails to spawn", async () => {
|
||||
vi.clearAllMocks();
|
||||
peekGatewayRestartReason.mockReturnValue(undefined);
|
||||
try {
|
||||
setPlatform("darwin");
|
||||
process.env.OPENCLAW_LAUNCHD_LABEL = "ai.openclaw.gateway";
|
||||
restartGatewayProcessWithFreshPid.mockReturnValueOnce({
|
||||
mode: "supervised",
|
||||
handoffSpawned: Promise.resolve(false),
|
||||
});
|
||||
|
||||
await withIsolatedSignals(async ({ captureSignal }) => {
|
||||
const { start, runtime, exited } = await createSignaledLoopHarness();
|
||||
const restartSignal = captureSignal("SIGUSR2");
|
||||
const sigint = captureSignal("SIGINT");
|
||||
|
||||
vi.useFakeTimers();
|
||||
restartSignal();
|
||||
await vi.advanceTimersByTimeAsync(1500);
|
||||
|
||||
expect(start).toHaveBeenCalledTimes(2);
|
||||
expect(runtime.exit).not.toHaveBeenCalled();
|
||||
expect(acquireGatewayLock).toHaveBeenCalledTimes(2);
|
||||
expect(gatewayLog.warn).toHaveBeenCalledWith(
|
||||
"launchd restart handoff failed to spawn; falling back to in-process restart",
|
||||
);
|
||||
|
||||
sigint();
|
||||
await expect(exited).resolves.toBe(0);
|
||||
});
|
||||
} finally {
|
||||
vi.useRealTimers();
|
||||
delete process.env.OPENCLAW_LAUNCHD_LABEL;
|
||||
if (originalPlatformDescriptor) {
|
||||
Object.defineProperty(process, "platform", originalPlatformDescriptor);
|
||||
}
|
||||
}
|
||||
});
|
||||
|
||||
it("leaves the successor to launchd after a SIGTERM restart intent", async () => {
|
||||
vi.clearAllMocks();
|
||||
consumeGatewayRestartIntentPayloadSync.mockReturnValueOnce({ reason: "gateway.restart" });
|
||||
setPlatform("darwin");
|
||||
process.env.OPENCLAW_LAUNCHD_LABEL = "ai.openclaw.gateway";
|
||||
restartGatewayProcessWithFreshPid.mockReturnValueOnce({
|
||||
mode: "supervised",
|
||||
handoffSpawned: Promise.resolve(true),
|
||||
});
|
||||
|
||||
await withIsolatedSignals(async ({ captureSignal }) => {
|
||||
const { start, exited } = await createSignaledLoopHarness();
|
||||
captureSignal("SIGTERM")();
|
||||
await expect(exited).resolves.toBe(0);
|
||||
expect(start).toHaveBeenCalledOnce();
|
||||
expect(restartGatewayProcessWithFreshPid).not.toHaveBeenCalled();
|
||||
expect(respawnGatewayProcessForUpdate).not.toHaveBeenCalled();
|
||||
});
|
||||
});
|
||||
|
||||
it("records external ownership even when native supervisor markers are inherited", async () => {
|
||||
vi.clearAllMocks();
|
||||
peekGatewayRestartReason.mockReturnValue(undefined);
|
||||
|
|
|
|||
|
|
@ -609,7 +609,7 @@ export async function runGatewayLoop(params: {
|
|||
return timer;
|
||||
};
|
||||
const timer = arm(startupBudget.timeoutMs);
|
||||
if (process.platform === "linux") {
|
||||
if (process.platform === "linux" || process.platform === "darwin") {
|
||||
void resolveGatewayShutdownBudget(supervisorMode, gatewayLog, {
|
||||
previous: startupBudget,
|
||||
acceptedAtMs: pendingRequest.acceptedAtMs,
|
||||
|
|
@ -761,7 +761,7 @@ export async function runGatewayLoop(params: {
|
|||
}
|
||||
|
||||
const completion = (async () => {
|
||||
if (process.platform === "linux") {
|
||||
if (process.platform === "linux" || process.platform === "darwin") {
|
||||
if (budget.nativeStopBudget && !getManagedUpdateOwner()) {
|
||||
armForceExitTimer(budget.timeoutMs);
|
||||
}
|
||||
|
|
|
|||
|
|
@ -4,5 +4,9 @@ export {
|
|||
GATEWAY_SUPERVISOR_EXIT_MARGIN_MS,
|
||||
GATEWAY_SHUTDOWN_TIMEOUT_MS,
|
||||
GATEWAY_SERVICE_STOP_TIMEOUT_MS,
|
||||
isRespawnedByLauncher,
|
||||
LAUNCH_AGENT_EXIT_TIMEOUT_SECONDS,
|
||||
resolveLauncherStopTimeoutMs,
|
||||
resolveShutdownReserveMs,
|
||||
resolveSupervisorExitMarginMs,
|
||||
} from "../../gateway-shutdown-budget.mjs";
|
||||
|
|
|
|||
376
src/infra/launchd-stop-timeout.test.ts
Normal file
376
src/infra/launchd-stop-timeout.test.ts
Normal file
|
|
@ -0,0 +1,376 @@
|
|||
import { afterEach, beforeEach, describe, expect, it, vi } from "vitest";
|
||||
import { isRespawnedByLauncher } from "./gateway-shutdown-budget.js";
|
||||
import { readLaunchdStopTimeout } from "./launchd-stop-timeout.js";
|
||||
|
||||
const { execLaunchctl } = vi.hoisted(() => ({ execLaunchctl: vi.fn() }));
|
||||
vi.mock("../daemon/launchd-exec.js", async (importOriginal) => ({
|
||||
...(await importOriginal<typeof import("../daemon/launchd-exec.js")>()),
|
||||
execLaunchctl,
|
||||
}));
|
||||
|
||||
// The authoritative list is module-local beside the deadline it authorises in
|
||||
// gateway-shutdown-budget.mjs, so it is restated here and tied back to the
|
||||
// implementation by the recognition cases at the end of this file.
|
||||
const RESPAWN_MARKERS = [
|
||||
"OPENCLAW_NODE_UPDATE_RESPAWNED",
|
||||
"OPENCLAW_COMPILE_CACHE_DISABLED_RESPAWNED",
|
||||
"OPENCLAW_PACKAGED_COMPILE_CACHE_RESPAWNED",
|
||||
] as const;
|
||||
const LAUNCHD_ENV = { XPC_SERVICE_NAME: "ai.openclaw.gateway" };
|
||||
// The service layout the recovery launcher bounds by the LaunchAgent exit timeout:
|
||||
// launchd names the job in XPC_SERVICE_NAME and the handoff carries the label, which
|
||||
// is the pair the launcher branches on and a respawned child inherits unchanged.
|
||||
const SERVICE_ENV = { ...LAUNCHD_ENV, OPENCLAW_LAUNCHD_LABEL: "ai.openclaw.gateway" };
|
||||
// Adds one of the markers a respawning launcher stamps on its child, which is what
|
||||
// separates a Gateway that launcher started from an unrelated parent.
|
||||
const RESPAWNED_SERVICE_ENV = { ...SERVICE_ENV, OPENCLAW_NODE_UPDATE_RESPAWNED: "1" };
|
||||
const result = (stdout: string) => ({ code: 0, stdout, stderr: "", termination: "exit" });
|
||||
|
||||
/**
|
||||
* Shaped like real `launchctl print` output rather than a bare field list: the
|
||||
* job's own `state` at one tab, then coalition blocks carrying their own
|
||||
* `state = active` at two tabs, then `job state`. A live Gateway LaunchDaemon
|
||||
* prints `state` three times in exactly this arrangement.
|
||||
*/
|
||||
const printed = (state: string, fields: string) =>
|
||||
result(
|
||||
`system/ai.openclaw.gateway = {\n\tactive count = 1\n\ttype = LaunchDaemon\n\tstate = ${state}\n\n${fields}\tresource coalition = {\n\t\tID = 18110\n\t\tstate = active\n\t}\n\n\tjetsam coalition = {\n\t\tID = 18111\n\t\tstate = active\n\t}\n\n\tjob state = running\n}\n`,
|
||||
);
|
||||
const stopping = (fields: string) => printed("SIGTERMed", fields);
|
||||
|
||||
beforeEach(() => {
|
||||
vi.stubGlobal("process", {
|
||||
...process,
|
||||
platform: "darwin",
|
||||
pid: 4242,
|
||||
ppid: 4241,
|
||||
getuid: () => 501,
|
||||
// The reader only runs inside a serving Gateway, and reconstructing an
|
||||
// undeclared launcher's timer replays the argv branch the launcher took.
|
||||
argv: ["/usr/local/bin/node", "/usr/local/bin/openclaw", "gateway", "run"],
|
||||
});
|
||||
execLaunchctl.mockReset();
|
||||
});
|
||||
afterEach(() => vi.unstubAllGlobals());
|
||||
|
||||
describe("launchd stop timeout reads the job launchd is stopping", () => {
|
||||
it("uses the operator job's effective exit timeout, not the template constant", async () => {
|
||||
execLaunchctl.mockResolvedValue(
|
||||
stopping("\tminimum runtime = 10\n\texit timeout = 5\n\tpid = 4242\n"),
|
||||
);
|
||||
await expect(readLaunchdStopTimeout(LAUNCHD_ENV)).resolves.toEqual({
|
||||
stop: { timeoutMs: 5_000, source: "launchd system/ai.openclaw.gateway exit timeout" },
|
||||
});
|
||||
expect(execLaunchctl).toHaveBeenCalledExactlyOnceWith(
|
||||
["print", "system/ai.openclaw.gateway"],
|
||||
2_000,
|
||||
);
|
||||
});
|
||||
|
||||
// The shared key-value parser keeps the LAST occurrence of a repeated key, and
|
||||
// `launchctl print` repeats `state` inside coalition blocks, so asking it for
|
||||
// `state` answers `active` and never sees the job at all. This case fails if
|
||||
// the reader ever goes back to that parser for the job's state.
|
||||
it("reads the job's own state, not a nested coalition's", async () => {
|
||||
execLaunchctl.mockResolvedValue(stopping("\texit timeout = 47\n\tpid = 4242\n"));
|
||||
await expect(readLaunchdStopTimeout(LAUNCHD_ENV)).resolves.toEqual({
|
||||
stop: { timeoutMs: 47_000, source: "launchd system/ai.openclaw.gateway exit timeout" },
|
||||
});
|
||||
});
|
||||
|
||||
// Measured on macOS 27 with ExitTimeOut 47: a plain `kill -TERM` left the job
|
||||
// printing `state = running` while the process handled the signal, and it was
|
||||
// still alive 85 seconds later. launchd never started its clock, so its
|
||||
// deadline bounds nothing and the caller keeps the budget it already had.
|
||||
it("declines the deadline when launchd is not the one stopping the job", async () => {
|
||||
execLaunchctl.mockResolvedValue(printed("running", "\texit timeout = 5\n\tpid = 4242\n"));
|
||||
// No deadline and nothing to warn about: this is the ordinary shape of a stop
|
||||
// that some other sender delivered.
|
||||
await expect(readLaunchdStopTimeout(LAUNCHD_ENV)).resolves.toEqual({ stop: null });
|
||||
// Our job was found in the first domain, so there is nothing left to search.
|
||||
expect(execLaunchctl).toHaveBeenCalledTimes(1);
|
||||
});
|
||||
|
||||
it("preserves direct-child drain when a marked launcher has not started its timer", async () => {
|
||||
execLaunchctl.mockResolvedValue(printed("running", "\texit timeout = 60\n\tpid = 4241\n"));
|
||||
await expect(readLaunchdStopTimeout(RESPAWNED_SERVICE_ENV)).resolves.toEqual({ stop: null });
|
||||
});
|
||||
|
||||
it("keeps launchd's unlimited exit timeout distinct from a missing one", async () => {
|
||||
execLaunchctl.mockResolvedValue(stopping("\texit timeout = 0\n\tpid = 4242\n"));
|
||||
await expect(readLaunchdStopTimeout(LAUNCHD_ENV)).resolves.toEqual({
|
||||
stop: {
|
||||
timeoutMs: Infinity,
|
||||
source: "launchd system/ai.openclaw.gateway unlimited exit timeout",
|
||||
},
|
||||
});
|
||||
});
|
||||
|
||||
it("caps an unlimited launchd stop at the launcher's finite timer", async () => {
|
||||
execLaunchctl.mockResolvedValue(stopping("\texit timeout = 0\n\tpid = 4241\n"));
|
||||
await expect(readLaunchdStopTimeout(RESPAWNED_SERVICE_ENV)).resolves.toEqual({
|
||||
stop: {
|
||||
timeoutMs: 19_000,
|
||||
source:
|
||||
"launchd system/ai.openclaw.gateway unlimited exit timeout capped at the launcher's 19000ms stop timer",
|
||||
},
|
||||
});
|
||||
});
|
||||
|
||||
// An empty value is deliberately not in this table: there is no state to fail to
|
||||
// recognise, so it belongs with the missing-state warning below.
|
||||
it.each(["waiting", "exited", "not running", "SIGTERM", "sigtermed"])(
|
||||
"treats the unrecognised state %j as not stopping",
|
||||
async (state) => {
|
||||
execLaunchctl.mockResolvedValue(printed(state, "\texit timeout = 5\n\tpid = 4242\n"));
|
||||
await expect(readLaunchdStopTimeout(LAUNCHD_ENV)).resolves.toEqual({ stop: null });
|
||||
},
|
||||
);
|
||||
|
||||
it("accepts any signal launchd reports having delivered", async () => {
|
||||
execLaunchctl.mockResolvedValue(printed("SIGKILLed", "\texit timeout = 9\n\tpid = 4242\n"));
|
||||
await expect(readLaunchdStopTimeout(LAUNCHD_ENV)).resolves.toEqual({
|
||||
stop: { timeoutMs: 9_000, source: "launchd system/ai.openclaw.gateway exit timeout" },
|
||||
});
|
||||
});
|
||||
|
||||
// An absent `state` is indistinguishable from a job launchd is not stopping, so the
|
||||
// budget would silently revert to the platform-neutral policy on a macOS that printed
|
||||
// this block differently. The deadline parsed from the same block is what makes the
|
||||
// case reportable, and the warning is what keeps it from being silent. An empty value
|
||||
// reaches the same place as a missing line, because neither yields a state to read.
|
||||
it.each([
|
||||
[
|
||||
"no state line",
|
||||
`system/ai.openclaw.gateway = {\n\tactive count = 1\n\ttype = LaunchDaemon\n\n\texit timeout = 47\n\tpid = 4242\n\tjob state = running\n}\n`,
|
||||
],
|
||||
["an empty state value", printed("", "\texit timeout = 47\n\tpid = 4242\n").stdout],
|
||||
])("warns when the job printed a deadline but %s", async (_label, stdout) => {
|
||||
execLaunchctl.mockResolvedValue(result(stdout));
|
||||
const read = await readLaunchdStopTimeout(LAUNCHD_ENV);
|
||||
expect(read.stop).toBeNull();
|
||||
expect(read.warning).toBe(
|
||||
"launchd system/ai.openclaw.gateway printed an exit timeout but no job state, so it is treated as not stopping and the Gateway stop policy is kept. Check the running job with launchctl print.",
|
||||
);
|
||||
});
|
||||
|
||||
// A job with neither field is an ordinary not-stopping answer and must stay quiet,
|
||||
// otherwise every in-process restart of a launchd-owned Gateway would warn.
|
||||
it("stays silent when the job is simply not stopping", async () => {
|
||||
execLaunchctl.mockResolvedValue(printed("running", "\texit timeout = 47\n\tpid = 4242\n"));
|
||||
await expect(readLaunchdStopTimeout(LAUNCHD_ENV)).resolves.toEqual({ stop: null });
|
||||
});
|
||||
|
||||
it("falls back to the gui domain when the job is not a LaunchDaemon", async () => {
|
||||
execLaunchctl
|
||||
.mockResolvedValueOnce({
|
||||
code: 113,
|
||||
stdout: "",
|
||||
stderr: "Could not find service",
|
||||
termination: "exit",
|
||||
})
|
||||
.mockResolvedValueOnce(stopping("\texit timeout = 20\n\tpid = 4242\n"));
|
||||
await expect(readLaunchdStopTimeout(LAUNCHD_ENV)).resolves.toEqual({
|
||||
stop: { timeoutMs: 20_000, source: "launchd gui/501/ai.openclaw.gateway exit timeout" },
|
||||
});
|
||||
});
|
||||
|
||||
// A service account with no logged-in session has no gui domain: launchctl
|
||||
// answers 125 "Domain does not support specified action" for gui/<uid> while
|
||||
// user/<uid> prints normally. Verified on a headless macOS service account.
|
||||
it("reaches the user domain when the account has no gui session", async () => {
|
||||
execLaunchctl
|
||||
.mockResolvedValueOnce({
|
||||
code: 113,
|
||||
stdout: "",
|
||||
stderr: "Could not find service",
|
||||
termination: "exit",
|
||||
})
|
||||
.mockResolvedValueOnce({
|
||||
code: 125,
|
||||
stdout: "",
|
||||
stderr: "Domain does not support specified action",
|
||||
termination: "exit",
|
||||
})
|
||||
.mockResolvedValueOnce(stopping("\texit timeout = 30\n\tpid = 4242\n"));
|
||||
await expect(readLaunchdStopTimeout(LAUNCHD_ENV)).resolves.toEqual({
|
||||
stop: { timeoutMs: 30_000, source: "launchd user/501/ai.openclaw.gateway exit timeout" },
|
||||
});
|
||||
expect(execLaunchctl).toHaveBeenCalledTimes(3);
|
||||
});
|
||||
|
||||
it("refuses a same-named job in every other domain and says why", async () => {
|
||||
execLaunchctl.mockResolvedValue(stopping("\texit timeout = 300\n\tpid = 99\n"));
|
||||
const read = await readLaunchdStopTimeout(LAUNCHD_ENV);
|
||||
// Nothing was established, so no deadline is reported at all and the caller
|
||||
// keeps the platform-neutral policy it already resolved.
|
||||
expect(read.stop).toBeNull();
|
||||
for (const target of [
|
||||
"system/ai.openclaw.gateway",
|
||||
"gui/501/ai.openclaw.gateway",
|
||||
"user/501/ai.openclaw.gateway",
|
||||
]) {
|
||||
expect(read.warning).toContain(`${target}: pid 99 is neither this process nor its launcher`);
|
||||
}
|
||||
});
|
||||
|
||||
// The installed service can keep a launcher parent while the serving Gateway
|
||||
// runs as its child, so the job prints the launcher's pid. Requiring
|
||||
// pid === process.pid there would reject the job that enforces the deadline.
|
||||
it("accepts the job when it is this process's launcher parent", async () => {
|
||||
execLaunchctl.mockResolvedValue(stopping("\texit timeout = 12\n\tpid = 4241\n"));
|
||||
await expect(readLaunchdStopTimeout(LAUNCHD_ENV)).resolves.toEqual({
|
||||
stop: { timeoutMs: 12_000, source: "launchd system/ai.openclaw.gateway exit timeout" },
|
||||
});
|
||||
});
|
||||
|
||||
// The launcher declares nothing, so its reap timer is derived from the expression
|
||||
// the launcher itself arms from. A longer operator deadline cannot be spent under
|
||||
// it: the parent force-kills this process first. Deriving rather than being told is
|
||||
// what makes this hold when an already-running older launcher started this Gateway,
|
||||
// which is the only shape an upgrade can take.
|
||||
//
|
||||
// Every marker is covered because all three respawn call sites reach the same
|
||||
// launcher through the same function and arm the same timer. Gating on the
|
||||
// Node-recovery marker alone left the two compile-cache respawns uncapped, and the
|
||||
// packaged one can wrap a foreground `gateway run` on an installed service.
|
||||
it.each(RESPAWN_MARKERS)(
|
||||
"caps the job deadline at the launcher's derived reap timer for %s",
|
||||
async (marker) => {
|
||||
execLaunchctl.mockResolvedValue(stopping("\texit timeout = 55\n\tpid = 4241\n"));
|
||||
await expect(readLaunchdStopTimeout({ ...SERVICE_ENV, [marker]: "1" })).resolves.toEqual({
|
||||
stop: {
|
||||
timeoutMs: 19_000,
|
||||
source:
|
||||
"launchd system/ai.openclaw.gateway exit timeout capped at the launcher's 19000ms stop timer",
|
||||
},
|
||||
});
|
||||
},
|
||||
);
|
||||
|
||||
// Holding the parent slot is not evidence of a reap timer. An operator wrapper can
|
||||
// keep the job's pid and start the Gateway itself while running none, and capping
|
||||
// its deadline would cut a drain nothing was going to interrupt.
|
||||
it("leaves a parent that did not respawn this process capping nothing", async () => {
|
||||
execLaunchctl.mockResolvedValue(stopping("\texit timeout = 55\n\tpid = 4241\n"));
|
||||
await expect(readLaunchdStopTimeout(SERVICE_ENV)).resolves.toEqual({
|
||||
stop: { timeoutMs: 55_000, source: "launchd system/ai.openclaw.gateway exit timeout" },
|
||||
});
|
||||
});
|
||||
|
||||
// Outside the service layout the launcher bounds its child by the platform-neutral
|
||||
// stop policy, which outlasts any job deadline launchd will enforce, so nothing is
|
||||
// cut even though the marker says a launcher is there.
|
||||
it("derives the policy deadline when the respawning job is not the service", async () => {
|
||||
execLaunchctl.mockResolvedValue(stopping("\texit timeout = 55\n\tpid = 4241\n"));
|
||||
await expect(
|
||||
readLaunchdStopTimeout({ ...LAUNCHD_ENV, OPENCLAW_NODE_UPDATE_RESPAWNED: "1" }),
|
||||
).resolves.toEqual({
|
||||
stop: { timeoutMs: 55_000, source: "launchd system/ai.openclaw.gateway exit timeout" },
|
||||
});
|
||||
});
|
||||
|
||||
// launchd reaps the whole job first, so a shorter job deadline still wins.
|
||||
it("keeps a job deadline shorter than the launcher's reap timer", async () => {
|
||||
execLaunchctl.mockResolvedValue(stopping("\texit timeout = 5\n\tpid = 4241\n"));
|
||||
await expect(readLaunchdStopTimeout(RESPAWNED_SERVICE_ENV)).resolves.toEqual({
|
||||
stop: { timeoutMs: 5_000, source: "launchd system/ai.openclaw.gateway exit timeout" },
|
||||
});
|
||||
});
|
||||
|
||||
// The launcher timer belongs to a parent, so a Gateway that is the job itself is
|
||||
// never shortened by one even while carrying an inherited respawn marker.
|
||||
it("ignores the launcher timer when this process is the job", async () => {
|
||||
execLaunchctl.mockResolvedValue(stopping("\texit timeout = 55\n\tpid = 4242\n"));
|
||||
await expect(readLaunchdStopTimeout(RESPAWNED_SERVICE_ENV)).resolves.toEqual({
|
||||
stop: { timeoutMs: 55_000, source: "launchd system/ai.openclaw.gateway exit timeout" },
|
||||
});
|
||||
});
|
||||
|
||||
// launchd is stopping the job, so a deadline is running even though its value
|
||||
// is unreadable. Guess short here: guessing long is what lets the drain die.
|
||||
it.each(["\tpid = 4242\n", "\texit timeout = not-a-number\n\tpid = 4242\n"])(
|
||||
"uses the conservative default when a running stop has no readable deadline",
|
||||
async (fields) => {
|
||||
execLaunchctl.mockResolvedValue(stopping(fields));
|
||||
const read = await readLaunchdStopTimeout(LAUNCHD_ENV);
|
||||
// A deadline IS running here, so this one is a real native budget even
|
||||
// though its value had to be defaulted.
|
||||
expect(read.stop?.timeoutMs).toBe(20_000);
|
||||
expect(read.stop?.source).toBe(
|
||||
"launchd system/ai.openclaw.gateway exit timeout unavailable; default ExitTimeOut",
|
||||
);
|
||||
expect(read.warning).toContain(
|
||||
"launchd is stopping system/ai.openclaw.gateway but its exit timeout is missing or invalid",
|
||||
);
|
||||
},
|
||||
);
|
||||
|
||||
// THE REGRESSION GUARD for the failed-inspection path. A probe that established
|
||||
// nothing must report no deadline: handing back the Gateway's own stop policy
|
||||
// here is what let the caller classify an unverified number as a native stop
|
||||
// budget, capping a longer requested restart drain and arming a forced exit.
|
||||
it("reports no deadline when the job cannot be inspected", async () => {
|
||||
execLaunchctl.mockResolvedValue({
|
||||
code: 1,
|
||||
stdout: "",
|
||||
stderr: "permission denied",
|
||||
termination: "exit",
|
||||
});
|
||||
const read = await readLaunchdStopTimeout(LAUNCHD_ENV);
|
||||
expect(read.stop).toBeNull();
|
||||
expect(read.warning).toContain(
|
||||
"system/ai.openclaw.gateway: launchctl print exited 1: permission denied",
|
||||
);
|
||||
expect(read.warning).toContain("keeping the Gateway stop policy");
|
||||
expect(read.warning).toContain("Check the running job with launchctl print.");
|
||||
});
|
||||
|
||||
it("survives launchctl throwing rather than exiting nonzero", async () => {
|
||||
execLaunchctl.mockRejectedValue(new Error("spawn ENOENT"));
|
||||
const read = await readLaunchdStopTimeout(LAUNCHD_ENV);
|
||||
expect(read.stop).toBeNull();
|
||||
expect(read.warning).toContain("launchctl print threw");
|
||||
});
|
||||
|
||||
it("honours an explicit label override", async () => {
|
||||
execLaunchctl.mockResolvedValue(stopping("\texit timeout = 45\n\tpid = 4242\n"));
|
||||
await expect(
|
||||
readLaunchdStopTimeout({ ...LAUNCHD_ENV, OPENCLAW_LAUNCHD_LABEL: "com.example.gw" }),
|
||||
).resolves.toEqual({
|
||||
stop: { timeoutMs: 45_000, source: "launchd system/com.example.gw exit timeout" },
|
||||
});
|
||||
});
|
||||
|
||||
it("warns instead of throwing when the configured label is invalid", async () => {
|
||||
const read = await readLaunchdStopTimeout({
|
||||
...LAUNCHD_ENV,
|
||||
OPENCLAW_LAUNCHD_LABEL: "bad label/../etc",
|
||||
});
|
||||
expect(read.stop).toBeNull();
|
||||
expect(read.warning).toContain("label could not be resolved");
|
||||
expect(execLaunchctl).not.toHaveBeenCalled();
|
||||
});
|
||||
|
||||
it("stays out of the way when this process is not a launchd job", async () => {
|
||||
await expect(readLaunchdStopTimeout({})).resolves.toEqual({ stop: null });
|
||||
expect(execLaunchctl).not.toHaveBeenCalled();
|
||||
});
|
||||
});
|
||||
|
||||
// Ties the names this file caps on back to the implementation's own list. Without
|
||||
// this, dropping a marker from that list would leave the cap cases above still
|
||||
// passing against the two that remained.
|
||||
describe("launcher respawn markers", () => {
|
||||
it.each(RESPAWN_MARKERS)("recognises %s as a respawning launcher", (marker) => {
|
||||
expect(isRespawnedByLauncher({ [marker]: "1" })).toBe(true);
|
||||
});
|
||||
|
||||
it("recognises nothing else", () => {
|
||||
expect(isRespawnedByLauncher({})).toBe(false);
|
||||
expect(isRespawnedByLauncher({ OPENCLAW_SUPERVISOR_MODE: "external" })).toBe(false);
|
||||
// Only the set value counts, so an emptied or disabled marker caps nothing.
|
||||
expect(isRespawnedByLauncher({ OPENCLAW_NODE_UPDATE_RESPAWNED: "0" })).toBe(false);
|
||||
expect(isRespawnedByLauncher({ OPENCLAW_NODE_UPDATE_RESPAWNED: "" })).toBe(false);
|
||||
});
|
||||
});
|
||||
187
src/infra/launchd-stop-timeout.ts
Normal file
187
src/infra/launchd-stop-timeout.ts
Normal file
|
|
@ -0,0 +1,187 @@
|
|||
import { parseStrictPositiveInteger } from "@openclaw/normalization-core/number-coercion";
|
||||
import { truncateUtf16Safe } from "@openclaw/normalization-core/utf16-slice";
|
||||
import { isForegroundGatewayRunArgv } from "../cli/gateway-run-argv.js";
|
||||
import { execLaunchctl, formatLaunchctlResultDetail } from "../daemon/launchd-exec.js";
|
||||
import { resolveLaunchAgentLabel } from "../daemon/launchd-label.js";
|
||||
import { parseKeyValueOutput } from "../daemon/runtime-parse.js";
|
||||
import { formatErrorMessage } from "./errors.js";
|
||||
import {
|
||||
isRespawnedByLauncher,
|
||||
LAUNCH_AGENT_EXIT_TIMEOUT_SECONDS,
|
||||
resolveLauncherStopTimeoutMs,
|
||||
} from "./gateway-shutdown-budget.js";
|
||||
import { detectRespawnSupervisor } from "./supervisor-markers.js";
|
||||
|
||||
type LaunchdStopTimeout = { timeoutMs: number; source: string };
|
||||
|
||||
// A warning without a deadline means inspection was inconclusive. Only a
|
||||
// confirmed launchd stop or our launcher's independent reap timer yields a
|
||||
// native budget; the caller owns the fallback policy.
|
||||
export type LaunchdStopRead = { stop: LaunchdStopTimeout | null; warning?: string };
|
||||
|
||||
const LAUNCHCTL_PRINT_TIMEOUT_MS = 2_000;
|
||||
|
||||
// Check all domains: a service account may lack a GUI session, and a matching
|
||||
// label in a different domain must not supply this process's deadline.
|
||||
function resolveLaunchdDomains(label: string): string[] {
|
||||
const uid = typeof process.getuid === "function" ? process.getuid() : 501;
|
||||
// A LaunchDaemon and a LaunchAgent can carry the same label in different
|
||||
// domains, so each is a candidate and the pid decides which one is ours.
|
||||
// A service account with no logged-in session has no gui domain at all, and
|
||||
// its per-user jobs live in `user/<uid>`, so both user domains are checked.
|
||||
return [`system/${label}`, `gui/${uid}/${label}`, `user/${uid}/${label}`];
|
||||
}
|
||||
|
||||
// launchctl can report the Gateway or a launcher parent that holds the job PID.
|
||||
function resolveJobRelation(pid: number | undefined): "self" | "launcher" | null {
|
||||
if (pid === undefined) {
|
||||
return null;
|
||||
}
|
||||
if (pid === process.pid) {
|
||||
return "self";
|
||||
}
|
||||
return pid === process.ppid ? "launcher" : null;
|
||||
}
|
||||
|
||||
// Nested coalition blocks also contain state: use the single-tab job field,
|
||||
// not the last state selected by the generic key-value parser.
|
||||
function readJobState(printed: string): string | undefined {
|
||||
return /^\tstate = (?<state>.+)$/mu.exec(printed)?.groups?.state?.trim();
|
||||
}
|
||||
|
||||
// launchd reports SIGTERMed during bootout/kickstart, but keeps running on an
|
||||
// externally delivered SIGTERM. Only the former starts ExitTimeOut's clock.
|
||||
function isLaunchdStoppingJob(state: string | undefined): boolean {
|
||||
return state !== undefined && /^SIG[A-Z0-9]+ed$/u.test(state);
|
||||
}
|
||||
|
||||
// A respawn marker, not parenthood alone, proves the launcher runs a reap timer.
|
||||
// Derive that timer from the same shared expression: an already-running parent
|
||||
// cannot be taught a new announced deadline by upgrading the child.
|
||||
function resolveParentLauncherStopTimeoutMs(env: NodeJS.ProcessEnv): number | undefined {
|
||||
// One of these is set on every child the launcher respawns, and all of them predate
|
||||
// this deadline being derived, so a marker is present for a Gateway that launcher
|
||||
// started and absent for any other parent. An operator wrapper that keeps the job's
|
||||
// pid and starts the Gateway itself runs no such timer, and capping its deadline
|
||||
// would cut a drain nothing was going to interrupt.
|
||||
if (!isRespawnedByLauncher(env)) {
|
||||
return undefined;
|
||||
}
|
||||
return resolveLauncherStopTimeoutMs({
|
||||
env,
|
||||
platform: process.platform,
|
||||
// The launcher branched on its own argv, and it respawns the child with the same
|
||||
// user arguments, so testing ours reproduces the branch it took.
|
||||
foreground: isForegroundGatewayRunArgv(process.argv),
|
||||
});
|
||||
}
|
||||
|
||||
// The job is stopping but its value is missing; use launchd's 20s default.
|
||||
function defaultStopDeadline(target: string, reason: string): LaunchdStopRead {
|
||||
const timeoutMs = LAUNCH_AGENT_EXIT_TIMEOUT_SECONDS * 1_000;
|
||||
return {
|
||||
stop: { timeoutMs, source: `launchd ${target} exit timeout unavailable; default ExitTimeOut` },
|
||||
warning: `launchd is stopping ${target} but ${reason}; using ${timeoutMs}ms default. Check the running job with launchctl print.`,
|
||||
};
|
||||
}
|
||||
|
||||
// Unknown ownership of the stop cannot justify shortening a potentially
|
||||
// unconstrained drain. Warn, but leave the caller's policy authoritative.
|
||||
function unresolved(failures: string[]): LaunchdStopRead {
|
||||
return {
|
||||
stop: null,
|
||||
// Each failure already names the target it came from, and the label-resolution
|
||||
// case has no label to name, so the prefix deliberately carries neither.
|
||||
warning: `Unable to inspect the launchd job; ${failures
|
||||
.map((failure) => truncateUtf16Safe(failure.replaceAll(/\s+/g, " "), 500))
|
||||
.join("; ")}; keeping the Gateway stop policy. Check the running job with launchctl print.`,
|
||||
};
|
||||
}
|
||||
|
||||
// Read the loaded job only on shutdown. Its ExitTimeOut applies only while
|
||||
// launchd stops it; a respawn launcher can enforce an independent deadline.
|
||||
export async function readLaunchdStopTimeout(
|
||||
env: NodeJS.ProcessEnv = process.env,
|
||||
): Promise<LaunchdStopRead> {
|
||||
if (detectRespawnSupervisor(env, "darwin") !== "launchd") {
|
||||
return { stop: null };
|
||||
}
|
||||
const failures: string[] = [];
|
||||
let label: string;
|
||||
try {
|
||||
label = resolveLaunchAgentLabel(env);
|
||||
} catch (error: unknown) {
|
||||
return unresolved([`label could not be resolved: ${formatErrorMessage(error)}`]);
|
||||
}
|
||||
for (const target of resolveLaunchdDomains(label)) {
|
||||
const failed = (reason: string) => failures.push(`${target}: ${reason}`);
|
||||
const result = await execLaunchctl(["print", target], LAUNCHCTL_PRINT_TIMEOUT_MS).catch(
|
||||
(error: unknown) => {
|
||||
failed(`launchctl print threw: ${formatErrorMessage(error)}`);
|
||||
return undefined;
|
||||
},
|
||||
);
|
||||
if (!result) {
|
||||
continue;
|
||||
}
|
||||
if (result.code !== 0) {
|
||||
failed(`launchctl print exited ${result.code}: ${formatLaunchctlResultDetail(result)}`);
|
||||
continue;
|
||||
}
|
||||
const printed = result.stdout || result.stderr || "";
|
||||
const entries = parseKeyValueOutput(printed, "=");
|
||||
// Adopting a deadline from a same-named job in the other domain would be
|
||||
// worse than the fallback, so the printed job must be ours.
|
||||
const pid = parseStrictPositiveInteger(entries.pid ?? "");
|
||||
const relation = resolveJobRelation(pid);
|
||||
if (!relation) {
|
||||
failed(`pid ${pid ?? "missing"} is neither this process nor its launcher`);
|
||||
continue;
|
||||
}
|
||||
// This is our job, so stop searching. Whether its deadline binds this stop is
|
||||
// a separate question from whether the job was found.
|
||||
const launcherMs =
|
||||
relation === "launcher" ? resolveParentLauncherStopTimeoutMs(env) : undefined;
|
||||
const state = readJobState(printed);
|
||||
if (!isLaunchdStoppingJob(state)) {
|
||||
// A direct signal to the Gateway starts neither launchd's clock nor its
|
||||
// parent launcher's timer. Node cannot identify the sender, so capping on
|
||||
// the parent marker would truncate that supported long drain. A signal
|
||||
// forwarded by the launcher is indistinguishable here (existing limitation).
|
||||
return !state && entries["exit timeout"] !== undefined
|
||||
? {
|
||||
stop: null,
|
||||
warning: `launchd ${target} printed an exit timeout but no job state, so it is treated as not stopping and the Gateway stop policy is kept. Check the running job with launchctl print.`,
|
||||
}
|
||||
: { stop: null };
|
||||
}
|
||||
const rawSeconds = entries["exit timeout"]?.trim();
|
||||
if (rawSeconds === "0") {
|
||||
return launcherMs !== undefined
|
||||
? {
|
||||
stop: {
|
||||
timeoutMs: launcherMs,
|
||||
source: `launchd ${target} unlimited exit timeout capped at the launcher's ${launcherMs}ms stop timer`,
|
||||
},
|
||||
}
|
||||
: { stop: { timeoutMs: Infinity, source: `launchd ${target} unlimited exit timeout` } };
|
||||
}
|
||||
const seconds = parseStrictPositiveInteger(rawSeconds ?? "");
|
||||
if (seconds === undefined) {
|
||||
return defaultStopDeadline(target, "its exit timeout is missing or invalid");
|
||||
}
|
||||
const jobMs = seconds * 1_000;
|
||||
// A parent that reaps this process on its own timer binds before the job's
|
||||
// ExitTimeOut, and spending the longer deadline would only get the drain
|
||||
// force-killed.
|
||||
return launcherMs !== undefined && launcherMs < jobMs
|
||||
? {
|
||||
stop: {
|
||||
timeoutMs: launcherMs,
|
||||
source: `launchd ${target} exit timeout capped at the launcher's ${launcherMs}ms stop timer`,
|
||||
},
|
||||
}
|
||||
: { stop: { timeoutMs: jobMs, source: `launchd ${target} exit timeout` } };
|
||||
}
|
||||
return unresolved(failures);
|
||||
}
|
||||
|
|
@ -18,6 +18,9 @@ beforeEach(() => {
|
|||
vi.useFakeTimers();
|
||||
child = new ChildProcess();
|
||||
kill = vi.spyOn(child, "kill").mockReturnValue(true);
|
||||
// `spawn` is hoisted once for the file, so its call log survives across cases
|
||||
// and `toHaveBeenCalledExactlyOnceWith` would only ever hold for the first one.
|
||||
spawn.mockClear();
|
||||
spawn.mockReturnValue(child);
|
||||
exit = vi.spyOn(process, "exit").mockImplementation(vi.fn<typeof process.exit>());
|
||||
vi.spyOn(process, "kill").mockReturnValue(true);
|
||||
|
|
@ -50,6 +53,21 @@ it.each([
|
|||
XPC_SERVICE_NAME: "ai.openclaw.fixture",
|
||||
});
|
||||
detach = () => child.emit("exit", 0, null);
|
||||
// The launcher tells the child nothing about the timer it armed: the serving
|
||||
// Gateway derives the same deadline from the same shared expression, which is what
|
||||
// lets a Gateway started by an already-running older launcher bound itself
|
||||
// correctly. So the env must reach the child unchanged, and the escalation
|
||||
// asserted below is what that derivation has to land on.
|
||||
expect(spawn).toHaveBeenCalledExactlyOnceWith(
|
||||
"node",
|
||||
["child.mjs"],
|
||||
expect.objectContaining({
|
||||
env: {
|
||||
OPENCLAW_LAUNCHD_LABEL: "ai.openclaw.fixture",
|
||||
XPC_SERVICE_NAME: "ai.openclaw.fixture",
|
||||
},
|
||||
}),
|
||||
);
|
||||
const signal = process.listeners("SIGTERM").find((listener) => !previous.has(listener));
|
||||
expect(signal).toBeDefined();
|
||||
signal!("SIGTERM");
|
||||
|
|
|
|||
|
|
@ -1291,6 +1291,7 @@ describe("frozen admission workflow barriers", () => {
|
|||
for (const [index, child] of children.entries()) {
|
||||
const toolingPaths = child.sources.tooling.map(({ path }) => path);
|
||||
const selectedPaths = child.sources.selected.map(({ path }) => path);
|
||||
expect(toolingPaths).toContain("scripts/lib/record-shared.mjs");
|
||||
if (index === 1) {
|
||||
expect(toolingPaths).toContain(extra);
|
||||
expect(selectedPaths).toContain("extensions/codex/package.json");
|
||||
|
|
|
|||
|
|
@ -92,7 +92,6 @@ function fixture(
|
|||
recursive: true,
|
||||
});
|
||||
for (const file of [
|
||||
"record-shared.mjs",
|
||||
"update-compat-contract.mjs",
|
||||
"openclaw-e2e-instance.sh",
|
||||
"docker-e2e-watchdog.mjs",
|
||||
|
|
@ -554,6 +553,7 @@ describe("frozen admission bootstrap repairs", () => {
|
|||
it.each([
|
||||
reader,
|
||||
"scripts/lib/docker-e2e-scenarios.mts",
|
||||
"scripts/lib/record-shared.mjs",
|
||||
shell,
|
||||
"scripts/lib/trusted-native-typescript.mjs",
|
||||
"scripts/lib/native-typescript.mts",
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue