fix(gateway): derive the darwin stop budget from the launchd job (#157007)

* fix(gateway): derive the darwin stop budget from the launchd job

resolveGatewayShutdownBudget derives the stop deadline from restart
ownership, so a darwin host running with OPENCLAW_SUPERVISOR_MODE=external
resolves drain=315000ms whatever launchd actually enforces.

The linux-gated systemd probe becomes a platform dispatch, and a new
readLaunchdStopTimeout mirrors readSystemdStopTimeout: it reads the running
job's effective exit timeout and accepts it only when the printed pid is
this process. run-loop.ts is untouched because the reader detects launchd
from the environment itself.

Closes #156968

* fix(gateway): accept the launchd launcher parent and cap its deadline

The darwin reader accepted the printed job only when its pid was this
process. The installed service can keep a launcher parent while the serving
Gateway runs as its child, so the job prints the launcher's pid, the
enforcing job was rejected, and the Gateway fell back to 20 seconds in the
one layout where an operator's ExitTimeOut was meant to apply.

Accepting that job at face value would be worse than rejecting it.
node-runtime-recovery.mjs builds the launcher's own reap timer from the
compile-time LAUNCH_AGENT_EXIT_TIMEOUT_SECONDS rather than from the job, so
on a forwarded stop it re-sends SIGTERM to its child at 18000ms and SIGKILLs
it at 19000ms. Adopting a 90 second job deadline in that layout would plan a
75000ms drain and then lose it to its own parent, which is the truncated
drain this change exists to prevent. The reader now accepts process.pid or
process.ppid and takes min(jobExitTimeout, 20000) in the launcher case,
naming the cap in the source string only when it binds. A shorter job
deadline still applies in both layouts, because launchd reaps the whole job
regardless of the launcher.

The docs claimed the deadline is read at startup and when accepting
shutdown. run-loop.ts gates the budget refresh on linux at both consumption
sites, so on darwin it is read once, at startup, and nativeStopBudget has no
shutdown-time consumer there. The page now says startup-only and drops the
sentence about a failed shutdown reread retaining the startup budget, which
described linux behaviour only.

* fix(gateway): scope the darwin stop budget to launchd-driven stops

Address the two review findings on the previous revision.

The launcher parent is now identified by an explicit
OPENCLAW_LAUNCHER_STOP_TIMEOUT_MS declaration rather than a bare ppid
match, so an unknown parent keeps the job's full deadline instead of
being capped on a guess.

The darwin reader inspects the launchd job only while a stop is under
way, and adopts that job's deadline only when the job reports the
SIGTERMed state. An external SIGTERM keeps the platform-neutral drain,
because launchd's ExitTimeOut bounds only a stop launchd itself is
running.

The job state is read with an anchored single-tab pattern. launchctl
print emits further "state = active" lines inside the resource and
jetsam coalition blocks, and the shared key/value parser keeps the last
occurrence, so the shared parser could never observe SIGTERMed. A test
fixture reproduces the nested coalition shape.

* fix(gateway): keep a failed launchd inspection off the native stop budget

`readLaunchdStopTimeout` answered a failed job inspection with the Gateway's own
330000ms stop policy wrapped in a non-null result. `resolveGatewayShutdownBudget`
classifies every non-null result as a native stop budget, so a probe that
established nothing still set `nativeStopBudget`. On an externally supervised
darwin Gateway that caps a longer requested restart drain and can arm a forced
exit against a deadline no supervisor was confirmed to be enforcing.

The reader now reports its two answers separately. `stop` carries a deadline only
while launchd is confirmed to be stopping the job; `warning` reaches the operator
either way. A failed inspection reports `stop: null` and leaves the caller on the
platform-neutral policy it had already resolved, which is the same number as
before, now classified honestly. The reader no longer imports
GATEWAY_SERVICE_STOP_TIMEOUT_MS, because choosing the Gateway's policy number was
never its job.

Tests assert the flag and the downstream restart drain in the failure case, paired
with a confirmed-deadline case so deleting the read outright cannot pass both.

Also fixes TS2532 on the job-state regex: the named group is optional under
`noUncheckedIndexedAccess`, so `.trim()` needed the optional chain.

* test(infra): isolate the hoisted spawn mock and ratchet the OPENCLAW_* count

`spawn` is hoisted once per file, so its call log survived across cases and
`toHaveBeenCalledExactlyOnceWith` could only hold for the first one. Clear it in
`beforeEach`.

CI reports `OPENCLAW_* count 488 exceeds budget 487; update
config/env-var-count-budget.txt` as a warning. The new name is the launcher's
OPENCLAW_LAUNCHER_STOP_TIMEOUT_MS declaration, so the count is correct and the
ratchet moves to 488 with the reason recorded in the file.

* style(infra): apply oxfmt to the new launchd reader assertions

Run by the repo's own `pnpm run format`; joins one over-wrapped assertion line.
`pnpm run format:check` then reports "All matched files use the correct format."

* test(cli): mock the launchd reader and split the darwin run-loop cases out

Activating the darwin stop-budget refresh made `run-loop.test.ts` exercise the real
`readLaunchdStopTimeout`, which spawns three `launchctl print` calls against live
subprocess I/O while the suite holds `vi.useFakeTimers()`. One case ran 120039ms
before failing and the abandoned loop leaked into later cases, for eight failures
in total. All eight clear by mocking the reader; no existing expectation was wrong.

The mock could not simply be added. `run-loop.test.ts` is 3813 lines against a 1000
line cap where `check-line-cap-ratchet` rejects any growth, and an earlier attempt
failed with `3501 -> 3510 counted lines`. Taking the remedy that ratchet names, the
darwin cases move to `run-loop.launchd.test.ts` and the original file shrinks by 82
lines. It keeps the mock too, because four of those cases are registered into it by
shared `register*` helpers whose `it.each` arrays interleave systemd and launchd
variants and cannot be split without touching those support files.

The new file adds one case the default mock would otherwise hide: a job reporting a
30000ms deadline bounds the stop at 25000ms and arms the force exit there.

`node-runtime-recovery.test.ts` asserts the replacement child's spawn env exactly,
so it gains the launcher's own OPENCLAW_LAUNCHER_STOP_TIMEOUT_MS declaration rather
than a loosened matcher. That argv is a non-foreground `doctor --fix`, so the value
is the 1s signal exit grace plus the 1s force-kill grace.

* fix(gateway): only retain a startup budget when the probe was inconclusive

Extending the stop-budget refresh to darwin gave the retained-budget safety net a
new and wrong trigger. On a launchd-owned Gateway the startup budget is already
native, and an in-process restart signals the process without launchd running the
stop, so the job still prints `running` and the reader reports `stop: null`. The
old condition treated any null as unconfirmed, so every in-process restart of a
default macOS install logged "Retaining the startup shutdown budget of 15000ms
because the current supervisor stop timeout could not be confirmed." The timeout
was confirmed; it was confirmed not to apply.

`readNativeStopTimeout` now reports `inconclusive` alongside the deadline, and only
an inconclusive probe retains. Linux keeps its exact previous condition, an absent
unit or a warned read. The budget number was already unaffected, so this is the
message and the wasted `launchctl print` calls, not the deadline.

Two cases pin it: a confirmed not-stopping read warns nothing, and a failed
inspection still retains and still says so.

Also fixes two wording defects found in review. `unresolved()` no longer takes a
label it only ever interpolated as "the launchd job the configured label"; each
failure already names its own target. And the `readJobState` comment said `state`
appears four times in `launchctl print` output when its own next clause counts
three, which is what a live LaunchDaemon prints.

* style(cli): apply oxfmt to the new retained-budget assertions

Run by the repo's own `pnpm run format`; joins one over-wrapped assertion line.
Whitespace only, in a test file, so it cannot change runtime behaviour.

* fix(gateway): a defaulted deadline is not an inconclusive probe

`inconclusive` keyed off the warning alone, so the defaulted-value case counted as a
failed probe. There launchd is confirmed to be stopping the job and only its
deadline had to be guessed, so a clock is genuinely running and there is nothing to
retain. Under launchd ownership that combination logged both "using 20000ms
default" and "could not be confirmed" for the same stop, the second being false.

Now only a warning with no deadline at all counts as inconclusive. A third case
pins it, asserting the defaulted read warns exactly once and about the deadline
rather than about confirmation.

Also corrects a test fixture comment that still said `launchctl print` emits `state`
four times; the arrangement it documents, and a live LaunchDaemon, both show three.

* test(cli): correct the regression guard's account of the reported drain

The comment said active work "had time to finish" during the reported 315 second
drain. The reporting host's log shows the opposite: that drain hit its own timeout
with five tasks still active, and the Sep 20 stop with six. The drain was being
used, which is a stronger reason for the guard, not a weaker one.

* fix(infra): stop exporting a type nothing outside the module reads

Knip's all-exports gate reads LaunchdStopTimeout as dead: the only
public surface is readLaunchdStopTimeout's LaunchdStopRead return, and
nothing imports the nested shape by name. Keep it module-local.

* fix(gateway): keep a proportional drain and cap an undeclared launcher

The fixed 10s reserve and 5s exit margin were sized against the 315s policy
drain. Subtracting them outright from a launchd job's own ExitTimeOut spent a
short deadline entirely on overhead: 5 seconds funded the margin alone and 15
funded margin plus reserve, so active work drained for zero milliseconds in
both. Each allowance now takes at most a share of what it is carved from, so a
deadline long enough to fund them is unchanged and a shorter one keeps a
proportional drain. Drain is now positive for every positive deadline and never
decreases as ExitTimeOut grows.

The shutdown log also subtracted the unscaled reserve constant when reporting
drain, which understated a short budget by the amount the scaled reserve gives
back, so it now reports the margin and drain actually spent.

An already-running launcher published before OPENCLAW_LAUNCHER_STOP_TIMEOUT_MS
existed arms the same reap timer and declares nothing, which is the upgrade
shape: replacing files cannot change a launcher that is already running. Reading
that silence as "no deadline" let a long custom ExitTimeOut be budgeted past a
force-kill the parent was already counting down. The launcher's arithmetic moves
into the shared budget module so the serving Gateway reconstructs the same timer,
used only where the printed job carries OpenClaw's own label and its pid is this
process's immediate parent. In that position the parent is the process launchd
started for OpenClaw's job and the recovery launcher is the only path that puts a
Gateway underneath it, so an unrelated process manager holding the parent slot
still caps nothing.

* fix(infra): gate the launcher reconstruction on the recovery respawn marker

Holding the launchd job's pid is not evidence of a reap timer. An operator
wrapper can keep that pid and start the Gateway itself while running none, so
reconstructing a cap from the parent relation alone would cut a drain nothing was
going to interrupt. The recovery launcher has stamped OPENCLAW_NODE_UPDATE_RESPAWNED
on every child it respawns since long before it declared a stop timer, so that
marker is present in exactly the upgrade case and absent for any other parent.

* fix(infra): declare the new shared budget helpers for the type checker

The root module is typed by a hand-written declaration file rather than being
compiled, so new exports are invisible to src without being declared there.

* docs(gateway): scope the positive-drain claim to the allocation

The elapsed cost of resolving the budget is debited after the allocation, so a
deadline shorter than that cost can still leave nothing to spend. Claiming every
positive deadline yields a positive drain overstated it.

* test(cli): accept the retained-budget attribution on a synthetic launchd label

The fixture declares a launchd label with no matching job, so the per-stop
re-inspection cannot read one and the startup budget is retained. The deadline is
unchanged at 15000ms; only the source attribution differs. The assertion still
fails if the budget comes from the platform-neutral policy instead.

* fix(gateway): derive the launcher's reap deadline instead of declaring it

Measurement showed the declared value and the derived value are the same number:
a published launcher and a candidate launcher both arm an 18000ms exit grace, and
a candidate Gateway resolved an identical 19000ms cap and 14231ms budget under
each. The declaration was therefore carrying no information the child could not
compute, while adding an OPENCLAW_* name and leaving the upgrade path conditional
on which build of the launcher happened to be running.

Both sides now read one expression in gateway-shutdown-budget.mjs, the launcher to
arm its escalation and the Gateway to bound its budget, so they cannot disagree and
there is no published-versus-candidate launcher distinction left to prove. The cap
stays gated on OPENCLAW_NODE_UPDATE_RESPAWNED so a parent that did not respawn this
process still caps nothing.

Drops the env-var count back to 486, so the name ratchet no longer needs an owner
waiver, and takes config/env-var-count-budget.txt out of the diff entirely.

Test deadlines move off 90 seconds onto 55, because launchd clamps ExitTimeOut at
60 and a 90-second case cannot occur on macOS 27.

* test(infra): restore the recovery spawn-env cases to their base form

The launcher no longer declares a deadline, so the assertions these cases grew have
nothing left to check and their explanatory comment described behavior that is gone.

* fix(gateway): cap every respawning launcher, not just the Node-recovery one

All three runRespawnedChild call sites reach the same launcher through the same
function and arm the same escalation, but only the Node-recovery one sets
OPENCLAW_NODE_UPDATE_RESPAWNED. Gating on that marker alone left the two
compile-cache respawns uncapped, and the packaged one can wrap a foreground
gateway run on an installed service: the job's deadline would then be budgeted past
a force-kill the parent had already armed, which is the hazard this cap exists to
prevent.

The marker set moves next to the deadline it authorises, so adding a respawn path
cannot silently escape the cap. Keeping the literals in the root module also keeps
them out of the env-count ratchet's scope: reading the marker from counted
production source promoted a previously test-only name and held the count at 487
even after the declared-timer variable was removed.

Also applies oxfmt's own wrap to the launcher-relation expression.

* fix(infra): re-export the respawn marker set through the infra barrel

The constant existed only on the root module, so the test importing it from the
barrel got undefined and it.each(undefined) threw during collection: the suite
reported zero tests rather than a failure, and typecheck flagged the missing member.
Re-exporting the binding adds no OPENCLAW_* literal under src, so the env-var count
stays at 486.

* fix(gateway): address the pre-review findings on the launchd stop budget

Docs were stale or wrong in four places. The systemd section still promised a fixed
10 second reserve and 5 second margin, but the share caps are platform-neutral and
do change a Linux unit whose TimeoutStopSec is under 25 seconds; that radius is now
stated with the 20 second and 15 second cases spelled out. The direct-SIGTERM
paragraph claimed the platform-neutral drain "is the deadline that actually governs
that stop", which the branch's own kill -TERM capture contradicts: no supervisor
deadline governs that stop at all and the fallback is the template value. The
updated-launcher requirement no longer applies to the stop budget and says so. The
marker sentence named one marker when three are honoured, and one measured figure
was stale.

Two claims in the code were too strong. The marker list asserted every setter arms
the shared escalation, which is false for the compile-cache name: entry.compile-cache
sets it for a runner that reaps on a fixed short grace. That runner refuses a
foreground Gateway run off Windows and this deadline is only read during a darwin
stop, so it cannot be the parent here, and the comment now says that instead of
implying exclusivity. The sync note next to the escalation now records that the other
runner keeps its own copies of the graces.

An absent job state was indistinguishable from a job launchd is not stopping, so a
macOS that printed the block differently would silently revert to the platform-neutral
policy. When a deadline parsed but the state did not, that now warns and marks the read
inconclusive rather than passing as a positive answer.

Tests: launchd cases move off 90 and 315 second deadlines, which launchd clamps to 60
and cannot produce; the marker list is restated locally and tied back to the
implementation so dropping a marker fails; and the real-process case now asserts the
darwin probe actually ran, which it could not before.

Retires five exports with no production consumer that the dead-export scan flagged.

* test(infra): move the empty job state onto the missing-state warning

An empty value yields no state to recognise, so it reaches the same branch as a
missing line and now carries the same warning. Keeping it in the unrecognised-state
table asserted the absence of a field the reader deliberately reports.

* fix(infra): treat a whitespace-only job state as missing, and sharpen two docs claims

The state value is trimmed after matching, so a line carrying only whitespace yielded
an empty string rather than undefined and skipped the warning while behaving exactly
like a missing line.

The systemd section now says the exit margin shrinks below a 20-second deadline, not
just the reserve. The direct-signal paragraph distinguished only the launchd-supervised
budget; external supervisor mode keeps the platform-neutral policy and arms no
force-exit timer, and a signal delivered to the job pid does reach a resident recovery
launcher, which arms its own reap timer even though launchd is not stopping the job.

* docs(gateway): correct the loaded launchd job deadline guidance

launchctl print reports the job launchd has loaded, not the plist on disk,
so editing a loaded job's ExitTimeOut changes nothing until the job is
reloaded. Measured on macOS 27 against a scratch job: the plist moves from
20 to 47 while the loaded job keeps reporting 20, and it still reports 20
after launchctl kickstart -k, which restarts the process without reloading
the job. Only bootout then bootstrap makes it report 47, and that restarts
the Gateway with it.

State what reading per stop actually buys instead: the Gateway never plans
against a deadline it cached at its own startup.

* fix(gateway): keep the full cleanup reserve wherever a deadline funds it

The reserve was capped at half the shutdown budget whenever the budget could
not fund it outright. That kept a drain on short deadlines, but it also
reallocated deadlines that already worked: a 20 second ExitTimeOut, which is
both the shipped LaunchAgent template and launchd's own default, moved from a
10 second reserve and a 5 second drain to 7.5 seconds of each. Cleanup costing
between 7.5 and 10 seconds finished before and would have been cut off.

Bound the reserve by what keeping a drain actually requires instead. The floor
is 5 seconds, which is the drain the 20 second template already yields once the
fixed margin and reserve are subtracted, so funding the full reserve alongside
it takes 15 seconds of budget: exactly what that template resolves to. Every
deadline from 20 seconds up therefore keeps the allocation it had, and only a
deadline the fixed subtraction had already driven under a 5 second drain gives
any reserve up, never below half the budget.

Both allocations stay monotone in the deadline, and the 5 second case is
unchanged at 1875ms of each.

* fix(gateway): correct shutdown-budget probe-debit overclaims

- gateway-shutdown-budget.mjs: replace the floor-invariant claim with the
  real bound (reserve unchanged only at a >=15000ms post-margin budget,
  i.e. a >=20s deadline once the launchd probe's own cost, up to three
  launchctl print calls at 2000ms each, is subtracted); name the 8-20s
  sub-band reserve loss the old comment denied, including 16-19s where a
  flat subtraction already left a positive drain while still funding the
  reserve in full
- docs/gateway/restart-recovery.md: drop the systemd section's "down to
  the millisecond" promise and the launchd section's "exact allocation
  ... before this change" claim; both now state the inspection cost that
  comes off instead
- run-loop-shutdown-budget.test.ts: add a real 13ms elapsed-debit case at
  a 20s exit timeout (asserts reserve=9987, drain=5000, contrasting with
  the existing acceptedAtMs=MAX_SAFE_INTEGER zero-debit fixtures) and
  mirror darwin's 19/16/15/10s sub-20s it.each against a systemd stub,
  which previously had no sub-20s coverage

* test(cli): wrap the flat-subtraction assertion to satisfy oxfmt

The sub-20-second systemd cases added in 60e5eb509b6d left the flat-subtraction
comparison on a single line past the formatter's width, which failed oxfmt --check
and took check-lint and check-docs down with it. Lint itself reported zero warnings
and zero errors. Formatting only, no assertion or value changed.

* docs(gateway): bound the 20-second allocation claim by the inspection debit

The previous wording paired 'gives up exactly what the probe cost' with a probe
bounded near 6 seconds, and those two cannot both hold. Which allowance pays the
debit turns on a 5 second threshold: at or under it the reserve absorbs the whole
cost and the 5 second drain floor is untouched, so a 20 second deadline resolves
10000 minus the debit. Past it the floor is share-bounded as well and the two
converge on half the remainder, so a 6 second debit splits 9000 into 4500/4500
rather than retaining the old reserve. Documents the threshold in the docs and the
root-module comment, and pins the over-threshold case with a test so the boundary
is measured rather than reasoned.

* docs(gateway): qualify the full-reserve claim by the probe debit

The drain-floor comment asserted that a whole 10 second reserve follows
for every ExitTimeOut of 20 seconds or more once the probe that reads the
deadline is subtracted. The probe cost is charged before the budget is
split, so a 20 second ExitTimeOut clears the 15 second budget that a whole
reserve needs only when that probe costs nothing. At the 13ms measured on
this rig the reserve is 9987, and a whole reserve at that cost needs 20013.

The same comment already states this correctly six lines down, where it
resolves a 20 second deadline to 10000 - cost. This removes the
contradiction by qualifying the claim at its first statement instead of
restating the arithmetic twice.

Comment-only: no non-comment line changes and the export list is unchanged.

* fix(gateway): honor unlimited launchd stops

Preserve the observed launcher cap only for launchd-driven stops, and shorten redundant budget documentation without changing the measured split.

Co-authored-by: Patrick-Erichsen <20157849+Patrick-Erichsen@users.noreply.github.com>

* test(gateway): use action-based drain budget

Co-authored-by: Patrick-Erichsen <20157849+Patrick-Erichsen@users.noreply.github.com>

* style(gateway): format launchd timeout assertions

Co-authored-by: Patrick-Erichsen <20157849+Patrick-Erichsen@users.noreply.github.com>

* fix(ci): complete frozen Docker planner import closures

---------

Co-authored-by: roboclaw-bot <309084314+roboclaw-bot@users.noreply.github.com>
Co-authored-by: Patrick-Erichsen <20157849+Patrick-Erichsen@users.noreply.github.com>
Co-authored-by: Patrick Erichsen <patrick.a.erichsen@gmail.com>
This commit is contained in:
Mitchell Etzel 2026-09-28 18:04:14 -07:00 • committed by GitHub
parent d97d619dba
commit 6eeb120d69
No known key found for this signature in database
GPG key ID: B5690EEEBB952194
17 changed files with 2027 additions and 144 deletions

View file

@ -159,9 +159,14 @@ Node version and the service manager tracks a launcher parent. The launcher
forwards the stop signal and waits for the serving Gateway to drain within the
shared service budget. Managed restart intent targets the live serving owner,
so unfinished work still follows restart recovery when its drain budget expires.
The launchd stop budget remains 20 seconds; Linux units use the deadlines below.
When launchd drives the stop, macOS uses the running job's own `ExitTimeOut`, capped
by the launcher's own stop timer when a launcher is in the path, and Linux units use
their own stop timeout; both are described in the deadline sections below.
This requires a Gateway started with the updated launcher: replacing files cannot
change a launcher that is already running.
change a launcher that is already running. The stop deadlines below are the
exception: the serving Gateway derives the launcher's reap timer rather than being
told it, precisely so a Gateway started by an already-running older launcher still
bounds itself correctly.
For these managed restarts, if the CLI cannot verify the service command, serving
owner, or restart-intent recording, it refuses the restart before signaling with
@ -214,11 +219,25 @@ unit's effective `TimeoutStopUSec`, including drop-ins. It logs the source and
reconciled stop budget at both points, so a repaired unit takes effect without
restarting first. Inspection and any wait for startup to finish consume the same
shutdown deadline. Active-work drain uses at most
315 seconds, with 10 seconds reserved for final chat writes and server cleanup
and another 5 seconds before systemd's deadline. A unit with the default
315 seconds, with up to 10 seconds reserved for final chat writes and server cleanup
and up to another 5 seconds before systemd's deadline. A unit with the default
90-second stop timeout therefore gets a 75-second drain and an 85-second Gateway
shutdown deadline. A shorter supervisor timeout also caps requested restart waits.
The drained work, ordering, and interruption behavior stay the same.
Both allowances are bounded when the deadline cannot fund them, by the same rule the
launchd budget uses and for the same reason: subtracting two values sized for a
315-second drain from a short stop timeout consumed it entirely and left active work
nothing. A unit whose `TimeoutStopSec` is 20 seconds or longer keeps the allocation it
already had, less whatever the stop inspection itself cost, rather than to the exact
millisecond: 20 seconds is only the least that funds the full 10-second reserve
alongside a 5-second drain before that cost comes off. Below that the deadline cannot fund
both, and the reserve yields to keep a drain: a unit at or below 15 seconds now drains
at all where it previously drained for zero milliseconds, and the four values between
trade part of a reserve for a drain that was under 5 seconds. The reserve never falls
below half the shutdown budget. The exit margin holds its full 5 seconds down to a
20-second deadline and shrinks below that, to 3.75 seconds at 15 seconds and a quarter
of anything shorter. The drained work, ordering, and interruption behavior stay the
same, and systemd's own 90-second default is unaffected.
Service-child cleanup uses the remaining Gateway shutdown budget, leaving time
for final exit bookkeeping. A forced restart drains admitted work within the same
@ -309,6 +328,59 @@ before acknowledging, without overwriting a newer turn in that session.
If another child cannot be stopped, the response still reports incomplete
cancellation; the captured parent's cancellation is persisted before that error.
### Launchd stop deadlines
On a launchd-driven stop, the macOS Gateway reads the **loaded** job's effective
`exit timeout` with `launchctl print`. It checks the system, GUI, and user
domains, and accepts only a job whose PID is the Gateway or its launcher. Restart
ownership can still be external: the supervisor that enforces the stop, not the
owner of the next start, determines the deadline. A direct SIGTERM does not start
launchd's stop clock, so it keeps the existing Gateway stop policy unless an
OpenClaw launcher independently enforces its own child reap timer.
The **stop budget** is time available until the supervisor can kill the process.
The Gateway sets aside an exit margin (up to 5 seconds) and plans its shutdown
inside the remainder. **Drain** is the first part of that shutdown: stop admitting
new work and wait for active turns and background tasks to settle. The rest is
reserved for final writes and service cleanup (up to 10 seconds). The Gateway
logs the chosen source, drain, shutdown deadline, reserve, and exit margin.
For the installed 20-second LaunchAgent job, the nominal split is 5 seconds of
drain, 10 seconds for cleanup, and 5 seconds before launchd's deadline. Time
spent inspecting the job is debited. For a shorter custom deadline the exit
margin is at most a quarter, and the cleanup reserve gives way to leave active
work up to 5 seconds of drain (at most half of the remaining shutdown budget
when it is very short). For example, a 5-second job nominally leaves 1.875
seconds each for drain and cleanup after its 1.25-second exit margin; a
15-second job leaves 5 seconds of drain, 6.25 seconds for cleanup, and a
3.75-second margin. **This reduces cleanup time for custom jobs below 20
seconds.** The default systemd 90-second deadline is unaffected. The probe can
consume up to three 2-second calls; a very slow inspection can leave no drain.
If a Node-recovery or compile-cache launcher is the job's PID, its own child
reap timer may be shorter than launchd's deadline. The Gateway caps its budget
at that timer only when its respawn markers establish that this launcher started
it; an unrelated parent does not shorten the budget.
A direct signal to a marked launcher can start its own timer while launchd
still reports `running`. The Gateway cannot distinguish that forwarded signal
from a direct signal to its child, where the launcher has no timer, so it does
not cap this case. Operators stopping a launcher directly should use the loaded
job stop instead to get an enforceable, reported deadline.
`ExitTimeOut=0` means no launchd stop deadline, not a missing value. Without a
launcher timer, the Gateway keeps its ordinary policy and does not classify the
job as a native deadline. An unreadable job leaves
the existing stop policy in place with a warning, while a confirmed stopping
job whose timeout is missing uses launchd's 20-second default with a warning.
The value comes from the job launchd has **loaded**, not the plist on disk.
Editing `ExitTimeOut` or running `launchctl kickstart -k` does not reload it;
`launchctl bootout` followed by `launchctl bootstrap` does, restarting the
Gateway. Reading per stop avoids a stale startup snapshot, not a stale loaded
job. On macOS 27, launchd reports at most 60 seconds even if the plist asks
for more; inspect the loaded job to confirm the effective value.
## Host sleep and process freezes
When a gateway host wakes from sleep, a virtual machine resumes, or the process

View file

@ -4,3 +4,13 @@ export const GATEWAY_SHUTDOWN_TIMEOUT_MS: number;
export const GATEWAY_SERVICE_STOP_TIMEOUT_MS: number;
export const LAUNCH_AGENT_EXIT_TIMEOUT_SECONDS: 20;
export const GATEWAY_RESTART_REPLACEMENT_TIMEOUT_MS: number;
export const RESPAWN_SIGNAL_FORCE_KILL_GRACE_MS: number;
export const RESPAWN_SIGNAL_HARD_EXIT_GRACE_MS: number;
export function resolveSupervisorExitMarginMs(stopTimeoutMs: number): number;
export function resolveShutdownReserveMs(shutdownTimeoutMs: number): number;
export function isRespawnedByLauncher(env: NodeJS.ProcessEnv): boolean;
export function resolveLauncherStopTimeoutMs(params: {
env: NodeJS.ProcessEnv;
platform: NodeJS.Platform;
foreground: boolean;
}): number;

View file

@ -9,3 +9,72 @@ export const GATEWAY_SERVICE_STOP_TIMEOUT_MS =
GATEWAY_SHUTDOWN_TIMEOUT_MS + GATEWAY_SUPERVISOR_EXIT_MARGIN_MS;
export const LAUNCH_AGENT_EXIT_TIMEOUT_SECONDS = 20;
// Keep a positive shutdown budget when a supervisor's deadline is under 20s.
const GATEWAY_SUPERVISOR_EXIT_MARGIN_SHARE = 0.25;
// Preserve the template's 5s drain where possible. Under 20s, trade some cleanup
// reserve for drain; on very short jobs, split the remaining budget in half.
// Inspection time is debited before this allocation, so a slow probe can also
// reduce the drain. The measured upgrade tradeoff is documented in restart-recovery.
const GATEWAY_SHUTDOWN_DRAIN_FLOOR_MS = 5_000;
const GATEWAY_SHUTDOWN_DRAIN_FLOOR_SHARE = 0.5;
/** The exit margin to hold back from a supervisor-enforced stop deadline. */
export const resolveSupervisorExitMarginMs = (stopTimeoutMs) =>
Math.min(
GATEWAY_SUPERVISOR_EXIT_MARGIN_MS,
Math.floor(Math.max(0, stopTimeoutMs) * GATEWAY_SUPERVISOR_EXIT_MARGIN_SHARE),
);
/** The post-drain reserve to hold back from a resolved shutdown budget. */
export const resolveShutdownReserveMs = (shutdownTimeoutMs) => {
const budgetMs = Math.max(0, shutdownTimeoutMs);
const drainFloorMs = Math.min(
GATEWAY_SHUTDOWN_DRAIN_FLOOR_MS,
Math.floor(budgetMs * GATEWAY_SHUTDOWN_DRAIN_FLOOR_SHARE),
);
return Math.min(GATEWAY_SHUTDOWN_RESERVE_MS, budgetMs - drainFloorMs);
};
// Escalation graces the Node recovery launcher applies to a stopping child. Kept
// here rather than in the launcher so the serving Gateway can derive the deadline
// its parent enforces from the same numbers the parent armed it from.
const RESPAWN_SIGNAL_EXIT_GRACE_MS = 1_000;
export const RESPAWN_SIGNAL_FORCE_KILL_GRACE_MS = 1_000;
export const RESPAWN_SIGNAL_HARD_EXIT_GRACE_MS = 1_000;
// A launcher inside OpenClaw's LaunchAgent gets its template deadline; otherwise
// it uses the generic service stop policy. Its child inherits these markers.
const resolveRespawnServiceStopTimeoutMs = (env, platform) => {
const launchdService = env.OPENCLAW_LAUNCHD_LABEL?.trim();
return platform === "darwin" && launchdService && env.XPC_SERVICE_NAME === launchdService
? LAUNCH_AGENT_EXIT_TIMEOUT_SECONDS * 1_000
: GATEWAY_SERVICE_STOP_TIMEOUT_MS;
};
// Every runRespawnedChild call site stamps one of these. The compile-cache
// marker is also used by a different respawner, but that one refuses foreground
// Gateway runs on darwin. Keep this list with the launcher deadline it authorizes.
const RESPAWN_LAUNCHER_MARKER_ENV_VARS = [
"OPENCLAW_NODE_UPDATE_RESPAWNED",
"OPENCLAW_COMPILE_CACHE_DISABLED_RESPAWNED",
"OPENCLAW_PACKAGED_COMPILE_CACHE_RESPAWNED",
];
/** Whether a `runRespawnedChild` parent started this process. */
export const isRespawnedByLauncher = (env) =>
RESPAWN_LAUNCHER_MARKER_ENV_VARS.some((name) => env[name] === "1");
// Shared with the serving Gateway: derive the parent's reap deadline from the
// same graces, including on the first update while the old launcher is still live.
export const resolveLauncherStopTimeoutMs = ({ env, platform, foreground }) => {
const serviceStopTimeoutMs = resolveRespawnServiceStopTimeoutMs(env, platform);
const signalExitGraceMs =
platform !== "win32" && foreground
? serviceStopTimeoutMs -
RESPAWN_SIGNAL_FORCE_KILL_GRACE_MS -
RESPAWN_SIGNAL_HARD_EXIT_GRACE_MS
: RESPAWN_SIGNAL_EXIT_GRACE_MS;
return signalExitGraceMs + RESPAWN_SIGNAL_FORCE_KILL_GRACE_MS;
};

View file

@ -6,8 +6,9 @@ import path from "node:path";
import { consumeRootOptionToken as consumeLauncherRootOptionToken } from "./cli-root-options.mjs";
import { isForegroundGatewayRunArgv } from "./gateway-run-argv.mjs";
import {
GATEWAY_SERVICE_STOP_TIMEOUT_MS,
LAUNCH_AGENT_EXIT_TIMEOUT_SECONDS,
RESPAWN_SIGNAL_FORCE_KILL_GRACE_MS,
RESPAWN_SIGNAL_HARD_EXIT_GRACE_MS,
resolveLauncherStopTimeoutMs,
} from "./gateway-shutdown-budget.mjs";
import {
detectCurrentSqliteCapabilities,
@ -39,22 +40,21 @@ const respawnSignals =
process.platform === "win32"
? ["SIGTERM", "SIGINT", "SIGBREAK"]
: ["SIGTERM", "SIGINT", "SIGHUP", "SIGQUIT"];
const respawnSignalExitGraceMs = 1_000;
const respawnSignalForceKillGraceMs = 1_000;
const respawnSignalHardExitGraceMs = 1_000;
const respawnSignalForceKillGraceMs = RESPAWN_SIGNAL_FORCE_KILL_GRACE_MS;
const respawnSignalHardExitGraceMs = RESPAWN_SIGNAL_HARD_EXIT_GRACE_MS;
export const runRespawnedChild = (command, args, env) => {
const launchdService = env.OPENCLAW_LAUNCHD_LABEL?.trim();
const serviceStopTimeoutMs =
process.platform === "darwin" && launchdService && env.XPC_SERVICE_NAME === launchdService
? LAUNCH_AGENT_EXIT_TIMEOUT_SECONDS * 1_000
: GATEWAY_SERVICE_STOP_TIMEOUT_MS;
// The serving Gateway owns drain and cleanup. Reap a stuck child only in the
// supervisor's exit margin, after that owner has had its full shutdown budget.
const signalExitGraceMs =
process.platform !== "win32" && isForegroundGatewayRunArgv(process.argv)
? serviceStopTimeoutMs - respawnSignalForceKillGraceMs - respawnSignalHardExitGraceMs
: respawnSignalExitGraceMs;
// The shared resolver owns this arithmetic so the serving Gateway derives the very
// same deadline from the same expression, which is what lets a Gateway started by
// any build of this launcher bound itself correctly without being told.
const launcherStopTimeoutMs = resolveLauncherStopTimeoutMs({
env,
platform: process.platform,
foreground: isForegroundGatewayRunArgv(process.argv),
});
const signalExitGraceMs = launcherStopTimeoutMs - respawnSignalForceKillGraceMs;
const stdioIsTerminal = process.stdin.isTTY || process.stdout.isTTY;
const child = spawn(command, args, {
stdio: "inherit",
@ -62,7 +62,10 @@ export const runRespawnedChild = (command, args, env) => {
windowsHide: !stdioIsTerminal,
});
const listeners = new Map();
// Keep signal forwarding and bounded shutdown in sync with src/entry.compile-cache.ts.
// Keep signal forwarding and bounded shutdown in sync with src/entry.compile-cache.ts,
// which drives src/process/respawn-child-runner.ts. That runner still holds its own
// copies of the escalation graces and reaps on a fixed short one, so only this
// launcher's deadline is the one the serving Gateway derives.
let signalExitTimer = null;
let signalForceKillTimer = null;
let signalHardExitTimer = null;

View file

@ -1180,11 +1180,6 @@ async function preflightFrozenTargetContracts(input, workflow = false, verifiedT
for (const path of supportFiles[consumer] ?? []) {
required(sources.tooling, `scripts/e2e/lib/${path}`);
}
if (
["npm-onboard-channel-agent", "codex-on-demand", "update-corrupt-plugin"].includes(consumer)
) {
required(sources.tooling, "scripts/lib/record-shared.mjs");
}
if (consumer === "update-corrupt-plugin") {
required(sources.tooling, "scripts/lib/update-compat-contract.mjs");
required(sources.tooling, "scripts/lib/openclaw-e2e-instance.sh");

View file

@ -1,16 +1,29 @@
import { performance } from "node:perf_hooks";
import { afterEach, beforeEach, describe, expect, it, vi } from "vitest";
import { resolveGatewayShutdownBudget } from "./run-loop-shutdown-budget.js";
import {
GATEWAY_SHUTDOWN_RESERVE_MS,
GATEWAY_SUPERVISOR_EXIT_MARGIN_MS,
} from "../../infra/gateway-shutdown-budget.js";
import {
resolveGatewayShutdownBudget,
resolveGatewayShutdownDrainBudget,
} from "./run-loop-shutdown-budget.js";
const { readFile, execUser, execSystem } = vi.hoisted(() => ({
const { readFile, execUser, execSystem, execLaunchctl } = vi.hoisted(() => ({
readFile: vi.fn(),
execUser: vi.fn(),
execSystem: vi.fn(),
execLaunchctl: vi.fn(),
}));
vi.mock("node:fs/promises", () => ({ default: { readFile } }));
vi.mock("../../daemon/systemd-exec.js", () => ({
execSystemctlUser: execUser,
execSystemctl: execSystem,
}));
vi.mock("../../daemon/launchd-exec.js", async (importOriginal) => ({
...(await importOriginal<typeof import("../../daemon/launchd-exec.js")>()),
execLaunchctl,
}));
beforeEach(() => {
vi.stubGlobal("process", { ...process, platform: "linux", getuid: () => 1000, env: {} });
@ -23,6 +36,7 @@ beforeEach(() => {
"User=openclaw\nType=simple\nNotifyAccess=none\nKillMode=control-group",
stderr: "",
});
execLaunchctl.mockReset();
});
afterEach(() => vi.unstubAllGlobals());
@ -93,4 +107,472 @@ describe("Gateway stop deadline independent of restart ownership", () => {
expect(execSystem).not.toHaveBeenCalled();
expect(execUser).not.toHaveBeenCalled();
});
// The systemd coverage above only pins TimeoutStopUSec at 90 seconds and 10 minutes,
// both comfortably above the 20 second threshold. This mirrors the darwin exit-timeout
// table's sub-20-second cases against a systemd stub: a unit deadline in this band
// gives up part of a reserve a flat subtraction would have funded in full, rather than
// draining for zero milliseconds as that flat subtraction did.
it.each([
{ seconds: 19, timeoutMs: 14_250, reserveMs: 9_250, drainMs: 5_000, fixedDrainMs: 4_000 },
{ seconds: 16, timeoutMs: 12_000, reserveMs: 7_000, drainMs: 5_000, fixedDrainMs: 1_000 },
{ seconds: 15, timeoutMs: 11_250, reserveMs: 6_250, drainMs: 5_000, fixedDrainMs: 0 },
{ seconds: 10, timeoutMs: 7_500, reserveMs: 3_750, drainMs: 3_750, fixedDrainMs: 0 },
])(
"keeps a drain a $seconds second systemd TimeoutStopUSec previously spent on overhead",
async ({ seconds, timeoutMs, reserveMs, drainMs, fixedDrainMs }) => {
process.env.OPENCLAW_SUPERVISOR_MODE = "external";
execSystem.mockResolvedValue({
code: 0,
stdout: `LoadState=loaded\nTimeoutStopUSec=${seconds}s\nInvocationID=own`,
stderr: "",
});
const budget = await resolveGatewayShutdownBudget("external", {
info: vi.fn(),
warn: vi.fn(),
});
expect(budget.timeoutMs).toBe(timeoutMs);
expect(budget.reserveMs).toBe(reserveMs);
expect(budget.timeoutMs - budget.reserveMs).toBe(drainMs);
// What a flat, unshared subtraction would have left active work at the same deadline.
expect(
Math.max(
0,
seconds * 1_000 - GATEWAY_SUPERVISOR_EXIT_MARGIN_MS - GATEWAY_SHUTDOWN_RESERVE_MS,
),
).toBe(fixedDrainMs);
expect(budget.reserveMs).toBeGreaterThanOrEqual(Math.floor(timeoutMs / 2));
},
);
});
describe("Gateway stop deadline follows the launchd stop that is actually running", () => {
// A stop that is under way. `previous` is the startup budget a darwin Gateway
// resolves before any stop exists, which is the platform-neutral policy.
// The budget subtracts `performance.now() - acceptedAtMs` and floors it at
// zero, so an acceptance stamped ahead of the clock records exactly no elapsed
// time and keeps the asserted numbers exact instead of off by a stray
// millisecond.
const stoppingNow = {
previous: { timeoutMs: 325_000, nativeStopBudget: false },
acceptedAtMs: Number.MAX_SAFE_INTEGER,
};
const printed = (state: string, fields: string) => ({
code: 0,
stdout: `system/ai.openclaw.gateway = {\n\tstate = ${state}\n\n${fields}\tresource coalition = {\n\t\tstate = active\n\t}\n}\n`,
stderr: "",
termination: "exit",
});
beforeEach(() => {
vi.stubGlobal("process", {
...process,
platform: "darwin",
pid: 4242,
getuid: () => 501,
env: {},
});
process.env.XPC_SERVICE_NAME = "ai.openclaw.gateway";
process.env.OPENCLAW_SUPERVISOR_MODE = "external";
});
// THE REGRESSION GUARD. The linked report is an externally delivered SIGTERM
// under a five second job, where launchd never starts its clock and the drain ran
// its full 315 seconds. Measured on the reporting host, it then hit its own
// timeout with work still active rather than finishing early, so the drain was
// being used. Adopting the job deadline there would hand that same supported
// setup a zero drain and cut that work off at once.
it("keeps the full drain when launchd did not initiate the stop", async () => {
execLaunchctl.mockResolvedValue(printed("running", "\texit timeout = 5\n\tpid = 4242\n"));
const info = vi.fn();
const budget = await resolveGatewayShutdownBudget(
"external",
{ info, warn: vi.fn() },
stoppingNow,
);
budget.log("shutdown");
expect(info).toHaveBeenCalledWith(
"shutdown budget at shutdown: drain=315000ms shutdown=325000ms reserve=10000ms exitMargin=5000ms; source=Gateway stop policy=330000ms",
);
expect(budget.timeoutMs).toBe(325_000);
expect(budget.nativeStopBudget).toBe(false);
});
it.each(["external", "launchd"])(
"does not treat a launchd ExitTimeOut of zero as a native deadline under %s ownership",
async (supervisor) => {
execLaunchctl.mockResolvedValue(printed("SIGTERMed", "\texit timeout = 0\n\tpid = 4242\n"));
const budget = await resolveGatewayShutdownBudget(
supervisor,
{ info: vi.fn(), warn: vi.fn() },
stoppingNow,
);
expect(budget.timeoutMs).toBe(325_000);
expect(budget.nativeStopBudget).toBe(false);
const drain = resolveGatewayShutdownDrainBudget({
budget,
action: "restart",
forceRestart: false,
restartWithoutSupervisor: false,
acceptedAtMs: performance.now(),
requestedRestartDrainTimeoutMs: 600_000,
});
expect(drain.drainTimeoutMs).toBeGreaterThan(590_000);
},
);
// A deadline this short cannot fund the fixed 5s margin and 10s reserve, and
// subtracting them outright left the job's whole 5 seconds spent on overhead with
// nothing to drain. Each allowance is capped at a share of what it is carved from,
// so the short job keeps a proportional drain that still fits inside the deadline.
it("keeps a proportional drain when the job's exit timeout cannot fund the fixed allowances", async () => {
execLaunchctl.mockResolvedValue(
printed("SIGTERMed", "\tminimum runtime = 10\n\texit timeout = 5\n\tpid = 4242\n"),
);
const info = vi.fn();
const budget = await resolveGatewayShutdownBudget(
"external",
{ info, warn: vi.fn() },
stoppingNow,
);
budget.log("shutdown");
expect(info).toHaveBeenCalledWith(
"shutdown budget at shutdown: drain=1875ms shutdown=3750ms reserve=1875ms exitMargin=1250ms; source=launchd system/ai.openclaw.gateway exit timeout=5000ms",
);
expect(budget.timeoutMs).toBe(3_750);
expect(budget.reserveMs).toBe(1_875);
expect(budget.nativeStopBudget).toBe(true);
});
// Drain must never be starved to zero by the allowances: every positive deadline
// leaves active work some time, and a longer deadline never yields less of it.
it.each([1, 2, 5, 10, 15, 20, 25, 47, 60])(
"leaves a positive drain for a %s second exit timeout",
async (seconds) => {
execLaunchctl.mockResolvedValue(
printed("SIGTERMed", `\texit timeout = ${seconds}\n\tpid = 4242\n`),
);
const budget = await resolveGatewayShutdownBudget(
"external",
{ info: vi.fn(), warn: vi.fn() },
stoppingNow,
);
expect(budget.timeoutMs - budget.reserveMs).toBeGreaterThan(0);
expect(budget.timeoutMs).toBeLessThanOrEqual(seconds * 1_000);
},
);
it.each([
{ seconds: 20, timeoutMs: 15_000, reserveMs: 10_000, drainMs: 5_000, exitMarginMs: 5_000 },
{ seconds: 47, timeoutMs: 42_000, reserveMs: 10_000, drainMs: 32_000, exitMarginMs: 5_000 },
{ seconds: 55, timeoutMs: 50_000, reserveMs: 10_000, drainMs: 40_000, exitMarginMs: 5_000 },
])(
"derives the budget from a $seconds second exit timeout",
async ({ seconds, timeoutMs, reserveMs, drainMs, exitMarginMs }) => {
execLaunchctl.mockResolvedValue(
printed("SIGTERMed", `\texit timeout = ${seconds}\n\tpid = 4242\n`),
);
const info = vi.fn();
const budget = await resolveGatewayShutdownBudget(
"external",
{ info, warn: vi.fn() },
stoppingNow,
);
budget.log("shutdown");
expect(budget.timeoutMs).toBe(timeoutMs);
expect(budget.reserveMs).toBe(reserveMs);
expect(info).toHaveBeenCalledWith(
`shutdown budget at shutdown: drain=${drainMs}ms shutdown=${timeoutMs}ms reserve=${reserveMs}ms exitMargin=${exitMarginMs}ms; source=launchd system/ai.openclaw.gateway exit timeout=${seconds * 1_000}ms`,
);
},
);
// Every case above pins `acceptedAtMs` to `Number.MAX_SAFE_INTEGER`, which floors
// elapsed at zero and proves nothing about a real, nonzero debit. `performance.now`
// is not faked anywhere in this file, so this pins it directly rather than trusting
// the wall clock to land on a specific millisecond by chance: the reserve gives up
// exactly that observed 13ms debit and the 5 second drain floor still funds in full.
it("debits the reserve by a real elapsed delay, leaving the drain floor untouched", async () => {
execLaunchctl.mockResolvedValue(printed("SIGTERMed", "\texit timeout = 20\n\tpid = 4242\n"));
const nowMs = performance.now();
const clock = vi.spyOn(performance, "now").mockReturnValue(nowMs);
try {
const budget = await resolveGatewayShutdownBudget(
"external",
{ info: vi.fn(), warn: vi.fn() },
{ previous: stoppingNow.previous, acceptedAtMs: nowMs - 13 },
);
expect(budget.timeoutMs).toBe(14_987);
expect(budget.reserveMs).toBe(9_987);
expect(budget.timeoutMs - budget.reserveMs).toBe(5_000);
} finally {
clock.mockRestore();
}
});
// The case above only covers a debit small enough that the reserve pays it alone. The
// probe can cost far more than 13ms: it is up to three `launchctl print` calls at a
// 2 second timeout each. Past a 5 second debit the drain floor is share-bounded too,
// so the claim that a 20 second deadline keeps its old allocation less the debit stops
// holding and both allowances converge on half the remainder. Pinning that boundary
// keeps the documented threshold honest instead of reasoned.
it("converges the reserve and the drain once the elapsed debit passes the floor", async () => {
execLaunchctl.mockResolvedValue(printed("SIGTERMed", "\texit timeout = 20\n\tpid = 4242\n"));
const nowMs = performance.now();
const clock = vi.spyOn(performance, "now").mockReturnValue(nowMs);
try {
const budget = await resolveGatewayShutdownBudget(
"external",
{ info: vi.fn(), warn: vi.fn() },
{ previous: stoppingNow.previous, acceptedAtMs: nowMs - 6_000 },
);
expect(budget.timeoutMs).toBe(9_000);
expect(budget.reserveMs).toBe(4_500);
expect(budget.timeoutMs - budget.reserveMs).toBe(4_500);
// Not the old allocation less the debit: that would have left the reserve at 4000.
expect(budget.reserveMs).not.toBe(GATEWAY_SHUTDOWN_RESERVE_MS - 6_000);
} finally {
clock.mockRestore();
}
});
// The allowances are only ever capped to keep a drain, never to reallocate a deadline
// that already worked. A deadline able to fund the 10s reserve alongside the 5s drain
// the 20s template yields needs 15s of shutdown budget, which every deadline from 20s
// up has, so all of them must resolve exactly what subtracting the fixed allowances
// outright resolved. Asserting against that arithmetic rather than against literals is
// what makes this a regression test: capping the reserve at a share of the budget, as
// an earlier revision did, drops a 20s job's reserve to 7500ms and fails here.
it.each([20, 21, 25, 30, 47, 55, 60])(
"allocates a %s second exit timeout exactly as the fixed allowances did",
async (seconds) => {
execLaunchctl.mockResolvedValue(
printed("SIGTERMed", `\texit timeout = ${seconds}\n\tpid = 4242\n`),
);
const budget = await resolveGatewayShutdownBudget(
"external",
{ info: vi.fn(), warn: vi.fn() },
stoppingNow,
);
const fixedTimeoutMs = seconds * 1_000 - GATEWAY_SUPERVISOR_EXIT_MARGIN_MS;
expect(budget.timeoutMs).toBe(fixedTimeoutMs);
expect(budget.reserveMs).toBe(GATEWAY_SHUTDOWN_RESERVE_MS);
expect(budget.timeoutMs - budget.reserveMs).toBe(
fixedTimeoutMs - GATEWAY_SHUTDOWN_RESERVE_MS,
);
},
);
// Under 20 seconds the budget cannot fund both allowances, so one has to give. The
// fixed subtraction gave up the drain: 15 seconds and below drained for 0ms, and the
// four deadlines between left active work under 5 seconds. These pin what is given up
// instead, and that the reserve never falls below half the budget doing it. No shipped
// template or platform default lands here: the LaunchAgent template and launchd's own
// default are both 20 seconds, and systemd's default stop timeout is 90.
it.each([
{ seconds: 19, timeoutMs: 14_250, reserveMs: 9_250, drainMs: 5_000, fixedDrainMs: 4_000 },
{ seconds: 16, timeoutMs: 12_000, reserveMs: 7_000, drainMs: 5_000, fixedDrainMs: 1_000 },
{ seconds: 15, timeoutMs: 11_250, reserveMs: 6_250, drainMs: 5_000, fixedDrainMs: 0 },
{ seconds: 10, timeoutMs: 7_500, reserveMs: 3_750, drainMs: 3_750, fixedDrainMs: 0 },
])(
"keeps a drain a $seconds second exit timeout previously spent on overhead",
async ({ seconds, timeoutMs, reserveMs, drainMs, fixedDrainMs }) => {
execLaunchctl.mockResolvedValue(
printed("SIGTERMed", `\texit timeout = ${seconds}\n\tpid = 4242\n`),
);
const budget = await resolveGatewayShutdownBudget(
"external",
{ info: vi.fn(), warn: vi.fn() },
stoppingNow,
);
expect(budget.timeoutMs).toBe(timeoutMs);
expect(budget.reserveMs).toBe(reserveMs);
expect(budget.timeoutMs - budget.reserveMs).toBe(drainMs);
// What the fixed subtraction left active work at the same deadline.
expect(
Math.max(
0,
seconds * 1_000 - GATEWAY_SUPERVISOR_EXIT_MARGIN_MS - GATEWAY_SHUTDOWN_RESERVE_MS,
),
).toBe(fixedDrainMs);
expect(budget.reserveMs).toBeGreaterThanOrEqual(Math.floor(timeoutMs / 2));
},
);
// Failing to inspect the job establishes nothing, so shortening the drain here
// would cut work that no launchd deadline was bounding.
it("warns and keeps the platform-neutral policy when the job cannot be inspected", async () => {
execLaunchctl.mockResolvedValue({
code: 1,
stdout: "",
stderr: "permission denied",
termination: "exit",
});
const warn = vi.fn();
const budget = await resolveGatewayShutdownBudget(
"external",
{ info: vi.fn(), warn },
stoppingNow,
);
expect(budget.timeoutMs).toBe(325_000);
// The number alone is not the contract. A failed probe confirmed no launchd
// deadline, so this must not be classified as a native stop budget either:
// that flag is what caps a restart drain and arms a forced exit.
expect(budget.nativeStopBudget).toBe(false);
expect(warn).toHaveBeenCalledExactlyOnceWith(
expect.stringContaining("Unable to inspect the launchd job"),
);
});
// The flag is only worth asserting because of what it does downstream, so drive
// the real consumer. An operator restart that asked to drain for ten minutes
// keeps that request when no launchd deadline was ever confirmed.
it("leaves a longer requested restart drain uncapped when the job cannot be inspected", async () => {
execLaunchctl.mockResolvedValue({
code: 1,
stdout: "",
stderr: "permission denied",
termination: "exit",
});
const budget = await resolveGatewayShutdownBudget(
"external",
{ info: vi.fn(), warn: vi.fn() },
stoppingNow,
);
expect(budget.nativeStopBudget).toBe(false);
const drain = resolveGatewayShutdownDrainBudget({
budget,
action: "restart",
forceRestart: false,
restartWithoutSupervisor: false,
acceptedAtMs: performance.now(),
requestedRestartDrainTimeoutMs: 600_000,
});
// Only elapsed time comes off the request; no supervisor ceiling applies.
expect(drain.drainTimeoutMs).toBeGreaterThan(590_000);
expect(drain.restartTimeoutMs()).toBe(325_000);
});
// The same consumer, with a deadline that WAS confirmed, still gets capped.
// Without this pair the test above would also pass if the launchd read were
// deleted outright.
it("caps that same restart drain at a confirmed job deadline", async () => {
execLaunchctl.mockResolvedValue(printed("SIGTERMed", "\texit timeout = 47\n\tpid = 4242\n"));
const budget = await resolveGatewayShutdownBudget(
"external",
{ info: vi.fn(), warn: vi.fn() },
stoppingNow,
);
expect(budget.nativeStopBudget).toBe(true);
const drain = resolveGatewayShutdownDrainBudget({
budget,
action: "restart",
forceRestart: false,
restartWithoutSupervisor: false,
acceptedAtMs: performance.now(),
requestedRestartDrainTimeoutMs: 600_000,
});
// 47s job - 5s exit margin = 42000ms shutdown, less the 10000ms reserve.
expect(drain.drainTimeoutMs).toBe(32_000);
});
// A launchd-OWNED Gateway, rather than an externally supervised one. Its startup
// budget is already native, so this is the configuration where the retained-budget
// safety net can fire. `previous` is what a 20 second template job resolves.
const launchdOwnedStop = {
previous: { timeoutMs: 15_000, nativeStopBudget: true },
acceptedAtMs: Number.MAX_SAFE_INTEGER,
};
// An in-process restart signals the Gateway without launchd running the stop, so
// the job still prints `running`. That is a confirmed answer, not a failed probe,
// and claiming the deadline "could not be confirmed" there would be false on every
// in-process restart of a default macOS install.
it("does not claim an unconfirmed timeout when launchd is confirmed not to be stopping", async () => {
delete process.env.OPENCLAW_SUPERVISOR_MODE;
execLaunchctl.mockResolvedValue(printed("running", "\texit timeout = 20\n\tpid = 4242\n"));
const warn = vi.fn();
const budget = await resolveGatewayShutdownBudget(
"launchd",
{ info: vi.fn(), warn },
launchdOwnedStop,
);
expect(warn).not.toHaveBeenCalled();
expect(budget.timeoutMs).toBe(15_000);
expect(budget.nativeStopBudget).toBe(true);
});
// A probe that established nothing is the case the safety net exists for, so the
// startup budget is held rather than widened.
it("retains the startup budget when the job could not be inspected at all", async () => {
delete process.env.OPENCLAW_SUPERVISOR_MODE;
execLaunchctl.mockResolvedValue({
code: 1,
stdout: "",
stderr: "permission denied",
termination: "exit",
});
const warn = vi.fn();
const budget = await resolveGatewayShutdownBudget(
"launchd",
{ info: vi.fn(), warn },
launchdOwnedStop,
);
expect(warn).toHaveBeenCalledWith(expect.stringContaining("Unable to inspect the launchd job"));
expect(warn).toHaveBeenCalledWith(
"Retaining the startup shutdown budget of 15000ms because the current supervisor stop timeout could not be confirmed.",
);
expect(budget.timeoutMs).toBe(15_000);
});
// Warned but non-null is the defaulted-value case, not a failed probe: launchd is
// confirmed to be stopping the job, so a clock is running and there is nothing to
// retain. Warning about a timeout that "could not be confirmed" here would be the
// same false statement in a different place.
it("does not retain when only the deadline's value had to be defaulted", async () => {
delete process.env.OPENCLAW_SUPERVISOR_MODE;
execLaunchctl.mockResolvedValue(printed("SIGTERMed", "\tpid = 4242\n"));
const warn = vi.fn();
const budget = await resolveGatewayShutdownBudget(
"launchd",
{ info: vi.fn(), warn },
launchdOwnedStop,
);
expect(warn).toHaveBeenCalledExactlyOnceWith(
expect.stringContaining("its exit timeout is missing or invalid"),
);
expect(budget.timeoutMs).toBe(15_000);
});
// No stop is running at startup, so there is no enforcing deadline to read and
// no reason to spend a launchctl print discovering that.
it("does not inspect the job at startup", async () => {
const info = vi.fn();
const budget = await resolveGatewayShutdownBudget("external", { info, warn: vi.fn() });
budget.log("startup");
expect(execLaunchctl).not.toHaveBeenCalled();
expect(budget.timeoutMs).toBe(325_000);
expect(budget.nativeStopBudget).toBe(false);
expect(info).toHaveBeenCalledWith(
"shutdown budget at startup: drain=315000ms shutdown=325000ms reserve=10000ms exitMargin=5000ms; source=Gateway stop policy=330000ms",
);
});
it("keeps the platform-neutral policy when darwin is not running a launchd job", async () => {
process.env = {};
const warn = vi.fn();
const budget = await resolveGatewayShutdownBudget(null, { info: vi.fn(), warn }, stoppingNow);
expect(budget.timeoutMs).toBe(325_000);
expect(budget.nativeStopBudget).toBe(false);
expect(warn).not.toHaveBeenCalled();
expect(execLaunchctl).not.toHaveBeenCalled();
});
it("never reads systemd on darwin", async () => {
execLaunchctl.mockResolvedValue(printed("SIGTERMed", "\texit timeout = 20\n\tpid = 4242\n"));
await resolveGatewayShutdownBudget("external", { info: vi.fn(), warn: vi.fn() }, stoppingNow);
expect(execSystem).not.toHaveBeenCalled();
expect(execUser).not.toHaveBeenCalled();
expect(readFile).not.toHaveBeenCalled();
});
});

View file

@ -2,13 +2,58 @@ import { performance } from "node:perf_hooks";
import { LAUNCH_AGENT_EXIT_TIMEOUT_SECONDS } from "../../daemon/launchd-plist.js";
import {
GATEWAY_SERVICE_STOP_TIMEOUT_MS,
GATEWAY_SHUTDOWN_RESERVE_MS,
GATEWAY_SHUTDOWN_TIMEOUT_MS,
GATEWAY_SUPERVISOR_EXIT_MARGIN_MS,
resolveShutdownReserveMs,
resolveSupervisorExitMarginMs,
} from "../../infra/gateway-shutdown-budget.js";
import { readLaunchdStopTimeout } from "../../infra/launchd-stop-timeout.js";
import { readSystemdStopTimeout } from "../../infra/systemd-stop-timeout.js";
import type { GatewayRunSignalAction } from "./run-loop-request.js";
type NativeStopTimeout = { timeoutMs: number; source: string };
/**
* Ask whichever supervisor actually enforces the deadline on this platform.
*
* Three independent answers. `stop` is a deadline that may be spent as a native
* stop budget. `warning` is what the operator needs to hear. `inconclusive` says
* the probe could not establish an answer at all, which is the only case the
* retained-budget safety net below is for: a read that positively determined no
* launchd deadline governs this stop is an answer, not a failure, so retaining a
* startup budget and warning that the timeout "could not be confirmed" would be
* false on every in-process restart of a launchd-owned Gateway.
*/
async function readNativeStopTimeout(stopping: boolean): Promise<{
stop: NativeStopTimeout | null;
warning?: string;
inconclusive: boolean;
}> {
if (process.platform === "linux") {
const systemd = await readSystemdStopTimeout();
// Unchanged from the linux-only original: absent unit or warned read both
// count as unconfirmed there.
return {
stop: systemd,
warning: systemd?.warning,
inconclusive: !systemd || Boolean(systemd.warning),
};
}
// launchd's ExitTimeOut bounds a stop that launchd is running and nothing else:
// an externally delivered SIGTERM never starts that clock, and the job outlives
// the deadline untouched. There is no enforcing deadline to read before a stop
// is under way, and reading one at startup would spend a launchctl print only
// to adopt a deadline that does not govern the stop the Gateway will get.
if (process.platform === "darwin" && stopping) {
const read = await readLaunchdStopTimeout();
// Warned but non-null is the defaulted-value case: launchd is confirmed to be
// stopping the job and only its deadline had to be guessed, so a clock is
// genuinely running and nothing needs retaining. Only a warning with no
// deadline at all means the probe established nothing.
return { ...read, inconclusive: read.stop === null && read.warning !== undefined };
}
return { stop: null, inconclusive: false };
}
export async function resolveGatewayShutdownBudget(
supervisor: string | null,
logger: { info(message: string): void; warn(message: string): void },
@ -17,37 +62,45 @@ export async function resolveGatewayShutdownBudget(
acceptedAtMs: number;
},
) {
// Restart ownership may be external while systemd still enforces the stop deadline.
const systemdStop = process.platform === "linux" ? await readSystemdStopTimeout() : null;
// Restart ownership may be external while the platform supervisor still
// enforces the stop deadline. That holds on darwin exactly as it does on linux.
const native = await readNativeStopTimeout(refresh !== undefined);
const nativeStop = native.stop;
const retained =
refresh?.previous.nativeStopBudget && (!systemdStop || systemdStop.warning)
? refresh.previous
: undefined;
if (systemdStop?.warning) {
logger.warn(systemdStop.warning);
refresh?.previous.nativeStopBudget && native.inconclusive ? refresh.previous : undefined;
if (native.warning) {
logger.warn(native.warning);
}
if (retained) {
logger.warn(
`Retaining the startup shutdown budget of ${retained.timeoutMs}ms because the current systemd stop timeout could not be confirmed.`,
`Retaining the startup shutdown budget of ${retained.timeoutMs}ms because the current supervisor stop timeout could not be confirmed.`,
);
}
const stop = systemdStop ?? {
const stop = nativeStop ?? {
timeoutMs:
supervisor === "launchd"
? LAUNCH_AGENT_EXIT_TIMEOUT_SECONDS * 1_000
: GATEWAY_SERVICE_STOP_TIMEOUT_MS,
source: supervisor === "launchd" ? "launchd ExitTimeOut" : "Gateway stop policy",
};
const nativeStopBudget = systemdStop !== null || supervisor === "launchd" || Boolean(retained);
// ExitTimeOut=0 is unlimited. It is an observed job value, not a native
// deadline that may cap a requested restart or arm a forced exit.
const nativeStopBudget = nativeStop
? Number.isFinite(nativeStop.timeoutMs)
: supervisor === "launchd" || Boolean(retained);
// An operator job may enforce a deadline far shorter than the policy these fixed
// allowances were sized against, so each is capped at a share of what it is carved
// from. A deadline long enough to fund them is unaffected; a short one keeps a
// proportional drain instead of surrendering all of it to margin and reserve.
const exitMarginMs = resolveSupervisorExitMarginMs(stop.timeoutMs);
const limitMs =
retained?.timeoutMs ??
Math.min(GATEWAY_SHUTDOWN_TIMEOUT_MS, stop.timeoutMs - GATEWAY_SUPERVISOR_EXIT_MARGIN_MS);
retained?.timeoutMs ?? Math.min(GATEWAY_SHUTDOWN_TIMEOUT_MS, stop.timeoutMs - exitMarginMs);
const elapsedMs =
refresh && nativeStopBudget
? Math.max(0, Math.ceil(performance.now() - refresh.acceptedAtMs))
: 0;
const timeoutMs = Math.max(0, limitMs - elapsedMs);
const reserveMs = Math.min(GATEWAY_SHUTDOWN_RESERVE_MS, timeoutMs);
const reserveMs = resolveShutdownReserveMs(timeoutMs);
return {
nativeStopBudget,
timeoutMs,
@ -64,7 +117,10 @@ export async function resolveGatewayShutdownBudget(
},
log: (phase: "startup" | "shutdown") => {
logger.info(
`shutdown budget at ${phase}: drain=${Math.max(0, timeoutMs - GATEWAY_SHUTDOWN_RESERVE_MS)}ms shutdown=${timeoutMs}ms reserve=${reserveMs}ms exitMargin=${GATEWAY_SUPERVISOR_EXIT_MARGIN_MS}ms; source=${retained ? `startup shutdown budget=${retained.timeoutMs}ms` : `${stop.source}=${stop.timeoutMs}ms`}`,
// Report the drain and margin actually spent. Subtracting the unscaled
// reserve constant here understated a short budget's drain by the amount the
// scaled reserve gave back.
`shutdown budget at ${phase}: drain=${Math.max(0, timeoutMs - reserveMs)}ms shutdown=${timeoutMs}ms reserve=${reserveMs}ms exitMargin=${exitMarginMs}ms; source=${retained ? `startup shutdown budget=${retained.timeoutMs}ms` : `${stop.source}=${stop.timeoutMs}ms`}`,
);
},
};

View file

@ -0,0 +1,678 @@
// darwin launchd-supervised run-loop cases. These live beside run-loop.test.ts because
// the darwin stop budget reads the launchd job on every stop/restart request, so each
// case here has to state the deadline launchd is enforcing instead of letting the real
// reader spawn launchctl print, and run-loop.test.ts is at its line cap.
import { performance } from "node:perf_hooks";
import { afterEach, beforeEach, describe, expect, it, vi } from "vitest";
import { useAutoCleanupTempDirTracker } from "../../../test/helpers/temp-dir.js";
import type { HostedGatewayStop } from "../../daemon/hosted-stop.js";
import type { GatewayActiveWorkSnapshot } from "../../infra/gateway-active-work.js";
import type { GatewayBootLifecycleCompletion } from "../../infra/gateway-boot-lifecycle.js";
import type { GatewayRestartIntent } from "../../infra/restart-intent.js";
import { SUPERVISOR_HINT_ENV_VARS } from "../../infra/supervisor-markers.js";
import type { RuntimeEnv } from "../../runtime.js";
import { captureEnv, deleteTestEnvValue } from "../../test-utils/env.js";
import type { GatewayRestartSnapshot } from "../daemon-cli/restart-health.js";
import {
createActiveWorkSnapshot,
createCloseMock,
createRuntimeWithExitSignal,
createSignaledStart,
originalPlatformDescriptor,
type UpdateRespawnResultFixture,
setPlatform,
waitForStart,
withIsolatedSignals,
} from "./run-loop.test-support.js";
useAutoCleanupTempDirTracker(afterEach);
const { readCgroup } = vi.hoisted(() => ({ readCgroup: vi.fn() }));
const spawnProcess = vi.hoisted(() => vi.fn<typeof import("node:child_process").spawn>());
vi.mock("node:child_process", async () => {
const actual = await vi.importActual<typeof import("node:child_process")>("node:child_process");
const { mockNodeChildProcessModule } =
await import("../../gateway/server-methods/node-child-process.test-support.js");
if (!spawnProcess.getMockImplementation()) {
spawnProcess.mockImplementation(actual.spawn);
}
const mocked = await mockNodeChildProcessModule({});
vi.spyOn(mocked, "spawn").mockImplementation(spawnProcess);
return mocked;
});
vi.mock("node:fs/promises", async (original) => {
const actual = await original<typeof import("node:fs/promises")>();
// Foreground fixtures must not inherit the CI runner's systemd service or filesystem timing.
const readFile = (...args: Parameters<typeof actual.readFile>) =>
args[0] === "/proc/self/cgroup" ? readCgroup() : actual.readFile(...args);
return { ...actual, readFile, default: { ...actual, readFile } };
});
const systemctl = vi.fn(async () => ({
code: 0,
stdout: "LoadState=loaded\nTimeoutStopUSec=5min 30s",
stderr: "",
}));
vi.mock("../../daemon/systemd-exec.js", () => ({
execSystemctl: () => systemctl(),
execSystemctlUser: () => systemctl(),
}));
// darwin now refreshes the stop budget on every stop/restart request, and the real
// reader spawns launchctl print against three domains through the mocked
// node:child_process module. Keep launchctl out of the loop and let each case state
// the deadline launchd is enforcing. The reader reports stop: null when launchd is not
// running the stop, which leaves the platform-neutral policy in force.
const readLaunchdStopTimeout = vi.fn<
typeof import("../../infra/launchd-stop-timeout.js").readLaunchdStopTimeout
>(async () => ({ stop: null }));
vi.mock("../../infra/launchd-stop-timeout.js", () => ({
readLaunchdStopTimeout: (...args: Parameters<typeof readLaunchdStopTimeout>) =>
readLaunchdStopTimeout(...args),
}));
const acquireGatewayLock = vi.fn(async (_opts?: { port?: number }) => ({
release: vi.fn(async () => {}),
}));
const hostedStopExecute = vi.fn<HostedGatewayStop["execute"]>();
const hostedStopDispose = vi.fn<HostedGatewayStop["dispose"]>();
const hostedStopPrepare =
vi.fn<typeof import("../../daemon/hosted-stop.js").prepareHostedGatewayStop>();
vi.mock("../../daemon/hosted-stop.js", () => ({
prepareHostedGatewayStop: (...args: Parameters<typeof hostedStopPrepare>) =>
hostedStopPrepare(...args),
}));
const consumeGatewayRestartIntentPayloadSync = vi.fn<
() => { reason?: string; force?: boolean; waitMs?: number } | null
>(() => null);
const consumeGatewayRestartIntent = vi.fn<() => GatewayRestartIntent | null>(() => null);
type ManagedUpdateOwner = NonNullable<GatewayRestartIntent["successorOwner"]>;
const cancelManagedServiceUpdateHandoff = vi.fn<
(_identity: ManagedUpdateOwner) => Promise<false | "restored-in-process" | "restart-after-exit">
>(async () => "restored-in-process");
const claimManagedServiceUpdateHandoff = vi.fn((_identity: ManagedUpdateOwner) => true);
const isForegroundUpdateHandoff = vi.fn((_identity: ManagedUpdateOwner) => false);
const completeForegroundUpdateHandoffAfterClose =
vi.fn<
typeof import("../../infra/update-managed-service-handoff.js").completeForegroundUpdateHandoffAfterClose
>();
const captureForegroundUpdateHandoffStop =
vi.fn<
typeof import("../../infra/update-managed-service-handoff.js").captureForegroundUpdateHandoffStop
>();
const requestManagedServiceUpdateHandoffPark = vi.fn(async (_identity: ManagedUpdateOwner) => true);
const waitForSystemServiceUpdateHandoffs = vi.fn<() => Promise<void> | undefined>();
const commitManagedServiceUpdateHandoff = vi.fn(
async (_identity: ManagedUpdateOwner, _outcome?: "update" | "restore") => true,
);
const consumeGatewayRestartAuthorization = vi.fn(() => true);
const consumeGatewayRestartIntentSync = vi.fn(() => false);
const isGatewayRestartExternallyAllowed = vi.fn(() => false);
const markGatewayRestartHandled = vi.fn();
const peekGatewayRestartReason = vi.fn<() => string | undefined>(() => undefined);
const resetGatewayRestartStateForInProcessRestart = vi.fn();
const resetGatewaySuspendCoordinatorForLifecycleRestart = vi.fn();
const consumeGatewaySuspendHandoff =
vi.fn<typeof import("../../infra/gateway-suspend-coordinator.js").consumeGatewaySuspendHandoff>();
const disarmGatewaySuspendHandoff = vi.fn();
const rollbackGatewayRestartSignalAdmission = vi.fn();
const requestGatewayRestartWithSignalAdmission = vi.fn(() => ({ status: "emitted" as const }));
const writeGatewayRestartHandoffSync = vi.fn(
(
_opts: unknown,
): {
kind: "gateway-supervisor-restart-handoff";
version: 1;
intentId: string;
pid: number;
createdAt: number;
expiresAt: number;
source: "unknown";
restartKind: "full-process";
supervisorMode: "external";
} | null => ({
kind: "gateway-supervisor-restart-handoff",
version: 1,
intentId: "test-intent",
pid: process.pid,
createdAt: Date.now(),
expiresAt: Date.now() + 60_000,
source: "unknown",
restartKind: "full-process",
supervisorMode: "external",
}),
);
const scheduleGatewayRestart = vi.fn((_opts?: { delayMs?: number; reason?: string }) => ({
ok: true,
pid: process.pid,
signal: "SIGUSR2" as const,
delayMs: 0,
mode: "emit" as const,
coalesced: false,
cooldownMsApplied: 0,
}));
const idleActiveWorkSnapshot = createActiveWorkSnapshot();
const createGatewayActiveWorkSnapshot = vi.fn(() => idleActiveWorkSnapshot);
const waitForGatewayActiveWork = vi.fn(
async (
_timeoutMs?: number,
options?: { onSnapshot?: (snapshot: GatewayActiveWorkSnapshot) => void },
) => {
const snapshot = createGatewayActiveWorkSnapshot();
options?.onSnapshot?.(snapshot);
return { drained: snapshot.idle, snapshot };
},
);
const advanceCronActiveJobGeneration = vi.fn();
const resetCronActiveJobs = vi.fn();
const abortActiveCronTaskRuns = vi.fn((_reason?: string) => 0);
const retireActiveCronTaskRunTracking = vi.fn();
const waitForActiveCronTaskRuns = vi.fn(async (_timeoutMs?: number) => ({
drained: true,
active: 0,
}));
const waitForActiveCronJobs = vi.fn(async (_timeoutMs?: number) => ({
drained: true,
active: 0,
}));
const reloadTaskRuntimeStateFromStore = vi.fn();
const clearRuntimeConfigSnapshot = vi.fn();
const restartGatewayProcessWithFreshPid = vi.fn<
(_opts?: { env?: NodeJS.ProcessEnv }) => {
mode: "supervised" | "disabled" | "failed";
detail?: string;
exitCode?: number;
handoffSpawned?: Promise<boolean>;
}
>(() => ({ mode: "disabled" }));
const respawnGatewayProcessForUpdate = vi.fn<
(_opts?: { env?: NodeJS.ProcessEnv }) => UpdateRespawnResultFixture
>(() => ({ mode: "disabled", detail: "OPENCLAW_NO_RESPAWN" }));
const { killProcessTree } = vi.hoisted(() => ({ killProcessTree: vi.fn() }));
vi.mock("../../process/kill-tree.js", async (importOriginal) => ({
...(await importOriginal<typeof import("../../process/kill-tree.js")>()),
killProcessTree,
}));
const markUpdateRestartSentinelFailure = vi.fn<(reason: string) => Promise<null>>(async () => null);
const writeRestartSentinelIfUnchanged = vi.fn<
typeof import("../../infra/restart-sentinel.js").writeRestartSentinelIfUnchanged
>(async () => null);
const readRestartSentinelReadOnly =
vi.fn<typeof import("../../infra/restart-sentinel.js").readRestartSentinelReadOnly>();
const waitForGatewayHealthyRestart =
vi.fn<typeof import("../daemon-cli/restart-health.js").waitForGatewayHealthyRestart>();
vi.mock("../daemon-cli/restart-health.js", async (importOriginal) => ({
...(await importOriginal<typeof import("../daemon-cli/restart-health.js")>()),
waitForGatewayHealthyRestart: (...args: Parameters<typeof waitForGatewayHealthyRestart>) =>
waitForGatewayHealthyRestart(...args),
}));
const respawnHealth = (
overrides: Partial<GatewayRestartSnapshot> = {},
): GatewayRestartSnapshot => ({
runtime: { status: "running", pid: 7777 },
portUsage: { port: 18789, status: "busy", listeners: [{ pid: 7777 }], hints: [] },
healthy: true,
waitOutcome: "healthy",
staleGatewayPids: [],
...overrides,
});
const abortPendingChannelReloads = vi.fn();
const abortEmbeddedAgentRun = vi.fn(
(_sessionId?: string, _opts?: { mode?: "all" | "compacting"; reason?: "restart" }) => false,
);
const gatewayLog = {
debug: vi.fn(),
info: vi.fn(),
warn: vi.fn(),
error: vi.fn(),
};
const flushLogger = vi.fn(async () => {});
const writeDiagnosticStabilityBundleForFailureSync = vi.fn(() => ({
message: "stability bundle recorded",
}));
const hasManagedProviderLocalServices = vi.fn(() => false);
const stopManagedProviderLocalServices = vi.fn(async () => {});
const cancelShutdownHardExitWatchdog = vi.fn();
const armShutdownHardExitWatchdog = vi.fn(
(_params: { delayMs: number; onError: (error: unknown) => void }) => ({
cancel: cancelShutdownHardExitWatchdog,
}),
);
vi.mock("../../infra/gateway-lock.js", async (original) => ({
...(await original<typeof import("../../infra/gateway-lock.js")>()),
acquireGatewayLock: (opts?: { port?: number }) => acquireGatewayLock(opts),
}));
vi.mock("../../infra/restart.js", async (importOriginal) => {
const actual = await importOriginal<typeof import("../../infra/restart.js")>();
return {
...actual,
consumeGatewayRestartIntent: () => consumeGatewayRestartIntent(),
consumeGatewayRestartAuthorization: () => consumeGatewayRestartAuthorization(),
isGatewayRestartExternallyAllowed: () => isGatewayRestartExternallyAllowed(),
markGatewayRestartHandled: () => markGatewayRestartHandled(),
peekGatewayRestartReason: () => peekGatewayRestartReason(),
resetGatewayRestartStateForInProcessRestart: () =>
resetGatewayRestartStateForInProcessRestart(),
rollbackGatewayRestartSignalAdmission: () => rollbackGatewayRestartSignalAdmission(),
requestGatewayRestartWithSignalAdmission,
scheduleGatewayRestart: (opts?: { delayMs?: number; reason?: string }) =>
scheduleGatewayRestart(opts),
};
});
vi.mock("../../infra/restart-intent.js", () => ({
consumeGatewayRestartIntentPayloadSync: () => consumeGatewayRestartIntentPayloadSync(),
consumeGatewayRestartIntentSync: () => consumeGatewayRestartIntentSync(),
}));
vi.mock("../../infra/update-managed-service-handoff.js", () => ({
waitForSystemServiceUpdateHandoffs: () => waitForSystemServiceUpdateHandoffs(),
captureForegroundUpdateHandoffStop: (
params: Parameters<typeof captureForegroundUpdateHandoffStop>[0],
) => captureForegroundUpdateHandoffStop(params),
isForegroundUpdateHandoff: (identity: ManagedUpdateOwner) => isForegroundUpdateHandoff(identity),
completeForegroundUpdateHandoffAfterClose: (identity: ManagedUpdateOwner) =>
completeForegroundUpdateHandoffAfterClose(identity),
cancelManagedServiceUpdateHandoff: (identity: ManagedUpdateOwner) =>
cancelManagedServiceUpdateHandoff(identity),
claimManagedServiceUpdateHandoff: (identity: ManagedUpdateOwner) =>
claimManagedServiceUpdateHandoff(identity),
requestManagedServiceUpdateHandoffPark: (identity: ManagedUpdateOwner) =>
requestManagedServiceUpdateHandoffPark(identity),
commitManagedServiceUpdateHandoff: (
identity: ManagedUpdateOwner,
outcome?: "update" | "restore",
) => commitManagedServiceUpdateHandoff(identity, outcome),
}));
vi.mock("../../infra/gateway-suspend-coordinator.js", () => ({
consumeGatewaySuspendHandoff: (...args: Parameters<typeof consumeGatewaySuspendHandoff>) =>
consumeGatewaySuspendHandoff(...args),
disarmGatewaySuspendHandoff: (...args: unknown[]) => disarmGatewaySuspendHandoff(...args),
resetGatewaySuspendCoordinatorForLifecycleRestart: () =>
resetGatewaySuspendCoordinatorForLifecycleRestart(),
}));
vi.mock("../../infra/process-respawn.js", async (importOriginal) => ({
...(await importOriginal<typeof import("../../infra/process-respawn.js")>()),
respawnGatewayProcessForUpdate: (opts?: { env?: NodeJS.ProcessEnv }) =>
respawnGatewayProcessForUpdate(opts),
restartGatewayProcessWithFreshPid: (opts?: { env?: NodeJS.ProcessEnv }) =>
restartGatewayProcessWithFreshPid(opts),
}));
vi.mock("../../infra/restart-sentinel.js", () => ({
readRestartSentinelReadOnly: () => readRestartSentinelReadOnly(),
markUpdateRestartSentinelFailure: (reason: string) => markUpdateRestartSentinelFailure(reason),
writeRestartSentinelIfUnchanged: (...args: Parameters<typeof writeRestartSentinelIfUnchanged>) =>
writeRestartSentinelIfUnchanged(...args),
}));
vi.mock("../../infra/restart-handoff.js", () => ({
writeGatewayRestartHandoffSync: (opts: unknown) => writeGatewayRestartHandoffSync(opts),
}));
vi.mock("../../infra/gateway-active-work.js", () => ({
createGatewayActiveWorkSnapshot: () => createGatewayActiveWorkSnapshot(),
waitForGatewayActiveWork: (
timeoutMs?: number,
options?: { onSnapshot?: (snapshot: GatewayActiveWorkSnapshot) => void },
) => waitForGatewayActiveWork(timeoutMs, options),
}));
vi.mock("../../cron/active-jobs.js", () => ({
advanceCronActiveJobGeneration: () => advanceCronActiveJobGeneration(),
resetCronActiveJobs: () => resetCronActiveJobs(),
waitForActiveCronJobs: (timeoutMs: number) => waitForActiveCronJobs(timeoutMs),
}));
vi.mock("../../cron/service/active-run-cancellation.js", () => ({
abortActiveCronTaskRuns: (reason?: string) => abortActiveCronTaskRuns(reason),
retireActiveCronTaskRunTracking: () => retireActiveCronTaskRunTracking(),
waitForActiveCronTaskRuns: (timeoutMs: number) => waitForActiveCronTaskRuns(timeoutMs),
}));
vi.mock("../../tasks/runtime-internal.js", () => ({
reloadTaskRuntimeStateFromStore: () => reloadTaskRuntimeStateFromStore(),
}));
vi.mock("../../config/runtime-snapshot.js", () => ({
clearRuntimeConfigSnapshot: () => clearRuntimeConfigSnapshot(),
getRuntimeConfigSourceSnapshot: () => null,
registerRuntimeConfigSnapshotPreparer: vi.fn(),
}));
vi.mock("../../agents/embedded-agent-runner/runs.js", () => ({
abortEmbeddedAgentRun: (
sessionId?: string,
opts?: { mode?: "all" | "compacting"; reason?: "restart" },
) => abortEmbeddedAgentRun(sessionId, opts),
}));
vi.mock("../../logging/subsystem.js", () => ({
createSubsystemLogger: () => gatewayLog,
}));
vi.mock("../../logging/logger.js", () => ({
flushLogger: () => flushLogger(),
}));
vi.mock("../../logging/diagnostic-stability-bundle.js", () => ({
writeDiagnosticStabilityBundleForFailureSync,
}));
vi.mock("../../agents/provider-runtime-lifecycle.js", () => ({
hasManagedProviderLocalServices: () => hasManagedProviderLocalServices(),
}));
vi.mock("../../agents/provider-local-service.js", () => ({
stopManagedProviderLocalServices: () => stopManagedProviderLocalServices(),
}));
vi.mock("../../gateway/server-reload-generation.js", () => ({
abortPendingChannelReloads: () => abortPendingChannelReloads(),
}));
vi.mock("./shutdown-hard-exit.js", () => ({
armShutdownHardExitWatchdog: (params: { delayMs: number; onError: (error: unknown) => void }) =>
armShutdownHardExitWatchdog(params),
}));
async function runLoopWithStart(params: {
start: ReturnType<typeof vi.fn>;
runtime: RuntimeEnv;
ownsProcessLifecycle?: boolean;
lockPort?: number;
healthHost?: string;
beginBoot?: (startedAtMs: number) => void | Promise<void>;
completeBoot?: (completion: GatewayBootLifecycleCompletion) => void;
}) {
vi.resetModules();
const { runGatewayLoop } = await import("./run-loop.js");
const loopPromise = runGatewayLoop({
start: params.start as unknown as Parameters<typeof runGatewayLoop>[0]["start"],
runtime: params.runtime,
ownsProcessLifecycle: params.ownsProcessLifecycle,
lockPort: params.lockPort,
healthHost: params.healthHost,
beginBoot: params.beginBoot,
completeBoot: params.completeBoot,
});
return { loopPromise };
}
async function createSignaledLoopHarness(exitCallOrder?: string[], ownsProcessLifecycle = false) {
const close = createCloseMock();
const { start, started } = createSignaledStart(close);
const { runtime, exited } = createRuntimeWithExitSignal(exitCallOrder);
const { loopPromise } = await runLoopWithStart({ start, runtime, ownsProcessLifecycle });
await waitForStart(started);
return { close, start, runtime, exited, loopPromise };
}
function expectRestartHandoffCall(expected: {
restartKind: "full-process" | "update-process";
reason: string | undefined;
supervisorMode: "external" | "launchd";
}) {
expect(writeGatewayRestartHandoffSync).toHaveBeenCalledTimes(1);
const [handoff] = writeGatewayRestartHandoffSync.mock.calls[0] ?? [];
if (!handoff || typeof handoff !== "object" || Array.isArray(handoff)) {
throw new Error("expected restart handoff options object");
}
const processInstanceId = (handoff as { processInstanceId?: unknown }).processInstanceId;
expect(typeof processInstanceId).toBe("string");
if (typeof processInstanceId !== "string") {
throw new Error("expected restart handoff processInstanceId string");
}
expect(processInstanceId).toMatch(
/^[0-9a-f]{8}-[0-9a-f]{4}-[1-5][0-9a-f]{3}-[89ab][0-9a-f]{3}-[0-9a-f]{12}$/,
);
expect(handoff).toEqual({
...expected,
processInstanceId,
});
}
let gatewayWorkAdmissionActual: typeof import("../../process/gateway-work-admission.js");
let supervisorEnvSnapshot: ReturnType<typeof captureEnv> | undefined;
beforeEach(async () => {
vi.useRealTimers();
vi.clearAllMocks();
spawnProcess
.mockReset()
.mockImplementation(
(await vi.importActual<typeof import("node:child_process")>("node:child_process")).spawn,
);
acquireGatewayLock.mockReset().mockImplementation(async () => ({
release: vi.fn(async () => {}),
}));
setPlatform("linux");
readCgroup.mockReset().mockResolvedValue("0::/\n");
systemctl.mockReset().mockResolvedValue({
code: 0,
stdout: "LoadState=loaded\nTimeoutStopUSec=5min 30s",
stderr: "",
});
// mockReset also drops any one-shot launchd deadline a previous case left queued.
readLaunchdStopTimeout.mockReset().mockResolvedValue({ stop: null });
hostedStopExecute.mockReset().mockResolvedValue({ outcome: "accepted" });
hostedStopDispose.mockReset().mockResolvedValue(undefined);
hostedStopPrepare.mockReset().mockImplementation(async (_owner, assertCurrent) => {
assertCurrent();
return { execute: hostedStopExecute, dispose: hostedStopDispose };
});
supervisorEnvSnapshot = captureEnv([...SUPERVISOR_HINT_ENV_VARS, "OPENCLAW_NO_RESPAWN"]);
for (const key of [...SUPERVISOR_HINT_ENV_VARS, "OPENCLAW_NO_RESPAWN"]) {
deleteTestEnvValue(key);
}
// clearAllMocks preserves queued one-shot results. A skipped lifecycle branch
// must not shift a stale supervisor or respawn decision into the next case.
consumeGatewayRestartIntent.mockReset();
consumeGatewayRestartIntentPayloadSync.mockReset().mockReturnValue(null);
consumeGatewaySuspendHandoff.mockReset().mockReturnValue({ ok: true, value: false });
disarmGatewaySuspendHandoff.mockClear();
consumeGatewayRestartIntent.mockReturnValue(null);
peekGatewayRestartReason.mockReset();
peekGatewayRestartReason.mockReturnValue(undefined);
restartGatewayProcessWithFreshPid.mockReset();
restartGatewayProcessWithFreshPid.mockReturnValue({ mode: "disabled" });
respawnGatewayProcessForUpdate.mockReset();
waitForGatewayHealthyRestart.mockReset().mockResolvedValue(respawnHealth());
writeRestartSentinelIfUnchanged.mockReset().mockResolvedValue(null);
readRestartSentinelReadOnly.mockReset().mockResolvedValue(null);
respawnGatewayProcessForUpdate.mockReturnValue({
mode: "disabled",
detail: "OPENCLAW_NO_RESPAWN",
});
hasManagedProviderLocalServices.mockReset();
hasManagedProviderLocalServices.mockReturnValue(false);
stopManagedProviderLocalServices.mockReset();
stopManagedProviderLocalServices.mockResolvedValue(undefined);
gatewayWorkAdmissionActual = await vi.importActual("../../process/gateway-work-admission.js");
gatewayWorkAdmissionActual.resetGatewayWorkAdmission();
createGatewayActiveWorkSnapshot.mockReset();
createGatewayActiveWorkSnapshot.mockReturnValue(idleActiveWorkSnapshot);
waitForGatewayActiveWork.mockReset();
waitForGatewayActiveWork.mockImplementation(async (_timeoutMs, options) => {
const snapshot = createGatewayActiveWorkSnapshot();
options?.onSnapshot?.(snapshot);
return { drained: snapshot.idle, snapshot };
});
cancelManagedServiceUpdateHandoff.mockReset();
cancelManagedServiceUpdateHandoff.mockResolvedValue("restored-in-process");
claimManagedServiceUpdateHandoff.mockReset();
claimManagedServiceUpdateHandoff.mockReturnValue(true);
isForegroundUpdateHandoff.mockReset().mockReturnValue(false);
completeForegroundUpdateHandoffAfterClose.mockReset().mockResolvedValue({ respawn: true });
captureForegroundUpdateHandoffStop.mockReset().mockReturnValue(undefined);
requestManagedServiceUpdateHandoffPark.mockReset();
requestManagedServiceUpdateHandoffPark.mockResolvedValue(true);
waitForSystemServiceUpdateHandoffs.mockReset().mockReturnValue(undefined);
commitManagedServiceUpdateHandoff.mockReset();
commitManagedServiceUpdateHandoff.mockResolvedValue(true);
});
afterEach(() => {
supervisorEnvSnapshot?.restore();
supervisorEnvSnapshot = undefined;
vi.useRealTimers();
if (originalPlatformDescriptor) {
Object.defineProperty(process, "platform", originalPlatformDescriptor);
}
});
describe("runGatewayLoop darwin launchd supervision", () => {
it("waits briefly before exiting on launchd supervised restart", async () => {
vi.clearAllMocks();
peekGatewayRestartReason.mockReturnValue(undefined);
try {
setPlatform("darwin");
process.env.OPENCLAW_LAUNCHD_LABEL = "ai.openclaw.gateway";
restartGatewayProcessWithFreshPid.mockReturnValueOnce({
mode: "supervised",
handoffSpawned: Promise.resolve(true),
});
await withIsolatedSignals(async ({ captureSignal }) => {
const { runtime, exited } = await createSignaledLoopHarness();
const restartSignal = captureSignal("SIGUSR2");
vi.useFakeTimers();
restartSignal();
await vi.advanceTimersByTimeAsync(1499);
expect(runtime.exit).not.toHaveBeenCalled();
await vi.advanceTimersByTimeAsync(1);
await expect(exited).resolves.toBe(0);
expect(runtime.exit).toHaveBeenCalledWith(0);
expectRestartHandoffCall({
restartKind: "full-process",
reason: undefined,
supervisorMode: "launchd",
});
});
} finally {
vi.useRealTimers();
delete process.env.OPENCLAW_LAUNCHD_LABEL;
if (originalPlatformDescriptor) {
Object.defineProperty(process, "platform", originalPlatformDescriptor);
}
}
});
it("falls back in-process when the launchd restart handoff fails to spawn", async () => {
vi.clearAllMocks();
peekGatewayRestartReason.mockReturnValue(undefined);
try {
setPlatform("darwin");
process.env.OPENCLAW_LAUNCHD_LABEL = "ai.openclaw.gateway";
restartGatewayProcessWithFreshPid.mockReturnValueOnce({
mode: "supervised",
handoffSpawned: Promise.resolve(false),
});
await withIsolatedSignals(async ({ captureSignal }) => {
const { start, runtime, exited } = await createSignaledLoopHarness();
const restartSignal = captureSignal("SIGUSR2");
const sigint = captureSignal("SIGINT");
vi.useFakeTimers();
restartSignal();
await vi.advanceTimersByTimeAsync(1500);
expect(start).toHaveBeenCalledTimes(2);
expect(runtime.exit).not.toHaveBeenCalled();
expect(acquireGatewayLock).toHaveBeenCalledTimes(2);
expect(gatewayLog.warn).toHaveBeenCalledWith(
"launchd restart handoff failed to spawn; falling back to in-process restart",
);
sigint();
await expect(exited).resolves.toBe(0);
});
} finally {
vi.useRealTimers();
delete process.env.OPENCLAW_LAUNCHD_LABEL;
if (originalPlatformDescriptor) {
Object.defineProperty(process, "platform", originalPlatformDescriptor);
}
}
});
it("leaves the successor to launchd after a SIGTERM restart intent", async () => {
vi.clearAllMocks();
consumeGatewayRestartIntentPayloadSync.mockReturnValueOnce({ reason: "gateway.restart" });
setPlatform("darwin");
process.env.OPENCLAW_LAUNCHD_LABEL = "ai.openclaw.gateway";
restartGatewayProcessWithFreshPid.mockReturnValueOnce({
mode: "supervised",
handoffSpawned: Promise.resolve(true),
});
await withIsolatedSignals(async ({ captureSignal }) => {
const { start, exited } = await createSignaledLoopHarness();
captureSignal("SIGTERM")();
await expect(exited).resolves.toBe(0);
expect(start).toHaveBeenCalledOnce();
expect(restartGatewayProcessWithFreshPid).not.toHaveBeenCalled();
expect(respawnGatewayProcessForUpdate).not.toHaveBeenCalled();
});
});
// The stop budget refresh is the reason darwin reads the launchd job at all, so prove
// the job deadline it reports is what bounds the stop and arms the force exit.
it("bounds a launchd-supervised stop on the deadline the printed job reports", async () => {
vi.clearAllMocks();
try {
setPlatform("darwin");
process.env.OPENCLAW_LAUNCHD_LABEL = "ai.openclaw.gateway";
readLaunchdStopTimeout.mockResolvedValue({
stop: { timeoutMs: 30_000, source: "launchd system/ai.openclaw.gateway exit timeout" },
});
hasManagedProviderLocalServices.mockReturnValue(true);
stopManagedProviderLocalServices.mockReturnValue(new Promise<void>(() => {}));
await withIsolatedSignals(async ({ captureSignal }) => {
const { close, runtime } = await createSignaledLoopHarness();
vi.useFakeTimers();
const clock = vi.spyOn(performance, "now").mockImplementation(() => Date.now());
try {
captureSignal("SIGTERM")();
await vi.advanceTimersByTimeAsync(24_999);
expect(close).toHaveBeenCalledOnce();
expect(stopManagedProviderLocalServices).toHaveBeenCalledOnce();
expect(runtime.exit).not.toHaveBeenCalled();
await vi.advanceTimersByTimeAsync(1);
// launchd owns the successor, so an abandoned drain still exits 0 for the job.
expect(runtime.exit).toHaveBeenCalledExactlyOnceWith(0);
expect(writeDiagnosticStabilityBundleForFailureSync).toHaveBeenCalledWith(
"gateway.stop_shutdown_timeout",
undefined,
);
expect(gatewayLog.info).toHaveBeenCalledWith(
"shutdown budget at shutdown: drain=15000ms shutdown=25000ms reserve=10000ms exitMargin=5000ms; source=launchd system/ai.openclaw.gateway exit timeout=30000ms",
);
} finally {
clock.mockRestore();
vi.clearAllTimers();
vi.useRealTimers();
}
});
} finally {
delete process.env.OPENCLAW_LAUNCHD_LABEL;
if (originalPlatformDescriptor) {
Object.defineProperty(process, "platform", originalPlatformDescriptor);
}
}
});
});

View file

@ -79,9 +79,23 @@ it
expect(exit, output).toEqual([0, null]);
expect(elapsed, output).toBeLessThan(stopTimeoutMs);
expect(output).toContain("process proof: acquisition-cancelled");
// The fixture's launchd label is synthetic, so the per-stop re-inspection cannot
// find a job to read and the startup budget is retained instead. Both
// attributions prove the same thing this case is guarding: the shutdown budget
// came from the 20 second native stop timeout and not from the platform-neutral
// policy, which would report a 330000ms source and a far longer deadline.
expect(output).toMatch(
new RegExp(`shutdown budget at shutdown:.*source=.*=${stopTimeoutMs}ms`),
new RegExp(
`shutdown budget at shutdown:.*source=(?:.*=${stopTimeoutMs}ms|startup shutdown budget=${shutdownTimeoutMs}ms)`,
),
);
if (process.platform === "darwin") {
// The label resolves to no real job, so this cannot assert a successful read.
// What it does assert is that the darwin probe ran inside a real spawned
// Gateway on a real stop: only the launchd reader emits this, and reverting the
// darwin dispatch removes it.
expect(output).toContain("Unable to inspect the launchd job");
}
if (mode === "cooperative") {
expect(output).toContain("process proof: acquisition-joined");
expect(output).not.toContain("shutdown deadline reached");

View file

@ -78,6 +78,19 @@ vi.mock("../../daemon/systemd-exec.js", () => ({
execSystemctlUser: () => systemctl(),
}));
// darwin now refreshes the stop budget on every stop/restart request, and the real
// reader spawns launchctl print against three domains through the mocked
// node:child_process module. Keep launchctl out of the loop and let each case state
// the deadline launchd is enforcing. The reader reports stop: null when launchd is not
// running the stop, which leaves the platform-neutral policy in force.
const readLaunchdStopTimeout = vi.fn<
typeof import("../../infra/launchd-stop-timeout.js").readLaunchdStopTimeout
>(async () => ({ stop: null }));
vi.mock("../../infra/launchd-stop-timeout.js", () => ({
readLaunchdStopTimeout: (...args: Parameters<typeof readLaunchdStopTimeout>) =>
readLaunchdStopTimeout(...args),
}));
const acquireGatewayLock = vi.fn(async (_opts?: { port?: number }) => ({
release: vi.fn(async () => {}),
}));
@ -468,6 +481,8 @@ beforeEach(async () => {
stdout: "LoadState=loaded\nTimeoutStopUSec=5min 30s",
stderr: "",
});
// mockReset also drops any one-shot launchd deadline a previous case left queued.
readLaunchdStopTimeout.mockReset().mockResolvedValue({ stop: null });
hostedStopExecute.mockReset().mockResolvedValue({ outcome: "accepted" });
hostedStopDispose.mockReset().mockResolvedValue(undefined);
hostedStopPrepare.mockReset().mockImplementation(async (_owner, assertCurrent) => {
@ -2794,103 +2809,6 @@ describe("runGatewayLoop", () => {
}
});
it("waits briefly before exiting on launchd supervised restart", async () => {
vi.clearAllMocks();
peekGatewayRestartReason.mockReturnValue(undefined);
try {
setPlatform("darwin");
process.env.OPENCLAW_LAUNCHD_LABEL = "ai.openclaw.gateway";
restartGatewayProcessWithFreshPid.mockReturnValueOnce({
mode: "supervised",
handoffSpawned: Promise.resolve(true),
});
await withIsolatedSignals(async ({ captureSignal }) => {
const { runtime, exited } = await createSignaledLoopHarness();
const restartSignal = captureSignal("SIGUSR2");
vi.useFakeTimers();
restartSignal();
await vi.advanceTimersByTimeAsync(1499);
expect(runtime.exit).not.toHaveBeenCalled();
await vi.advanceTimersByTimeAsync(1);
await expect(exited).resolves.toBe(0);
expect(runtime.exit).toHaveBeenCalledWith(0);
expectRestartHandoffCall({
restartKind: "full-process",
reason: undefined,
supervisorMode: "launchd",
});
});
} finally {
vi.useRealTimers();
delete process.env.OPENCLAW_LAUNCHD_LABEL;
if (originalPlatformDescriptor) {
Object.defineProperty(process, "platform", originalPlatformDescriptor);
}
}
});
it("falls back in-process when the launchd restart handoff fails to spawn", async () => {
vi.clearAllMocks();
peekGatewayRestartReason.mockReturnValue(undefined);
try {
setPlatform("darwin");
process.env.OPENCLAW_LAUNCHD_LABEL = "ai.openclaw.gateway";
restartGatewayProcessWithFreshPid.mockReturnValueOnce({
mode: "supervised",
handoffSpawned: Promise.resolve(false),
});
await withIsolatedSignals(async ({ captureSignal }) => {
const { start, runtime, exited } = await createSignaledLoopHarness();
const restartSignal = captureSignal("SIGUSR2");
const sigint = captureSignal("SIGINT");
vi.useFakeTimers();
restartSignal();
await vi.advanceTimersByTimeAsync(1500);
expect(start).toHaveBeenCalledTimes(2);
expect(runtime.exit).not.toHaveBeenCalled();
expect(acquireGatewayLock).toHaveBeenCalledTimes(2);
expect(gatewayLog.warn).toHaveBeenCalledWith(
"launchd restart handoff failed to spawn; falling back to in-process restart",
);
sigint();
await expect(exited).resolves.toBe(0);
});
} finally {
vi.useRealTimers();
delete process.env.OPENCLAW_LAUNCHD_LABEL;
if (originalPlatformDescriptor) {
Object.defineProperty(process, "platform", originalPlatformDescriptor);
}
}
});
it("leaves the successor to launchd after a SIGTERM restart intent", async () => {
vi.clearAllMocks();
consumeGatewayRestartIntentPayloadSync.mockReturnValueOnce({ reason: "gateway.restart" });
setPlatform("darwin");
process.env.OPENCLAW_LAUNCHD_LABEL = "ai.openclaw.gateway";
restartGatewayProcessWithFreshPid.mockReturnValueOnce({
mode: "supervised",
handoffSpawned: Promise.resolve(true),
});
await withIsolatedSignals(async ({ captureSignal }) => {
const { start, exited } = await createSignaledLoopHarness();
captureSignal("SIGTERM")();
await expect(exited).resolves.toBe(0);
expect(start).toHaveBeenCalledOnce();
expect(restartGatewayProcessWithFreshPid).not.toHaveBeenCalled();
expect(respawnGatewayProcessForUpdate).not.toHaveBeenCalled();
});
});
it("records external ownership even when native supervisor markers are inherited", async () => {
vi.clearAllMocks();
peekGatewayRestartReason.mockReturnValue(undefined);

View file

@ -609,7 +609,7 @@ export async function runGatewayLoop(params: {
return timer;
};
const timer = arm(startupBudget.timeoutMs);
if (process.platform === "linux") {
if (process.platform === "linux" || process.platform === "darwin") {
void resolveGatewayShutdownBudget(supervisorMode, gatewayLog, {
previous: startupBudget,
acceptedAtMs: pendingRequest.acceptedAtMs,
@ -761,7 +761,7 @@ export async function runGatewayLoop(params: {
}
const completion = (async () => {
if (process.platform === "linux") {
if (process.platform === "linux" || process.platform === "darwin") {
if (budget.nativeStopBudget && !getManagedUpdateOwner()) {
armForceExitTimer(budget.timeoutMs);
}

View file

@ -4,5 +4,9 @@ export {
GATEWAY_SUPERVISOR_EXIT_MARGIN_MS,
GATEWAY_SHUTDOWN_TIMEOUT_MS,
GATEWAY_SERVICE_STOP_TIMEOUT_MS,
isRespawnedByLauncher,
LAUNCH_AGENT_EXIT_TIMEOUT_SECONDS,
resolveLauncherStopTimeoutMs,
resolveShutdownReserveMs,
resolveSupervisorExitMarginMs,
} from "../../gateway-shutdown-budget.mjs";

View file

@ -0,0 +1,376 @@
import { afterEach, beforeEach, describe, expect, it, vi } from "vitest";
import { isRespawnedByLauncher } from "./gateway-shutdown-budget.js";
import { readLaunchdStopTimeout } from "./launchd-stop-timeout.js";
const { execLaunchctl } = vi.hoisted(() => ({ execLaunchctl: vi.fn() }));
vi.mock("../daemon/launchd-exec.js", async (importOriginal) => ({
...(await importOriginal<typeof import("../daemon/launchd-exec.js")>()),
execLaunchctl,
}));
// The authoritative list is module-local beside the deadline it authorises in
// gateway-shutdown-budget.mjs, so it is restated here and tied back to the
// implementation by the recognition cases at the end of this file.
const RESPAWN_MARKERS = [
"OPENCLAW_NODE_UPDATE_RESPAWNED",
"OPENCLAW_COMPILE_CACHE_DISABLED_RESPAWNED",
"OPENCLAW_PACKAGED_COMPILE_CACHE_RESPAWNED",
] as const;
const LAUNCHD_ENV = { XPC_SERVICE_NAME: "ai.openclaw.gateway" };
// The service layout the recovery launcher bounds by the LaunchAgent exit timeout:
// launchd names the job in XPC_SERVICE_NAME and the handoff carries the label, which
// is the pair the launcher branches on and a respawned child inherits unchanged.
const SERVICE_ENV = { ...LAUNCHD_ENV, OPENCLAW_LAUNCHD_LABEL: "ai.openclaw.gateway" };
// Adds one of the markers a respawning launcher stamps on its child, which is what
// separates a Gateway that launcher started from an unrelated parent.
const RESPAWNED_SERVICE_ENV = { ...SERVICE_ENV, OPENCLAW_NODE_UPDATE_RESPAWNED: "1" };
const result = (stdout: string) => ({ code: 0, stdout, stderr: "", termination: "exit" });
/**
* Shaped like real `launchctl print` output rather than a bare field list: the
* job's own `state` at one tab, then coalition blocks carrying their own
* `state = active` at two tabs, then `job state`. A live Gateway LaunchDaemon
* prints `state` three times in exactly this arrangement.
*/
const printed = (state: string, fields: string) =>
result(
`system/ai.openclaw.gateway = {\n\tactive count = 1\n\ttype = LaunchDaemon\n\tstate = ${state}\n\n${fields}\tresource coalition = {\n\t\tID = 18110\n\t\tstate = active\n\t}\n\n\tjetsam coalition = {\n\t\tID = 18111\n\t\tstate = active\n\t}\n\n\tjob state = running\n}\n`,
);
const stopping = (fields: string) => printed("SIGTERMed", fields);
beforeEach(() => {
vi.stubGlobal("process", {
...process,
platform: "darwin",
pid: 4242,
ppid: 4241,
getuid: () => 501,
// The reader only runs inside a serving Gateway, and reconstructing an
// undeclared launcher's timer replays the argv branch the launcher took.
argv: ["/usr/local/bin/node", "/usr/local/bin/openclaw", "gateway", "run"],
});
execLaunchctl.mockReset();
});
afterEach(() => vi.unstubAllGlobals());
describe("launchd stop timeout reads the job launchd is stopping", () => {
it("uses the operator job's effective exit timeout, not the template constant", async () => {
execLaunchctl.mockResolvedValue(
stopping("\tminimum runtime = 10\n\texit timeout = 5\n\tpid = 4242\n"),
);
await expect(readLaunchdStopTimeout(LAUNCHD_ENV)).resolves.toEqual({
stop: { timeoutMs: 5_000, source: "launchd system/ai.openclaw.gateway exit timeout" },
});
expect(execLaunchctl).toHaveBeenCalledExactlyOnceWith(
["print", "system/ai.openclaw.gateway"],
2_000,
);
});
// The shared key-value parser keeps the LAST occurrence of a repeated key, and
// `launchctl print` repeats `state` inside coalition blocks, so asking it for
// `state` answers `active` and never sees the job at all. This case fails if
// the reader ever goes back to that parser for the job's state.
it("reads the job's own state, not a nested coalition's", async () => {
execLaunchctl.mockResolvedValue(stopping("\texit timeout = 47\n\tpid = 4242\n"));
await expect(readLaunchdStopTimeout(LAUNCHD_ENV)).resolves.toEqual({
stop: { timeoutMs: 47_000, source: "launchd system/ai.openclaw.gateway exit timeout" },
});
});
// Measured on macOS 27 with ExitTimeOut 47: a plain `kill -TERM` left the job
// printing `state = running` while the process handled the signal, and it was
// still alive 85 seconds later. launchd never started its clock, so its
// deadline bounds nothing and the caller keeps the budget it already had.
it("declines the deadline when launchd is not the one stopping the job", async () => {
execLaunchctl.mockResolvedValue(printed("running", "\texit timeout = 5\n\tpid = 4242\n"));
// No deadline and nothing to warn about: this is the ordinary shape of a stop
// that some other sender delivered.
await expect(readLaunchdStopTimeout(LAUNCHD_ENV)).resolves.toEqual({ stop: null });
// Our job was found in the first domain, so there is nothing left to search.
expect(execLaunchctl).toHaveBeenCalledTimes(1);
});
it("preserves direct-child drain when a marked launcher has not started its timer", async () => {
execLaunchctl.mockResolvedValue(printed("running", "\texit timeout = 60\n\tpid = 4241\n"));
await expect(readLaunchdStopTimeout(RESPAWNED_SERVICE_ENV)).resolves.toEqual({ stop: null });
});
it("keeps launchd's unlimited exit timeout distinct from a missing one", async () => {
execLaunchctl.mockResolvedValue(stopping("\texit timeout = 0\n\tpid = 4242\n"));
await expect(readLaunchdStopTimeout(LAUNCHD_ENV)).resolves.toEqual({
stop: {
timeoutMs: Infinity,
source: "launchd system/ai.openclaw.gateway unlimited exit timeout",
},
});
});
it("caps an unlimited launchd stop at the launcher's finite timer", async () => {
execLaunchctl.mockResolvedValue(stopping("\texit timeout = 0\n\tpid = 4241\n"));
await expect(readLaunchdStopTimeout(RESPAWNED_SERVICE_ENV)).resolves.toEqual({
stop: {
timeoutMs: 19_000,
source:
"launchd system/ai.openclaw.gateway unlimited exit timeout capped at the launcher's 19000ms stop timer",
},
});
});
// An empty value is deliberately not in this table: there is no state to fail to
// recognise, so it belongs with the missing-state warning below.
it.each(["waiting", "exited", "not running", "SIGTERM", "sigtermed"])(
"treats the unrecognised state %j as not stopping",
async (state) => {
execLaunchctl.mockResolvedValue(printed(state, "\texit timeout = 5\n\tpid = 4242\n"));
await expect(readLaunchdStopTimeout(LAUNCHD_ENV)).resolves.toEqual({ stop: null });
},
);
it("accepts any signal launchd reports having delivered", async () => {
execLaunchctl.mockResolvedValue(printed("SIGKILLed", "\texit timeout = 9\n\tpid = 4242\n"));
await expect(readLaunchdStopTimeout(LAUNCHD_ENV)).resolves.toEqual({
stop: { timeoutMs: 9_000, source: "launchd system/ai.openclaw.gateway exit timeout" },
});
});
// An absent `state` is indistinguishable from a job launchd is not stopping, so the
// budget would silently revert to the platform-neutral policy on a macOS that printed
// this block differently. The deadline parsed from the same block is what makes the
// case reportable, and the warning is what keeps it from being silent. An empty value
// reaches the same place as a missing line, because neither yields a state to read.
it.each([
[
"no state line",
`system/ai.openclaw.gateway = {\n\tactive count = 1\n\ttype = LaunchDaemon\n\n\texit timeout = 47\n\tpid = 4242\n\tjob state = running\n}\n`,
],
["an empty state value", printed("", "\texit timeout = 47\n\tpid = 4242\n").stdout],
])("warns when the job printed a deadline but %s", async (_label, stdout) => {
execLaunchctl.mockResolvedValue(result(stdout));
const read = await readLaunchdStopTimeout(LAUNCHD_ENV);
expect(read.stop).toBeNull();
expect(read.warning).toBe(
"launchd system/ai.openclaw.gateway printed an exit timeout but no job state, so it is treated as not stopping and the Gateway stop policy is kept. Check the running job with launchctl print.",
);
});
// A job with neither field is an ordinary not-stopping answer and must stay quiet,
// otherwise every in-process restart of a launchd-owned Gateway would warn.
it("stays silent when the job is simply not stopping", async () => {
execLaunchctl.mockResolvedValue(printed("running", "\texit timeout = 47\n\tpid = 4242\n"));
await expect(readLaunchdStopTimeout(LAUNCHD_ENV)).resolves.toEqual({ stop: null });
});
it("falls back to the gui domain when the job is not a LaunchDaemon", async () => {
execLaunchctl
.mockResolvedValueOnce({
code: 113,
stdout: "",
stderr: "Could not find service",
termination: "exit",
})
.mockResolvedValueOnce(stopping("\texit timeout = 20\n\tpid = 4242\n"));
await expect(readLaunchdStopTimeout(LAUNCHD_ENV)).resolves.toEqual({
stop: { timeoutMs: 20_000, source: "launchd gui/501/ai.openclaw.gateway exit timeout" },
});
});
// A service account with no logged-in session has no gui domain: launchctl
// answers 125 "Domain does not support specified action" for gui/<uid> while
// user/<uid> prints normally. Verified on a headless macOS service account.
it("reaches the user domain when the account has no gui session", async () => {
execLaunchctl
.mockResolvedValueOnce({
code: 113,
stdout: "",
stderr: "Could not find service",
termination: "exit",
})
.mockResolvedValueOnce({
code: 125,
stdout: "",
stderr: "Domain does not support specified action",
termination: "exit",
})
.mockResolvedValueOnce(stopping("\texit timeout = 30\n\tpid = 4242\n"));
await expect(readLaunchdStopTimeout(LAUNCHD_ENV)).resolves.toEqual({
stop: { timeoutMs: 30_000, source: "launchd user/501/ai.openclaw.gateway exit timeout" },
});
expect(execLaunchctl).toHaveBeenCalledTimes(3);
});
it("refuses a same-named job in every other domain and says why", async () => {
execLaunchctl.mockResolvedValue(stopping("\texit timeout = 300\n\tpid = 99\n"));
const read = await readLaunchdStopTimeout(LAUNCHD_ENV);
// Nothing was established, so no deadline is reported at all and the caller
// keeps the platform-neutral policy it already resolved.
expect(read.stop).toBeNull();
for (const target of [
"system/ai.openclaw.gateway",
"gui/501/ai.openclaw.gateway",
"user/501/ai.openclaw.gateway",
]) {
expect(read.warning).toContain(`${target}: pid 99 is neither this process nor its launcher`);
}
});
// The installed service can keep a launcher parent while the serving Gateway
// runs as its child, so the job prints the launcher's pid. Requiring
// pid === process.pid there would reject the job that enforces the deadline.
it("accepts the job when it is this process's launcher parent", async () => {
execLaunchctl.mockResolvedValue(stopping("\texit timeout = 12\n\tpid = 4241\n"));
await expect(readLaunchdStopTimeout(LAUNCHD_ENV)).resolves.toEqual({
stop: { timeoutMs: 12_000, source: "launchd system/ai.openclaw.gateway exit timeout" },
});
});
// The launcher declares nothing, so its reap timer is derived from the expression
// the launcher itself arms from. A longer operator deadline cannot be spent under
// it: the parent force-kills this process first. Deriving rather than being told is
// what makes this hold when an already-running older launcher started this Gateway,
// which is the only shape an upgrade can take.
//
// Every marker is covered because all three respawn call sites reach the same
// launcher through the same function and arm the same timer. Gating on the
// Node-recovery marker alone left the two compile-cache respawns uncapped, and the
// packaged one can wrap a foreground `gateway run` on an installed service.
it.each(RESPAWN_MARKERS)(
"caps the job deadline at the launcher's derived reap timer for %s",
async (marker) => {
execLaunchctl.mockResolvedValue(stopping("\texit timeout = 55\n\tpid = 4241\n"));
await expect(readLaunchdStopTimeout({ ...SERVICE_ENV, [marker]: "1" })).resolves.toEqual({
stop: {
timeoutMs: 19_000,
source:
"launchd system/ai.openclaw.gateway exit timeout capped at the launcher's 19000ms stop timer",
},
});
},
);
// Holding the parent slot is not evidence of a reap timer. An operator wrapper can
// keep the job's pid and start the Gateway itself while running none, and capping
// its deadline would cut a drain nothing was going to interrupt.
it("leaves a parent that did not respawn this process capping nothing", async () => {
execLaunchctl.mockResolvedValue(stopping("\texit timeout = 55\n\tpid = 4241\n"));
await expect(readLaunchdStopTimeout(SERVICE_ENV)).resolves.toEqual({
stop: { timeoutMs: 55_000, source: "launchd system/ai.openclaw.gateway exit timeout" },
});
});
// Outside the service layout the launcher bounds its child by the platform-neutral
// stop policy, which outlasts any job deadline launchd will enforce, so nothing is
// cut even though the marker says a launcher is there.
it("derives the policy deadline when the respawning job is not the service", async () => {
execLaunchctl.mockResolvedValue(stopping("\texit timeout = 55\n\tpid = 4241\n"));
await expect(
readLaunchdStopTimeout({ ...LAUNCHD_ENV, OPENCLAW_NODE_UPDATE_RESPAWNED: "1" }),
).resolves.toEqual({
stop: { timeoutMs: 55_000, source: "launchd system/ai.openclaw.gateway exit timeout" },
});
});
// launchd reaps the whole job first, so a shorter job deadline still wins.
it("keeps a job deadline shorter than the launcher's reap timer", async () => {
execLaunchctl.mockResolvedValue(stopping("\texit timeout = 5\n\tpid = 4241\n"));
await expect(readLaunchdStopTimeout(RESPAWNED_SERVICE_ENV)).resolves.toEqual({
stop: { timeoutMs: 5_000, source: "launchd system/ai.openclaw.gateway exit timeout" },
});
});
// The launcher timer belongs to a parent, so a Gateway that is the job itself is
// never shortened by one even while carrying an inherited respawn marker.
it("ignores the launcher timer when this process is the job", async () => {
execLaunchctl.mockResolvedValue(stopping("\texit timeout = 55\n\tpid = 4242\n"));
await expect(readLaunchdStopTimeout(RESPAWNED_SERVICE_ENV)).resolves.toEqual({
stop: { timeoutMs: 55_000, source: "launchd system/ai.openclaw.gateway exit timeout" },
});
});
// launchd is stopping the job, so a deadline is running even though its value
// is unreadable. Guess short here: guessing long is what lets the drain die.
it.each(["\tpid = 4242\n", "\texit timeout = not-a-number\n\tpid = 4242\n"])(
"uses the conservative default when a running stop has no readable deadline",
async (fields) => {
execLaunchctl.mockResolvedValue(stopping(fields));
const read = await readLaunchdStopTimeout(LAUNCHD_ENV);
// A deadline IS running here, so this one is a real native budget even
// though its value had to be defaulted.
expect(read.stop?.timeoutMs).toBe(20_000);
expect(read.stop?.source).toBe(
"launchd system/ai.openclaw.gateway exit timeout unavailable; default ExitTimeOut",
);
expect(read.warning).toContain(
"launchd is stopping system/ai.openclaw.gateway but its exit timeout is missing or invalid",
);
},
);
// THE REGRESSION GUARD for the failed-inspection path. A probe that established
// nothing must report no deadline: handing back the Gateway's own stop policy
// here is what let the caller classify an unverified number as a native stop
// budget, capping a longer requested restart drain and arming a forced exit.
it("reports no deadline when the job cannot be inspected", async () => {
execLaunchctl.mockResolvedValue({
code: 1,
stdout: "",
stderr: "permission denied",
termination: "exit",
});
const read = await readLaunchdStopTimeout(LAUNCHD_ENV);
expect(read.stop).toBeNull();
expect(read.warning).toContain(
"system/ai.openclaw.gateway: launchctl print exited 1: permission denied",
);
expect(read.warning).toContain("keeping the Gateway stop policy");
expect(read.warning).toContain("Check the running job with launchctl print.");
});
it("survives launchctl throwing rather than exiting nonzero", async () => {
execLaunchctl.mockRejectedValue(new Error("spawn ENOENT"));
const read = await readLaunchdStopTimeout(LAUNCHD_ENV);
expect(read.stop).toBeNull();
expect(read.warning).toContain("launchctl print threw");
});
it("honours an explicit label override", async () => {
execLaunchctl.mockResolvedValue(stopping("\texit timeout = 45\n\tpid = 4242\n"));
await expect(
readLaunchdStopTimeout({ ...LAUNCHD_ENV, OPENCLAW_LAUNCHD_LABEL: "com.example.gw" }),
).resolves.toEqual({
stop: { timeoutMs: 45_000, source: "launchd system/com.example.gw exit timeout" },
});
});
it("warns instead of throwing when the configured label is invalid", async () => {
const read = await readLaunchdStopTimeout({
...LAUNCHD_ENV,
OPENCLAW_LAUNCHD_LABEL: "bad label/../etc",
});
expect(read.stop).toBeNull();
expect(read.warning).toContain("label could not be resolved");
expect(execLaunchctl).not.toHaveBeenCalled();
});
it("stays out of the way when this process is not a launchd job", async () => {
await expect(readLaunchdStopTimeout({})).resolves.toEqual({ stop: null });
expect(execLaunchctl).not.toHaveBeenCalled();
});
});
// Ties the names this file caps on back to the implementation's own list. Without
// this, dropping a marker from that list would leave the cap cases above still
// passing against the two that remained.
describe("launcher respawn markers", () => {
it.each(RESPAWN_MARKERS)("recognises %s as a respawning launcher", (marker) => {
expect(isRespawnedByLauncher({ [marker]: "1" })).toBe(true);
});
it("recognises nothing else", () => {
expect(isRespawnedByLauncher({})).toBe(false);
expect(isRespawnedByLauncher({ OPENCLAW_SUPERVISOR_MODE: "external" })).toBe(false);
// Only the set value counts, so an emptied or disabled marker caps nothing.
expect(isRespawnedByLauncher({ OPENCLAW_NODE_UPDATE_RESPAWNED: "0" })).toBe(false);
expect(isRespawnedByLauncher({ OPENCLAW_NODE_UPDATE_RESPAWNED: "" })).toBe(false);
});
});

View file

@ -0,0 +1,187 @@
import { parseStrictPositiveInteger } from "@openclaw/normalization-core/number-coercion";
import { truncateUtf16Safe } from "@openclaw/normalization-core/utf16-slice";
import { isForegroundGatewayRunArgv } from "../cli/gateway-run-argv.js";
import { execLaunchctl, formatLaunchctlResultDetail } from "../daemon/launchd-exec.js";
import { resolveLaunchAgentLabel } from "../daemon/launchd-label.js";
import { parseKeyValueOutput } from "../daemon/runtime-parse.js";
import { formatErrorMessage } from "./errors.js";
import {
isRespawnedByLauncher,
LAUNCH_AGENT_EXIT_TIMEOUT_SECONDS,
resolveLauncherStopTimeoutMs,
} from "./gateway-shutdown-budget.js";
import { detectRespawnSupervisor } from "./supervisor-markers.js";
type LaunchdStopTimeout = { timeoutMs: number; source: string };
// A warning without a deadline means inspection was inconclusive. Only a
// confirmed launchd stop or our launcher's independent reap timer yields a
// native budget; the caller owns the fallback policy.
export type LaunchdStopRead = { stop: LaunchdStopTimeout | null; warning?: string };
const LAUNCHCTL_PRINT_TIMEOUT_MS = 2_000;
// Check all domains: a service account may lack a GUI session, and a matching
// label in a different domain must not supply this process's deadline.
function resolveLaunchdDomains(label: string): string[] {
const uid = typeof process.getuid === "function" ? process.getuid() : 501;
// A LaunchDaemon and a LaunchAgent can carry the same label in different
// domains, so each is a candidate and the pid decides which one is ours.
// A service account with no logged-in session has no gui domain at all, and
// its per-user jobs live in `user/<uid>`, so both user domains are checked.
return [`system/${label}`, `gui/${uid}/${label}`, `user/${uid}/${label}`];
}
// launchctl can report the Gateway or a launcher parent that holds the job PID.
function resolveJobRelation(pid: number | undefined): "self" | "launcher" | null {
if (pid === undefined) {
return null;
}
if (pid === process.pid) {
return "self";
}
return pid === process.ppid ? "launcher" : null;
}
// Nested coalition blocks also contain state: use the single-tab job field,
// not the last state selected by the generic key-value parser.
function readJobState(printed: string): string | undefined {
return /^\tstate = (?<state>.+)$/mu.exec(printed)?.groups?.state?.trim();
}
// launchd reports SIGTERMed during bootout/kickstart, but keeps running on an
// externally delivered SIGTERM. Only the former starts ExitTimeOut's clock.
function isLaunchdStoppingJob(state: string | undefined): boolean {
return state !== undefined && /^SIG[A-Z0-9]+ed$/u.test(state);
}
// A respawn marker, not parenthood alone, proves the launcher runs a reap timer.
// Derive that timer from the same shared expression: an already-running parent
// cannot be taught a new announced deadline by upgrading the child.
function resolveParentLauncherStopTimeoutMs(env: NodeJS.ProcessEnv): number | undefined {
// One of these is set on every child the launcher respawns, and all of them predate
// this deadline being derived, so a marker is present for a Gateway that launcher
// started and absent for any other parent. An operator wrapper that keeps the job's
// pid and starts the Gateway itself runs no such timer, and capping its deadline
// would cut a drain nothing was going to interrupt.
if (!isRespawnedByLauncher(env)) {
return undefined;
}
return resolveLauncherStopTimeoutMs({
env,
platform: process.platform,
// The launcher branched on its own argv, and it respawns the child with the same
// user arguments, so testing ours reproduces the branch it took.
foreground: isForegroundGatewayRunArgv(process.argv),
});
}
// The job is stopping but its value is missing; use launchd's 20s default.
function defaultStopDeadline(target: string, reason: string): LaunchdStopRead {
const timeoutMs = LAUNCH_AGENT_EXIT_TIMEOUT_SECONDS * 1_000;
return {
stop: { timeoutMs, source: `launchd ${target} exit timeout unavailable; default ExitTimeOut` },
warning: `launchd is stopping ${target} but ${reason}; using ${timeoutMs}ms default. Check the running job with launchctl print.`,
};
}
// Unknown ownership of the stop cannot justify shortening a potentially
// unconstrained drain. Warn, but leave the caller's policy authoritative.
function unresolved(failures: string[]): LaunchdStopRead {
return {
stop: null,
// Each failure already names the target it came from, and the label-resolution
// case has no label to name, so the prefix deliberately carries neither.
warning: `Unable to inspect the launchd job; ${failures
.map((failure) => truncateUtf16Safe(failure.replaceAll(/\s+/g, " "), 500))
.join("; ")}; keeping the Gateway stop policy. Check the running job with launchctl print.`,
};
}
// Read the loaded job only on shutdown. Its ExitTimeOut applies only while
// launchd stops it; a respawn launcher can enforce an independent deadline.
export async function readLaunchdStopTimeout(
env: NodeJS.ProcessEnv = process.env,
): Promise<LaunchdStopRead> {
if (detectRespawnSupervisor(env, "darwin") !== "launchd") {
return { stop: null };
}
const failures: string[] = [];
let label: string;
try {
label = resolveLaunchAgentLabel(env);
} catch (error: unknown) {
return unresolved([`label could not be resolved: ${formatErrorMessage(error)}`]);
}
for (const target of resolveLaunchdDomains(label)) {
const failed = (reason: string) => failures.push(`${target}: ${reason}`);
const result = await execLaunchctl(["print", target], LAUNCHCTL_PRINT_TIMEOUT_MS).catch(
(error: unknown) => {
failed(`launchctl print threw: ${formatErrorMessage(error)}`);
return undefined;
},
);
if (!result) {
continue;
}
if (result.code !== 0) {
failed(`launchctl print exited ${result.code}: ${formatLaunchctlResultDetail(result)}`);
continue;
}
const printed = result.stdout || result.stderr || "";
const entries = parseKeyValueOutput(printed, "=");
// Adopting a deadline from a same-named job in the other domain would be
// worse than the fallback, so the printed job must be ours.
const pid = parseStrictPositiveInteger(entries.pid ?? "");
const relation = resolveJobRelation(pid);
if (!relation) {
failed(`pid ${pid ?? "missing"} is neither this process nor its launcher`);
continue;
}
// This is our job, so stop searching. Whether its deadline binds this stop is
// a separate question from whether the job was found.
const launcherMs =
relation === "launcher" ? resolveParentLauncherStopTimeoutMs(env) : undefined;
const state = readJobState(printed);
if (!isLaunchdStoppingJob(state)) {
// A direct signal to the Gateway starts neither launchd's clock nor its
// parent launcher's timer. Node cannot identify the sender, so capping on
// the parent marker would truncate that supported long drain. A signal
// forwarded by the launcher is indistinguishable here (existing limitation).
return !state && entries["exit timeout"] !== undefined
? {
stop: null,
warning: `launchd ${target} printed an exit timeout but no job state, so it is treated as not stopping and the Gateway stop policy is kept. Check the running job with launchctl print.`,
}
: { stop: null };
}
const rawSeconds = entries["exit timeout"]?.trim();
if (rawSeconds === "0") {
return launcherMs !== undefined
? {
stop: {
timeoutMs: launcherMs,
source: `launchd ${target} unlimited exit timeout capped at the launcher's ${launcherMs}ms stop timer`,
},
}
: { stop: { timeoutMs: Infinity, source: `launchd ${target} unlimited exit timeout` } };
}
const seconds = parseStrictPositiveInteger(rawSeconds ?? "");
if (seconds === undefined) {
return defaultStopDeadline(target, "its exit timeout is missing or invalid");
}
const jobMs = seconds * 1_000;
// A parent that reaps this process on its own timer binds before the job's
// ExitTimeOut, and spending the longer deadline would only get the drain
// force-killed.
return launcherMs !== undefined && launcherMs < jobMs
? {
stop: {
timeoutMs: launcherMs,
source: `launchd ${target} exit timeout capped at the launcher's ${launcherMs}ms stop timer`,
},
}
: { stop: { timeoutMs: jobMs, source: `launchd ${target} exit timeout` } };
}
return unresolved(failures);
}

View file

@ -18,6 +18,9 @@ beforeEach(() => {
vi.useFakeTimers();
child = new ChildProcess();
kill = vi.spyOn(child, "kill").mockReturnValue(true);
// `spawn` is hoisted once for the file, so its call log survives across cases
// and `toHaveBeenCalledExactlyOnceWith` would only ever hold for the first one.
spawn.mockClear();
spawn.mockReturnValue(child);
exit = vi.spyOn(process, "exit").mockImplementation(vi.fn<typeof process.exit>());
vi.spyOn(process, "kill").mockReturnValue(true);
@ -50,6 +53,21 @@ it.each([
XPC_SERVICE_NAME: "ai.openclaw.fixture",
});
detach = () => child.emit("exit", 0, null);
// The launcher tells the child nothing about the timer it armed: the serving
// Gateway derives the same deadline from the same shared expression, which is what
// lets a Gateway started by an already-running older launcher bound itself
// correctly. So the env must reach the child unchanged, and the escalation
// asserted below is what that derivation has to land on.
expect(spawn).toHaveBeenCalledExactlyOnceWith(
"node",
["child.mjs"],
expect.objectContaining({
env: {
OPENCLAW_LAUNCHD_LABEL: "ai.openclaw.fixture",
XPC_SERVICE_NAME: "ai.openclaw.fixture",
},
}),
);
const signal = process.listeners("SIGTERM").find((listener) => !previous.has(listener));
expect(signal).toBeDefined();
signal!("SIGTERM");

View file

@ -1291,6 +1291,7 @@ describe("frozen admission workflow barriers", () => {
for (const [index, child] of children.entries()) {
const toolingPaths = child.sources.tooling.map(({ path }) => path);
const selectedPaths = child.sources.selected.map(({ path }) => path);
expect(toolingPaths).toContain("scripts/lib/record-shared.mjs");
if (index === 1) {
expect(toolingPaths).toContain(extra);
expect(selectedPaths).toContain("extensions/codex/package.json");

View file

@ -92,7 +92,6 @@ function fixture(
recursive: true,
});
for (const file of [
"record-shared.mjs",
"update-compat-contract.mjs",
"openclaw-e2e-instance.sh",
"docker-e2e-watchdog.mjs",
@ -554,6 +553,7 @@ describe("frozen admission bootstrap repairs", () => {
it.each([
reader,
"scripts/lib/docker-e2e-scenarios.mts",
"scripts/lib/record-shared.mjs",
shell,
"scripts/lib/trusted-native-typescript.mjs",
"scripts/lib/native-typescript.mts",