diff --git a/docs/gateway/restart-recovery.md b/docs/gateway/restart-recovery.md index e519a1441965..5efbf132be71 100644 --- a/docs/gateway/restart-recovery.md +++ b/docs/gateway/restart-recovery.md @@ -159,9 +159,14 @@ Node version and the service manager tracks a launcher parent. The launcher forwards the stop signal and waits for the serving Gateway to drain within the shared service budget. Managed restart intent targets the live serving owner, so unfinished work still follows restart recovery when its drain budget expires. -The launchd stop budget remains 20 seconds; Linux units use the deadlines below. +When launchd drives the stop, macOS uses the running job's own `ExitTimeOut`, capped +by the launcher's own stop timer when a launcher is in the path, and Linux units use +their own stop timeout; both are described in the deadline sections below. This requires a Gateway started with the updated launcher: replacing files cannot -change a launcher that is already running. +change a launcher that is already running. The stop deadlines below are the +exception: the serving Gateway derives the launcher's reap timer rather than being +told it, precisely so a Gateway started by an already-running older launcher still +bounds itself correctly. For these managed restarts, if the CLI cannot verify the service command, serving owner, or restart-intent recording, it refuses the restart before signaling with @@ -214,11 +219,25 @@ unit's effective `TimeoutStopUSec`, including drop-ins. It logs the source and reconciled stop budget at both points, so a repaired unit takes effect without restarting first. Inspection and any wait for startup to finish consume the same shutdown deadline. Active-work drain uses at most -315 seconds, with 10 seconds reserved for final chat writes and server cleanup -and another 5 seconds before systemd's deadline. A unit with the default +315 seconds, with up to 10 seconds reserved for final chat writes and server cleanup +and up to another 5 seconds before systemd's deadline. A unit with the default 90-second stop timeout therefore gets a 75-second drain and an 85-second Gateway shutdown deadline. A shorter supervisor timeout also caps requested restart waits. -The drained work, ordering, and interruption behavior stay the same. + +Both allowances are bounded when the deadline cannot fund them, by the same rule the +launchd budget uses and for the same reason: subtracting two values sized for a +315-second drain from a short stop timeout consumed it entirely and left active work +nothing. A unit whose `TimeoutStopSec` is 20 seconds or longer keeps the allocation it +already had, less whatever the stop inspection itself cost, rather than to the exact +millisecond: 20 seconds is only the least that funds the full 10-second reserve +alongside a 5-second drain before that cost comes off. Below that the deadline cannot fund +both, and the reserve yields to keep a drain: a unit at or below 15 seconds now drains +at all where it previously drained for zero milliseconds, and the four values between +trade part of a reserve for a drain that was under 5 seconds. The reserve never falls +below half the shutdown budget. The exit margin holds its full 5 seconds down to a +20-second deadline and shrinks below that, to 3.75 seconds at 15 seconds and a quarter +of anything shorter. The drained work, ordering, and interruption behavior stay the +same, and systemd's own 90-second default is unaffected. Service-child cleanup uses the remaining Gateway shutdown budget, leaving time for final exit bookkeeping. A forced restart drains admitted work within the same @@ -309,6 +328,59 @@ before acknowledging, without overwriting a newer turn in that session. If another child cannot be stopped, the response still reports incomplete cancellation; the captured parent's cancellation is persisted before that error. +### Launchd stop deadlines + +On a launchd-driven stop, the macOS Gateway reads the **loaded** job's effective +`exit timeout` with `launchctl print`. It checks the system, GUI, and user +domains, and accepts only a job whose PID is the Gateway or its launcher. Restart +ownership can still be external: the supervisor that enforces the stop, not the +owner of the next start, determines the deadline. A direct SIGTERM does not start +launchd's stop clock, so it keeps the existing Gateway stop policy unless an +OpenClaw launcher independently enforces its own child reap timer. + +The **stop budget** is time available until the supervisor can kill the process. +The Gateway sets aside an exit margin (up to 5 seconds) and plans its shutdown +inside the remainder. **Drain** is the first part of that shutdown: stop admitting +new work and wait for active turns and background tasks to settle. The rest is +reserved for final writes and service cleanup (up to 10 seconds). The Gateway +logs the chosen source, drain, shutdown deadline, reserve, and exit margin. + +For the installed 20-second LaunchAgent job, the nominal split is 5 seconds of +drain, 10 seconds for cleanup, and 5 seconds before launchd's deadline. Time +spent inspecting the job is debited. For a shorter custom deadline the exit +margin is at most a quarter, and the cleanup reserve gives way to leave active +work up to 5 seconds of drain (at most half of the remaining shutdown budget +when it is very short). For example, a 5-second job nominally leaves 1.875 +seconds each for drain and cleanup after its 1.25-second exit margin; a +15-second job leaves 5 seconds of drain, 6.25 seconds for cleanup, and a +3.75-second margin. **This reduces cleanup time for custom jobs below 20 +seconds.** The default systemd 90-second deadline is unaffected. The probe can +consume up to three 2-second calls; a very slow inspection can leave no drain. + +If a Node-recovery or compile-cache launcher is the job's PID, its own child +reap timer may be shorter than launchd's deadline. The Gateway caps its budget +at that timer only when its respawn markers establish that this launcher started +it; an unrelated parent does not shorten the budget. + +A direct signal to a marked launcher can start its own timer while launchd +still reports `running`. The Gateway cannot distinguish that forwarded signal +from a direct signal to its child, where the launcher has no timer, so it does +not cap this case. Operators stopping a launcher directly should use the loaded +job stop instead to get an enforceable, reported deadline. + +`ExitTimeOut=0` means no launchd stop deadline, not a missing value. Without a +launcher timer, the Gateway keeps its ordinary policy and does not classify the +job as a native deadline. An unreadable job leaves +the existing stop policy in place with a warning, while a confirmed stopping +job whose timeout is missing uses launchd's 20-second default with a warning. + +The value comes from the job launchd has **loaded**, not the plist on disk. +Editing `ExitTimeOut` or running `launchctl kickstart -k` does not reload it; +`launchctl bootout` followed by `launchctl bootstrap` does, restarting the +Gateway. Reading per stop avoids a stale startup snapshot, not a stale loaded +job. On macOS 27, launchd reports at most 60 seconds even if the plist asks +for more; inspect the loaded job to confirm the effective value. + ## Host sleep and process freezes When a gateway host wakes from sleep, a virtual machine resumes, or the process diff --git a/gateway-shutdown-budget.d.mts b/gateway-shutdown-budget.d.mts index 013678346a31..b69d5aa1b637 100644 --- a/gateway-shutdown-budget.d.mts +++ b/gateway-shutdown-budget.d.mts @@ -4,3 +4,13 @@ export const GATEWAY_SHUTDOWN_TIMEOUT_MS: number; export const GATEWAY_SERVICE_STOP_TIMEOUT_MS: number; export const LAUNCH_AGENT_EXIT_TIMEOUT_SECONDS: 20; export const GATEWAY_RESTART_REPLACEMENT_TIMEOUT_MS: number; +export const RESPAWN_SIGNAL_FORCE_KILL_GRACE_MS: number; +export const RESPAWN_SIGNAL_HARD_EXIT_GRACE_MS: number; +export function resolveSupervisorExitMarginMs(stopTimeoutMs: number): number; +export function resolveShutdownReserveMs(shutdownTimeoutMs: number): number; +export function isRespawnedByLauncher(env: NodeJS.ProcessEnv): boolean; +export function resolveLauncherStopTimeoutMs(params: { + env: NodeJS.ProcessEnv; + platform: NodeJS.Platform; + foreground: boolean; +}): number; diff --git a/gateway-shutdown-budget.mjs b/gateway-shutdown-budget.mjs index c6469692105c..2609fc89630d 100644 --- a/gateway-shutdown-budget.mjs +++ b/gateway-shutdown-budget.mjs @@ -9,3 +9,72 @@ export const GATEWAY_SERVICE_STOP_TIMEOUT_MS = GATEWAY_SHUTDOWN_TIMEOUT_MS + GATEWAY_SUPERVISOR_EXIT_MARGIN_MS; export const LAUNCH_AGENT_EXIT_TIMEOUT_SECONDS = 20; + +// Keep a positive shutdown budget when a supervisor's deadline is under 20s. +const GATEWAY_SUPERVISOR_EXIT_MARGIN_SHARE = 0.25; + +// Preserve the template's 5s drain where possible. Under 20s, trade some cleanup +// reserve for drain; on very short jobs, split the remaining budget in half. +// Inspection time is debited before this allocation, so a slow probe can also +// reduce the drain. The measured upgrade tradeoff is documented in restart-recovery. +const GATEWAY_SHUTDOWN_DRAIN_FLOOR_MS = 5_000; +const GATEWAY_SHUTDOWN_DRAIN_FLOOR_SHARE = 0.5; + +/** The exit margin to hold back from a supervisor-enforced stop deadline. */ +export const resolveSupervisorExitMarginMs = (stopTimeoutMs) => + Math.min( + GATEWAY_SUPERVISOR_EXIT_MARGIN_MS, + Math.floor(Math.max(0, stopTimeoutMs) * GATEWAY_SUPERVISOR_EXIT_MARGIN_SHARE), + ); + +/** The post-drain reserve to hold back from a resolved shutdown budget. */ +export const resolveShutdownReserveMs = (shutdownTimeoutMs) => { + const budgetMs = Math.max(0, shutdownTimeoutMs); + const drainFloorMs = Math.min( + GATEWAY_SHUTDOWN_DRAIN_FLOOR_MS, + Math.floor(budgetMs * GATEWAY_SHUTDOWN_DRAIN_FLOOR_SHARE), + ); + return Math.min(GATEWAY_SHUTDOWN_RESERVE_MS, budgetMs - drainFloorMs); +}; + +// Escalation graces the Node recovery launcher applies to a stopping child. Kept +// here rather than in the launcher so the serving Gateway can derive the deadline +// its parent enforces from the same numbers the parent armed it from. +const RESPAWN_SIGNAL_EXIT_GRACE_MS = 1_000; +export const RESPAWN_SIGNAL_FORCE_KILL_GRACE_MS = 1_000; +export const RESPAWN_SIGNAL_HARD_EXIT_GRACE_MS = 1_000; + +// A launcher inside OpenClaw's LaunchAgent gets its template deadline; otherwise +// it uses the generic service stop policy. Its child inherits these markers. +const resolveRespawnServiceStopTimeoutMs = (env, platform) => { + const launchdService = env.OPENCLAW_LAUNCHD_LABEL?.trim(); + return platform === "darwin" && launchdService && env.XPC_SERVICE_NAME === launchdService + ? LAUNCH_AGENT_EXIT_TIMEOUT_SECONDS * 1_000 + : GATEWAY_SERVICE_STOP_TIMEOUT_MS; +}; + +// Every runRespawnedChild call site stamps one of these. The compile-cache +// marker is also used by a different respawner, but that one refuses foreground +// Gateway runs on darwin. Keep this list with the launcher deadline it authorizes. +const RESPAWN_LAUNCHER_MARKER_ENV_VARS = [ + "OPENCLAW_NODE_UPDATE_RESPAWNED", + "OPENCLAW_COMPILE_CACHE_DISABLED_RESPAWNED", + "OPENCLAW_PACKAGED_COMPILE_CACHE_RESPAWNED", +]; + +/** Whether a `runRespawnedChild` parent started this process. */ +export const isRespawnedByLauncher = (env) => + RESPAWN_LAUNCHER_MARKER_ENV_VARS.some((name) => env[name] === "1"); + +// Shared with the serving Gateway: derive the parent's reap deadline from the +// same graces, including on the first update while the old launcher is still live. +export const resolveLauncherStopTimeoutMs = ({ env, platform, foreground }) => { + const serviceStopTimeoutMs = resolveRespawnServiceStopTimeoutMs(env, platform); + const signalExitGraceMs = + platform !== "win32" && foreground + ? serviceStopTimeoutMs - + RESPAWN_SIGNAL_FORCE_KILL_GRACE_MS - + RESPAWN_SIGNAL_HARD_EXIT_GRACE_MS + : RESPAWN_SIGNAL_EXIT_GRACE_MS; + return signalExitGraceMs + RESPAWN_SIGNAL_FORCE_KILL_GRACE_MS; +}; diff --git a/node-runtime-recovery.mjs b/node-runtime-recovery.mjs index 91c7f8c895c5..b784f43238aa 100644 --- a/node-runtime-recovery.mjs +++ b/node-runtime-recovery.mjs @@ -6,8 +6,9 @@ import path from "node:path"; import { consumeRootOptionToken as consumeLauncherRootOptionToken } from "./cli-root-options.mjs"; import { isForegroundGatewayRunArgv } from "./gateway-run-argv.mjs"; import { - GATEWAY_SERVICE_STOP_TIMEOUT_MS, - LAUNCH_AGENT_EXIT_TIMEOUT_SECONDS, + RESPAWN_SIGNAL_FORCE_KILL_GRACE_MS, + RESPAWN_SIGNAL_HARD_EXIT_GRACE_MS, + resolveLauncherStopTimeoutMs, } from "./gateway-shutdown-budget.mjs"; import { detectCurrentSqliteCapabilities, @@ -39,22 +40,21 @@ const respawnSignals = process.platform === "win32" ? ["SIGTERM", "SIGINT", "SIGBREAK"] : ["SIGTERM", "SIGINT", "SIGHUP", "SIGQUIT"]; -const respawnSignalExitGraceMs = 1_000; -const respawnSignalForceKillGraceMs = 1_000; -const respawnSignalHardExitGraceMs = 1_000; +const respawnSignalForceKillGraceMs = RESPAWN_SIGNAL_FORCE_KILL_GRACE_MS; +const respawnSignalHardExitGraceMs = RESPAWN_SIGNAL_HARD_EXIT_GRACE_MS; export const runRespawnedChild = (command, args, env) => { - const launchdService = env.OPENCLAW_LAUNCHD_LABEL?.trim(); - const serviceStopTimeoutMs = - process.platform === "darwin" && launchdService && env.XPC_SERVICE_NAME === launchdService - ? LAUNCH_AGENT_EXIT_TIMEOUT_SECONDS * 1_000 - : GATEWAY_SERVICE_STOP_TIMEOUT_MS; // The serving Gateway owns drain and cleanup. Reap a stuck child only in the // supervisor's exit margin, after that owner has had its full shutdown budget. - const signalExitGraceMs = - process.platform !== "win32" && isForegroundGatewayRunArgv(process.argv) - ? serviceStopTimeoutMs - respawnSignalForceKillGraceMs - respawnSignalHardExitGraceMs - : respawnSignalExitGraceMs; + // The shared resolver owns this arithmetic so the serving Gateway derives the very + // same deadline from the same expression, which is what lets a Gateway started by + // any build of this launcher bound itself correctly without being told. + const launcherStopTimeoutMs = resolveLauncherStopTimeoutMs({ + env, + platform: process.platform, + foreground: isForegroundGatewayRunArgv(process.argv), + }); + const signalExitGraceMs = launcherStopTimeoutMs - respawnSignalForceKillGraceMs; const stdioIsTerminal = process.stdin.isTTY || process.stdout.isTTY; const child = spawn(command, args, { stdio: "inherit", @@ -62,7 +62,10 @@ export const runRespawnedChild = (command, args, env) => { windowsHide: !stdioIsTerminal, }); const listeners = new Map(); - // Keep signal forwarding and bounded shutdown in sync with src/entry.compile-cache.ts. + // Keep signal forwarding and bounded shutdown in sync with src/entry.compile-cache.ts, + // which drives src/process/respawn-child-runner.ts. That runner still holds its own + // copies of the escalation graces and reaps on a fixed short one, so only this + // launcher's deadline is the one the serving Gateway derives. let signalExitTimer = null; let signalForceKillTimer = null; let signalHardExitTimer = null; diff --git a/scripts/preflight-frozen-target-contracts.mjs b/scripts/preflight-frozen-target-contracts.mjs index a508fba840b1..60236ac75ee7 100644 --- a/scripts/preflight-frozen-target-contracts.mjs +++ b/scripts/preflight-frozen-target-contracts.mjs @@ -1180,11 +1180,6 @@ async function preflightFrozenTargetContracts(input, workflow = false, verifiedT for (const path of supportFiles[consumer] ?? []) { required(sources.tooling, `scripts/e2e/lib/${path}`); } - if ( - ["npm-onboard-channel-agent", "codex-on-demand", "update-corrupt-plugin"].includes(consumer) - ) { - required(sources.tooling, "scripts/lib/record-shared.mjs"); - } if (consumer === "update-corrupt-plugin") { required(sources.tooling, "scripts/lib/update-compat-contract.mjs"); required(sources.tooling, "scripts/lib/openclaw-e2e-instance.sh"); diff --git a/src/cli/gateway-cli/run-loop-shutdown-budget.test.ts b/src/cli/gateway-cli/run-loop-shutdown-budget.test.ts index 2b46a2247f34..417afd2923c8 100644 --- a/src/cli/gateway-cli/run-loop-shutdown-budget.test.ts +++ b/src/cli/gateway-cli/run-loop-shutdown-budget.test.ts @@ -1,16 +1,29 @@ +import { performance } from "node:perf_hooks"; import { afterEach, beforeEach, describe, expect, it, vi } from "vitest"; -import { resolveGatewayShutdownBudget } from "./run-loop-shutdown-budget.js"; +import { + GATEWAY_SHUTDOWN_RESERVE_MS, + GATEWAY_SUPERVISOR_EXIT_MARGIN_MS, +} from "../../infra/gateway-shutdown-budget.js"; +import { + resolveGatewayShutdownBudget, + resolveGatewayShutdownDrainBudget, +} from "./run-loop-shutdown-budget.js"; -const { readFile, execUser, execSystem } = vi.hoisted(() => ({ +const { readFile, execUser, execSystem, execLaunchctl } = vi.hoisted(() => ({ readFile: vi.fn(), execUser: vi.fn(), execSystem: vi.fn(), + execLaunchctl: vi.fn(), })); vi.mock("node:fs/promises", () => ({ default: { readFile } })); vi.mock("../../daemon/systemd-exec.js", () => ({ execSystemctlUser: execUser, execSystemctl: execSystem, })); +vi.mock("../../daemon/launchd-exec.js", async (importOriginal) => ({ + ...(await importOriginal()), + execLaunchctl, +})); beforeEach(() => { vi.stubGlobal("process", { ...process, platform: "linux", getuid: () => 1000, env: {} }); @@ -23,6 +36,7 @@ beforeEach(() => { "User=openclaw\nType=simple\nNotifyAccess=none\nKillMode=control-group", stderr: "", }); + execLaunchctl.mockReset(); }); afterEach(() => vi.unstubAllGlobals()); @@ -93,4 +107,472 @@ describe("Gateway stop deadline independent of restart ownership", () => { expect(execSystem).not.toHaveBeenCalled(); expect(execUser).not.toHaveBeenCalled(); }); + + // The systemd coverage above only pins TimeoutStopUSec at 90 seconds and 10 minutes, + // both comfortably above the 20 second threshold. This mirrors the darwin exit-timeout + // table's sub-20-second cases against a systemd stub: a unit deadline in this band + // gives up part of a reserve a flat subtraction would have funded in full, rather than + // draining for zero milliseconds as that flat subtraction did. + it.each([ + { seconds: 19, timeoutMs: 14_250, reserveMs: 9_250, drainMs: 5_000, fixedDrainMs: 4_000 }, + { seconds: 16, timeoutMs: 12_000, reserveMs: 7_000, drainMs: 5_000, fixedDrainMs: 1_000 }, + { seconds: 15, timeoutMs: 11_250, reserveMs: 6_250, drainMs: 5_000, fixedDrainMs: 0 }, + { seconds: 10, timeoutMs: 7_500, reserveMs: 3_750, drainMs: 3_750, fixedDrainMs: 0 }, + ])( + "keeps a drain a $seconds second systemd TimeoutStopUSec previously spent on overhead", + async ({ seconds, timeoutMs, reserveMs, drainMs, fixedDrainMs }) => { + process.env.OPENCLAW_SUPERVISOR_MODE = "external"; + execSystem.mockResolvedValue({ + code: 0, + stdout: `LoadState=loaded\nTimeoutStopUSec=${seconds}s\nInvocationID=own`, + stderr: "", + }); + const budget = await resolveGatewayShutdownBudget("external", { + info: vi.fn(), + warn: vi.fn(), + }); + expect(budget.timeoutMs).toBe(timeoutMs); + expect(budget.reserveMs).toBe(reserveMs); + expect(budget.timeoutMs - budget.reserveMs).toBe(drainMs); + // What a flat, unshared subtraction would have left active work at the same deadline. + expect( + Math.max( + 0, + seconds * 1_000 - GATEWAY_SUPERVISOR_EXIT_MARGIN_MS - GATEWAY_SHUTDOWN_RESERVE_MS, + ), + ).toBe(fixedDrainMs); + expect(budget.reserveMs).toBeGreaterThanOrEqual(Math.floor(timeoutMs / 2)); + }, + ); +}); + +describe("Gateway stop deadline follows the launchd stop that is actually running", () => { + // A stop that is under way. `previous` is the startup budget a darwin Gateway + // resolves before any stop exists, which is the platform-neutral policy. + // The budget subtracts `performance.now() - acceptedAtMs` and floors it at + // zero, so an acceptance stamped ahead of the clock records exactly no elapsed + // time and keeps the asserted numbers exact instead of off by a stray + // millisecond. + const stoppingNow = { + previous: { timeoutMs: 325_000, nativeStopBudget: false }, + acceptedAtMs: Number.MAX_SAFE_INTEGER, + }; + const printed = (state: string, fields: string) => ({ + code: 0, + stdout: `system/ai.openclaw.gateway = {\n\tstate = ${state}\n\n${fields}\tresource coalition = {\n\t\tstate = active\n\t}\n}\n`, + stderr: "", + termination: "exit", + }); + + beforeEach(() => { + vi.stubGlobal("process", { + ...process, + platform: "darwin", + pid: 4242, + getuid: () => 501, + env: {}, + }); + process.env.XPC_SERVICE_NAME = "ai.openclaw.gateway"; + process.env.OPENCLAW_SUPERVISOR_MODE = "external"; + }); + + // THE REGRESSION GUARD. The linked report is an externally delivered SIGTERM + // under a five second job, where launchd never starts its clock and the drain ran + // its full 315 seconds. Measured on the reporting host, it then hit its own + // timeout with work still active rather than finishing early, so the drain was + // being used. Adopting the job deadline there would hand that same supported + // setup a zero drain and cut that work off at once. + it("keeps the full drain when launchd did not initiate the stop", async () => { + execLaunchctl.mockResolvedValue(printed("running", "\texit timeout = 5\n\tpid = 4242\n")); + const info = vi.fn(); + const budget = await resolveGatewayShutdownBudget( + "external", + { info, warn: vi.fn() }, + stoppingNow, + ); + budget.log("shutdown"); + expect(info).toHaveBeenCalledWith( + "shutdown budget at shutdown: drain=315000ms shutdown=325000ms reserve=10000ms exitMargin=5000ms; source=Gateway stop policy=330000ms", + ); + expect(budget.timeoutMs).toBe(325_000); + expect(budget.nativeStopBudget).toBe(false); + }); + + it.each(["external", "launchd"])( + "does not treat a launchd ExitTimeOut of zero as a native deadline under %s ownership", + async (supervisor) => { + execLaunchctl.mockResolvedValue(printed("SIGTERMed", "\texit timeout = 0\n\tpid = 4242\n")); + const budget = await resolveGatewayShutdownBudget( + supervisor, + { info: vi.fn(), warn: vi.fn() }, + stoppingNow, + ); + expect(budget.timeoutMs).toBe(325_000); + expect(budget.nativeStopBudget).toBe(false); + const drain = resolveGatewayShutdownDrainBudget({ + budget, + action: "restart", + forceRestart: false, + restartWithoutSupervisor: false, + acceptedAtMs: performance.now(), + requestedRestartDrainTimeoutMs: 600_000, + }); + expect(drain.drainTimeoutMs).toBeGreaterThan(590_000); + }, + ); + + // A deadline this short cannot fund the fixed 5s margin and 10s reserve, and + // subtracting them outright left the job's whole 5 seconds spent on overhead with + // nothing to drain. Each allowance is capped at a share of what it is carved from, + // so the short job keeps a proportional drain that still fits inside the deadline. + it("keeps a proportional drain when the job's exit timeout cannot fund the fixed allowances", async () => { + execLaunchctl.mockResolvedValue( + printed("SIGTERMed", "\tminimum runtime = 10\n\texit timeout = 5\n\tpid = 4242\n"), + ); + const info = vi.fn(); + const budget = await resolveGatewayShutdownBudget( + "external", + { info, warn: vi.fn() }, + stoppingNow, + ); + budget.log("shutdown"); + expect(info).toHaveBeenCalledWith( + "shutdown budget at shutdown: drain=1875ms shutdown=3750ms reserve=1875ms exitMargin=1250ms; source=launchd system/ai.openclaw.gateway exit timeout=5000ms", + ); + expect(budget.timeoutMs).toBe(3_750); + expect(budget.reserveMs).toBe(1_875); + expect(budget.nativeStopBudget).toBe(true); + }); + + // Drain must never be starved to zero by the allowances: every positive deadline + // leaves active work some time, and a longer deadline never yields less of it. + it.each([1, 2, 5, 10, 15, 20, 25, 47, 60])( + "leaves a positive drain for a %s second exit timeout", + async (seconds) => { + execLaunchctl.mockResolvedValue( + printed("SIGTERMed", `\texit timeout = ${seconds}\n\tpid = 4242\n`), + ); + const budget = await resolveGatewayShutdownBudget( + "external", + { info: vi.fn(), warn: vi.fn() }, + stoppingNow, + ); + expect(budget.timeoutMs - budget.reserveMs).toBeGreaterThan(0); + expect(budget.timeoutMs).toBeLessThanOrEqual(seconds * 1_000); + }, + ); + + it.each([ + { seconds: 20, timeoutMs: 15_000, reserveMs: 10_000, drainMs: 5_000, exitMarginMs: 5_000 }, + { seconds: 47, timeoutMs: 42_000, reserveMs: 10_000, drainMs: 32_000, exitMarginMs: 5_000 }, + { seconds: 55, timeoutMs: 50_000, reserveMs: 10_000, drainMs: 40_000, exitMarginMs: 5_000 }, + ])( + "derives the budget from a $seconds second exit timeout", + async ({ seconds, timeoutMs, reserveMs, drainMs, exitMarginMs }) => { + execLaunchctl.mockResolvedValue( + printed("SIGTERMed", `\texit timeout = ${seconds}\n\tpid = 4242\n`), + ); + const info = vi.fn(); + const budget = await resolveGatewayShutdownBudget( + "external", + { info, warn: vi.fn() }, + stoppingNow, + ); + budget.log("shutdown"); + expect(budget.timeoutMs).toBe(timeoutMs); + expect(budget.reserveMs).toBe(reserveMs); + expect(info).toHaveBeenCalledWith( + `shutdown budget at shutdown: drain=${drainMs}ms shutdown=${timeoutMs}ms reserve=${reserveMs}ms exitMargin=${exitMarginMs}ms; source=launchd system/ai.openclaw.gateway exit timeout=${seconds * 1_000}ms`, + ); + }, + ); + + // Every case above pins `acceptedAtMs` to `Number.MAX_SAFE_INTEGER`, which floors + // elapsed at zero and proves nothing about a real, nonzero debit. `performance.now` + // is not faked anywhere in this file, so this pins it directly rather than trusting + // the wall clock to land on a specific millisecond by chance: the reserve gives up + // exactly that observed 13ms debit and the 5 second drain floor still funds in full. + it("debits the reserve by a real elapsed delay, leaving the drain floor untouched", async () => { + execLaunchctl.mockResolvedValue(printed("SIGTERMed", "\texit timeout = 20\n\tpid = 4242\n")); + const nowMs = performance.now(); + const clock = vi.spyOn(performance, "now").mockReturnValue(nowMs); + try { + const budget = await resolveGatewayShutdownBudget( + "external", + { info: vi.fn(), warn: vi.fn() }, + { previous: stoppingNow.previous, acceptedAtMs: nowMs - 13 }, + ); + expect(budget.timeoutMs).toBe(14_987); + expect(budget.reserveMs).toBe(9_987); + expect(budget.timeoutMs - budget.reserveMs).toBe(5_000); + } finally { + clock.mockRestore(); + } + }); + + // The case above only covers a debit small enough that the reserve pays it alone. The + // probe can cost far more than 13ms: it is up to three `launchctl print` calls at a + // 2 second timeout each. Past a 5 second debit the drain floor is share-bounded too, + // so the claim that a 20 second deadline keeps its old allocation less the debit stops + // holding and both allowances converge on half the remainder. Pinning that boundary + // keeps the documented threshold honest instead of reasoned. + it("converges the reserve and the drain once the elapsed debit passes the floor", async () => { + execLaunchctl.mockResolvedValue(printed("SIGTERMed", "\texit timeout = 20\n\tpid = 4242\n")); + const nowMs = performance.now(); + const clock = vi.spyOn(performance, "now").mockReturnValue(nowMs); + try { + const budget = await resolveGatewayShutdownBudget( + "external", + { info: vi.fn(), warn: vi.fn() }, + { previous: stoppingNow.previous, acceptedAtMs: nowMs - 6_000 }, + ); + expect(budget.timeoutMs).toBe(9_000); + expect(budget.reserveMs).toBe(4_500); + expect(budget.timeoutMs - budget.reserveMs).toBe(4_500); + // Not the old allocation less the debit: that would have left the reserve at 4000. + expect(budget.reserveMs).not.toBe(GATEWAY_SHUTDOWN_RESERVE_MS - 6_000); + } finally { + clock.mockRestore(); + } + }); + + // The allowances are only ever capped to keep a drain, never to reallocate a deadline + // that already worked. A deadline able to fund the 10s reserve alongside the 5s drain + // the 20s template yields needs 15s of shutdown budget, which every deadline from 20s + // up has, so all of them must resolve exactly what subtracting the fixed allowances + // outright resolved. Asserting against that arithmetic rather than against literals is + // what makes this a regression test: capping the reserve at a share of the budget, as + // an earlier revision did, drops a 20s job's reserve to 7500ms and fails here. + it.each([20, 21, 25, 30, 47, 55, 60])( + "allocates a %s second exit timeout exactly as the fixed allowances did", + async (seconds) => { + execLaunchctl.mockResolvedValue( + printed("SIGTERMed", `\texit timeout = ${seconds}\n\tpid = 4242\n`), + ); + const budget = await resolveGatewayShutdownBudget( + "external", + { info: vi.fn(), warn: vi.fn() }, + stoppingNow, + ); + const fixedTimeoutMs = seconds * 1_000 - GATEWAY_SUPERVISOR_EXIT_MARGIN_MS; + expect(budget.timeoutMs).toBe(fixedTimeoutMs); + expect(budget.reserveMs).toBe(GATEWAY_SHUTDOWN_RESERVE_MS); + expect(budget.timeoutMs - budget.reserveMs).toBe( + fixedTimeoutMs - GATEWAY_SHUTDOWN_RESERVE_MS, + ); + }, + ); + + // Under 20 seconds the budget cannot fund both allowances, so one has to give. The + // fixed subtraction gave up the drain: 15 seconds and below drained for 0ms, and the + // four deadlines between left active work under 5 seconds. These pin what is given up + // instead, and that the reserve never falls below half the budget doing it. No shipped + // template or platform default lands here: the LaunchAgent template and launchd's own + // default are both 20 seconds, and systemd's default stop timeout is 90. + it.each([ + { seconds: 19, timeoutMs: 14_250, reserveMs: 9_250, drainMs: 5_000, fixedDrainMs: 4_000 }, + { seconds: 16, timeoutMs: 12_000, reserveMs: 7_000, drainMs: 5_000, fixedDrainMs: 1_000 }, + { seconds: 15, timeoutMs: 11_250, reserveMs: 6_250, drainMs: 5_000, fixedDrainMs: 0 }, + { seconds: 10, timeoutMs: 7_500, reserveMs: 3_750, drainMs: 3_750, fixedDrainMs: 0 }, + ])( + "keeps a drain a $seconds second exit timeout previously spent on overhead", + async ({ seconds, timeoutMs, reserveMs, drainMs, fixedDrainMs }) => { + execLaunchctl.mockResolvedValue( + printed("SIGTERMed", `\texit timeout = ${seconds}\n\tpid = 4242\n`), + ); + const budget = await resolveGatewayShutdownBudget( + "external", + { info: vi.fn(), warn: vi.fn() }, + stoppingNow, + ); + expect(budget.timeoutMs).toBe(timeoutMs); + expect(budget.reserveMs).toBe(reserveMs); + expect(budget.timeoutMs - budget.reserveMs).toBe(drainMs); + // What the fixed subtraction left active work at the same deadline. + expect( + Math.max( + 0, + seconds * 1_000 - GATEWAY_SUPERVISOR_EXIT_MARGIN_MS - GATEWAY_SHUTDOWN_RESERVE_MS, + ), + ).toBe(fixedDrainMs); + expect(budget.reserveMs).toBeGreaterThanOrEqual(Math.floor(timeoutMs / 2)); + }, + ); + + // Failing to inspect the job establishes nothing, so shortening the drain here + // would cut work that no launchd deadline was bounding. + it("warns and keeps the platform-neutral policy when the job cannot be inspected", async () => { + execLaunchctl.mockResolvedValue({ + code: 1, + stdout: "", + stderr: "permission denied", + termination: "exit", + }); + const warn = vi.fn(); + const budget = await resolveGatewayShutdownBudget( + "external", + { info: vi.fn(), warn }, + stoppingNow, + ); + expect(budget.timeoutMs).toBe(325_000); + // The number alone is not the contract. A failed probe confirmed no launchd + // deadline, so this must not be classified as a native stop budget either: + // that flag is what caps a restart drain and arms a forced exit. + expect(budget.nativeStopBudget).toBe(false); + expect(warn).toHaveBeenCalledExactlyOnceWith( + expect.stringContaining("Unable to inspect the launchd job"), + ); + }); + + // The flag is only worth asserting because of what it does downstream, so drive + // the real consumer. An operator restart that asked to drain for ten minutes + // keeps that request when no launchd deadline was ever confirmed. + it("leaves a longer requested restart drain uncapped when the job cannot be inspected", async () => { + execLaunchctl.mockResolvedValue({ + code: 1, + stdout: "", + stderr: "permission denied", + termination: "exit", + }); + const budget = await resolveGatewayShutdownBudget( + "external", + { info: vi.fn(), warn: vi.fn() }, + stoppingNow, + ); + expect(budget.nativeStopBudget).toBe(false); + const drain = resolveGatewayShutdownDrainBudget({ + budget, + action: "restart", + forceRestart: false, + restartWithoutSupervisor: false, + acceptedAtMs: performance.now(), + requestedRestartDrainTimeoutMs: 600_000, + }); + // Only elapsed time comes off the request; no supervisor ceiling applies. + expect(drain.drainTimeoutMs).toBeGreaterThan(590_000); + expect(drain.restartTimeoutMs()).toBe(325_000); + }); + + // The same consumer, with a deadline that WAS confirmed, still gets capped. + // Without this pair the test above would also pass if the launchd read were + // deleted outright. + it("caps that same restart drain at a confirmed job deadline", async () => { + execLaunchctl.mockResolvedValue(printed("SIGTERMed", "\texit timeout = 47\n\tpid = 4242\n")); + const budget = await resolveGatewayShutdownBudget( + "external", + { info: vi.fn(), warn: vi.fn() }, + stoppingNow, + ); + expect(budget.nativeStopBudget).toBe(true); + const drain = resolveGatewayShutdownDrainBudget({ + budget, + action: "restart", + forceRestart: false, + restartWithoutSupervisor: false, + acceptedAtMs: performance.now(), + requestedRestartDrainTimeoutMs: 600_000, + }); + // 47s job - 5s exit margin = 42000ms shutdown, less the 10000ms reserve. + expect(drain.drainTimeoutMs).toBe(32_000); + }); + + // A launchd-OWNED Gateway, rather than an externally supervised one. Its startup + // budget is already native, so this is the configuration where the retained-budget + // safety net can fire. `previous` is what a 20 second template job resolves. + const launchdOwnedStop = { + previous: { timeoutMs: 15_000, nativeStopBudget: true }, + acceptedAtMs: Number.MAX_SAFE_INTEGER, + }; + + // An in-process restart signals the Gateway without launchd running the stop, so + // the job still prints `running`. That is a confirmed answer, not a failed probe, + // and claiming the deadline "could not be confirmed" there would be false on every + // in-process restart of a default macOS install. + it("does not claim an unconfirmed timeout when launchd is confirmed not to be stopping", async () => { + delete process.env.OPENCLAW_SUPERVISOR_MODE; + execLaunchctl.mockResolvedValue(printed("running", "\texit timeout = 20\n\tpid = 4242\n")); + const warn = vi.fn(); + const budget = await resolveGatewayShutdownBudget( + "launchd", + { info: vi.fn(), warn }, + launchdOwnedStop, + ); + expect(warn).not.toHaveBeenCalled(); + expect(budget.timeoutMs).toBe(15_000); + expect(budget.nativeStopBudget).toBe(true); + }); + + // A probe that established nothing is the case the safety net exists for, so the + // startup budget is held rather than widened. + it("retains the startup budget when the job could not be inspected at all", async () => { + delete process.env.OPENCLAW_SUPERVISOR_MODE; + execLaunchctl.mockResolvedValue({ + code: 1, + stdout: "", + stderr: "permission denied", + termination: "exit", + }); + const warn = vi.fn(); + const budget = await resolveGatewayShutdownBudget( + "launchd", + { info: vi.fn(), warn }, + launchdOwnedStop, + ); + expect(warn).toHaveBeenCalledWith(expect.stringContaining("Unable to inspect the launchd job")); + expect(warn).toHaveBeenCalledWith( + "Retaining the startup shutdown budget of 15000ms because the current supervisor stop timeout could not be confirmed.", + ); + expect(budget.timeoutMs).toBe(15_000); + }); + + // Warned but non-null is the defaulted-value case, not a failed probe: launchd is + // confirmed to be stopping the job, so a clock is running and there is nothing to + // retain. Warning about a timeout that "could not be confirmed" here would be the + // same false statement in a different place. + it("does not retain when only the deadline's value had to be defaulted", async () => { + delete process.env.OPENCLAW_SUPERVISOR_MODE; + execLaunchctl.mockResolvedValue(printed("SIGTERMed", "\tpid = 4242\n")); + const warn = vi.fn(); + const budget = await resolveGatewayShutdownBudget( + "launchd", + { info: vi.fn(), warn }, + launchdOwnedStop, + ); + expect(warn).toHaveBeenCalledExactlyOnceWith( + expect.stringContaining("its exit timeout is missing or invalid"), + ); + expect(budget.timeoutMs).toBe(15_000); + }); + + // No stop is running at startup, so there is no enforcing deadline to read and + // no reason to spend a launchctl print discovering that. + it("does not inspect the job at startup", async () => { + const info = vi.fn(); + const budget = await resolveGatewayShutdownBudget("external", { info, warn: vi.fn() }); + budget.log("startup"); + expect(execLaunchctl).not.toHaveBeenCalled(); + expect(budget.timeoutMs).toBe(325_000); + expect(budget.nativeStopBudget).toBe(false); + expect(info).toHaveBeenCalledWith( + "shutdown budget at startup: drain=315000ms shutdown=325000ms reserve=10000ms exitMargin=5000ms; source=Gateway stop policy=330000ms", + ); + }); + + it("keeps the platform-neutral policy when darwin is not running a launchd job", async () => { + process.env = {}; + const warn = vi.fn(); + const budget = await resolveGatewayShutdownBudget(null, { info: vi.fn(), warn }, stoppingNow); + expect(budget.timeoutMs).toBe(325_000); + expect(budget.nativeStopBudget).toBe(false); + expect(warn).not.toHaveBeenCalled(); + expect(execLaunchctl).not.toHaveBeenCalled(); + }); + + it("never reads systemd on darwin", async () => { + execLaunchctl.mockResolvedValue(printed("SIGTERMed", "\texit timeout = 20\n\tpid = 4242\n")); + await resolveGatewayShutdownBudget("external", { info: vi.fn(), warn: vi.fn() }, stoppingNow); + expect(execSystem).not.toHaveBeenCalled(); + expect(execUser).not.toHaveBeenCalled(); + expect(readFile).not.toHaveBeenCalled(); + }); }); diff --git a/src/cli/gateway-cli/run-loop-shutdown-budget.ts b/src/cli/gateway-cli/run-loop-shutdown-budget.ts index d8c56bd22e05..96eda1904f68 100644 --- a/src/cli/gateway-cli/run-loop-shutdown-budget.ts +++ b/src/cli/gateway-cli/run-loop-shutdown-budget.ts @@ -2,13 +2,58 @@ import { performance } from "node:perf_hooks"; import { LAUNCH_AGENT_EXIT_TIMEOUT_SECONDS } from "../../daemon/launchd-plist.js"; import { GATEWAY_SERVICE_STOP_TIMEOUT_MS, - GATEWAY_SHUTDOWN_RESERVE_MS, GATEWAY_SHUTDOWN_TIMEOUT_MS, - GATEWAY_SUPERVISOR_EXIT_MARGIN_MS, + resolveShutdownReserveMs, + resolveSupervisorExitMarginMs, } from "../../infra/gateway-shutdown-budget.js"; +import { readLaunchdStopTimeout } from "../../infra/launchd-stop-timeout.js"; import { readSystemdStopTimeout } from "../../infra/systemd-stop-timeout.js"; import type { GatewayRunSignalAction } from "./run-loop-request.js"; +type NativeStopTimeout = { timeoutMs: number; source: string }; + +/** + * Ask whichever supervisor actually enforces the deadline on this platform. + * + * Three independent answers. `stop` is a deadline that may be spent as a native + * stop budget. `warning` is what the operator needs to hear. `inconclusive` says + * the probe could not establish an answer at all, which is the only case the + * retained-budget safety net below is for: a read that positively determined no + * launchd deadline governs this stop is an answer, not a failure, so retaining a + * startup budget and warning that the timeout "could not be confirmed" would be + * false on every in-process restart of a launchd-owned Gateway. + */ +async function readNativeStopTimeout(stopping: boolean): Promise<{ + stop: NativeStopTimeout | null; + warning?: string; + inconclusive: boolean; +}> { + if (process.platform === "linux") { + const systemd = await readSystemdStopTimeout(); + // Unchanged from the linux-only original: absent unit or warned read both + // count as unconfirmed there. + return { + stop: systemd, + warning: systemd?.warning, + inconclusive: !systemd || Boolean(systemd.warning), + }; + } + // launchd's ExitTimeOut bounds a stop that launchd is running and nothing else: + // an externally delivered SIGTERM never starts that clock, and the job outlives + // the deadline untouched. There is no enforcing deadline to read before a stop + // is under way, and reading one at startup would spend a launchctl print only + // to adopt a deadline that does not govern the stop the Gateway will get. + if (process.platform === "darwin" && stopping) { + const read = await readLaunchdStopTimeout(); + // Warned but non-null is the defaulted-value case: launchd is confirmed to be + // stopping the job and only its deadline had to be guessed, so a clock is + // genuinely running and nothing needs retaining. Only a warning with no + // deadline at all means the probe established nothing. + return { ...read, inconclusive: read.stop === null && read.warning !== undefined }; + } + return { stop: null, inconclusive: false }; +} + export async function resolveGatewayShutdownBudget( supervisor: string | null, logger: { info(message: string): void; warn(message: string): void }, @@ -17,37 +62,45 @@ export async function resolveGatewayShutdownBudget( acceptedAtMs: number; }, ) { - // Restart ownership may be external while systemd still enforces the stop deadline. - const systemdStop = process.platform === "linux" ? await readSystemdStopTimeout() : null; + // Restart ownership may be external while the platform supervisor still + // enforces the stop deadline. That holds on darwin exactly as it does on linux. + const native = await readNativeStopTimeout(refresh !== undefined); + const nativeStop = native.stop; const retained = - refresh?.previous.nativeStopBudget && (!systemdStop || systemdStop.warning) - ? refresh.previous - : undefined; - if (systemdStop?.warning) { - logger.warn(systemdStop.warning); + refresh?.previous.nativeStopBudget && native.inconclusive ? refresh.previous : undefined; + if (native.warning) { + logger.warn(native.warning); } if (retained) { logger.warn( - `Retaining the startup shutdown budget of ${retained.timeoutMs}ms because the current systemd stop timeout could not be confirmed.`, + `Retaining the startup shutdown budget of ${retained.timeoutMs}ms because the current supervisor stop timeout could not be confirmed.`, ); } - const stop = systemdStop ?? { + const stop = nativeStop ?? { timeoutMs: supervisor === "launchd" ? LAUNCH_AGENT_EXIT_TIMEOUT_SECONDS * 1_000 : GATEWAY_SERVICE_STOP_TIMEOUT_MS, source: supervisor === "launchd" ? "launchd ExitTimeOut" : "Gateway stop policy", }; - const nativeStopBudget = systemdStop !== null || supervisor === "launchd" || Boolean(retained); + // ExitTimeOut=0 is unlimited. It is an observed job value, not a native + // deadline that may cap a requested restart or arm a forced exit. + const nativeStopBudget = nativeStop + ? Number.isFinite(nativeStop.timeoutMs) + : supervisor === "launchd" || Boolean(retained); + // An operator job may enforce a deadline far shorter than the policy these fixed + // allowances were sized against, so each is capped at a share of what it is carved + // from. A deadline long enough to fund them is unaffected; a short one keeps a + // proportional drain instead of surrendering all of it to margin and reserve. + const exitMarginMs = resolveSupervisorExitMarginMs(stop.timeoutMs); const limitMs = - retained?.timeoutMs ?? - Math.min(GATEWAY_SHUTDOWN_TIMEOUT_MS, stop.timeoutMs - GATEWAY_SUPERVISOR_EXIT_MARGIN_MS); + retained?.timeoutMs ?? Math.min(GATEWAY_SHUTDOWN_TIMEOUT_MS, stop.timeoutMs - exitMarginMs); const elapsedMs = refresh && nativeStopBudget ? Math.max(0, Math.ceil(performance.now() - refresh.acceptedAtMs)) : 0; const timeoutMs = Math.max(0, limitMs - elapsedMs); - const reserveMs = Math.min(GATEWAY_SHUTDOWN_RESERVE_MS, timeoutMs); + const reserveMs = resolveShutdownReserveMs(timeoutMs); return { nativeStopBudget, timeoutMs, @@ -64,7 +117,10 @@ export async function resolveGatewayShutdownBudget( }, log: (phase: "startup" | "shutdown") => { logger.info( - `shutdown budget at ${phase}: drain=${Math.max(0, timeoutMs - GATEWAY_SHUTDOWN_RESERVE_MS)}ms shutdown=${timeoutMs}ms reserve=${reserveMs}ms exitMargin=${GATEWAY_SUPERVISOR_EXIT_MARGIN_MS}ms; source=${retained ? `startup shutdown budget=${retained.timeoutMs}ms` : `${stop.source}=${stop.timeoutMs}ms`}`, + // Report the drain and margin actually spent. Subtracting the unscaled + // reserve constant here understated a short budget's drain by the amount the + // scaled reserve gave back. + `shutdown budget at ${phase}: drain=${Math.max(0, timeoutMs - reserveMs)}ms shutdown=${timeoutMs}ms reserve=${reserveMs}ms exitMargin=${exitMarginMs}ms; source=${retained ? `startup shutdown budget=${retained.timeoutMs}ms` : `${stop.source}=${stop.timeoutMs}ms`}`, ); }, }; diff --git a/src/cli/gateway-cli/run-loop.launchd.test.ts b/src/cli/gateway-cli/run-loop.launchd.test.ts new file mode 100644 index 000000000000..97264ba7a7b0 --- /dev/null +++ b/src/cli/gateway-cli/run-loop.launchd.test.ts @@ -0,0 +1,678 @@ +// darwin launchd-supervised run-loop cases. These live beside run-loop.test.ts because +// the darwin stop budget reads the launchd job on every stop/restart request, so each +// case here has to state the deadline launchd is enforcing instead of letting the real +// reader spawn launchctl print, and run-loop.test.ts is at its line cap. +import { performance } from "node:perf_hooks"; +import { afterEach, beforeEach, describe, expect, it, vi } from "vitest"; +import { useAutoCleanupTempDirTracker } from "../../../test/helpers/temp-dir.js"; +import type { HostedGatewayStop } from "../../daemon/hosted-stop.js"; +import type { GatewayActiveWorkSnapshot } from "../../infra/gateway-active-work.js"; +import type { GatewayBootLifecycleCompletion } from "../../infra/gateway-boot-lifecycle.js"; +import type { GatewayRestartIntent } from "../../infra/restart-intent.js"; +import { SUPERVISOR_HINT_ENV_VARS } from "../../infra/supervisor-markers.js"; +import type { RuntimeEnv } from "../../runtime.js"; +import { captureEnv, deleteTestEnvValue } from "../../test-utils/env.js"; +import type { GatewayRestartSnapshot } from "../daemon-cli/restart-health.js"; +import { + createActiveWorkSnapshot, + createCloseMock, + createRuntimeWithExitSignal, + createSignaledStart, + originalPlatformDescriptor, + type UpdateRespawnResultFixture, + setPlatform, + waitForStart, + withIsolatedSignals, +} from "./run-loop.test-support.js"; + +useAutoCleanupTempDirTracker(afterEach); +const { readCgroup } = vi.hoisted(() => ({ readCgroup: vi.fn() })); + +const spawnProcess = vi.hoisted(() => vi.fn()); +vi.mock("node:child_process", async () => { + const actual = await vi.importActual("node:child_process"); + const { mockNodeChildProcessModule } = + await import("../../gateway/server-methods/node-child-process.test-support.js"); + if (!spawnProcess.getMockImplementation()) { + spawnProcess.mockImplementation(actual.spawn); + } + const mocked = await mockNodeChildProcessModule({}); + vi.spyOn(mocked, "spawn").mockImplementation(spawnProcess); + return mocked; +}); + +vi.mock("node:fs/promises", async (original) => { + const actual = await original(); + // Foreground fixtures must not inherit the CI runner's systemd service or filesystem timing. + const readFile = (...args: Parameters) => + args[0] === "/proc/self/cgroup" ? readCgroup() : actual.readFile(...args); + return { ...actual, readFile, default: { ...actual, readFile } }; +}); + +const systemctl = vi.fn(async () => ({ + code: 0, + stdout: "LoadState=loaded\nTimeoutStopUSec=5min 30s", + stderr: "", +})); +vi.mock("../../daemon/systemd-exec.js", () => ({ + execSystemctl: () => systemctl(), + execSystemctlUser: () => systemctl(), +})); + +// darwin now refreshes the stop budget on every stop/restart request, and the real +// reader spawns launchctl print against three domains through the mocked +// node:child_process module. Keep launchctl out of the loop and let each case state +// the deadline launchd is enforcing. The reader reports stop: null when launchd is not +// running the stop, which leaves the platform-neutral policy in force. +const readLaunchdStopTimeout = vi.fn< + typeof import("../../infra/launchd-stop-timeout.js").readLaunchdStopTimeout +>(async () => ({ stop: null })); +vi.mock("../../infra/launchd-stop-timeout.js", () => ({ + readLaunchdStopTimeout: (...args: Parameters) => + readLaunchdStopTimeout(...args), +})); + +const acquireGatewayLock = vi.fn(async (_opts?: { port?: number }) => ({ + release: vi.fn(async () => {}), +})); +const hostedStopExecute = vi.fn(); +const hostedStopDispose = vi.fn(); +const hostedStopPrepare = + vi.fn(); +vi.mock("../../daemon/hosted-stop.js", () => ({ + prepareHostedGatewayStop: (...args: Parameters) => + hostedStopPrepare(...args), +})); +const consumeGatewayRestartIntentPayloadSync = vi.fn< + () => { reason?: string; force?: boolean; waitMs?: number } | null +>(() => null); +const consumeGatewayRestartIntent = vi.fn<() => GatewayRestartIntent | null>(() => null); +type ManagedUpdateOwner = NonNullable; +const cancelManagedServiceUpdateHandoff = vi.fn< + (_identity: ManagedUpdateOwner) => Promise +>(async () => "restored-in-process"); +const claimManagedServiceUpdateHandoff = vi.fn((_identity: ManagedUpdateOwner) => true); +const isForegroundUpdateHandoff = vi.fn((_identity: ManagedUpdateOwner) => false); +const completeForegroundUpdateHandoffAfterClose = + vi.fn< + typeof import("../../infra/update-managed-service-handoff.js").completeForegroundUpdateHandoffAfterClose + >(); +const captureForegroundUpdateHandoffStop = + vi.fn< + typeof import("../../infra/update-managed-service-handoff.js").captureForegroundUpdateHandoffStop + >(); +const requestManagedServiceUpdateHandoffPark = vi.fn(async (_identity: ManagedUpdateOwner) => true); +const waitForSystemServiceUpdateHandoffs = vi.fn<() => Promise | undefined>(); +const commitManagedServiceUpdateHandoff = vi.fn( + async (_identity: ManagedUpdateOwner, _outcome?: "update" | "restore") => true, +); +const consumeGatewayRestartAuthorization = vi.fn(() => true); +const consumeGatewayRestartIntentSync = vi.fn(() => false); +const isGatewayRestartExternallyAllowed = vi.fn(() => false); +const markGatewayRestartHandled = vi.fn(); +const peekGatewayRestartReason = vi.fn<() => string | undefined>(() => undefined); +const resetGatewayRestartStateForInProcessRestart = vi.fn(); +const resetGatewaySuspendCoordinatorForLifecycleRestart = vi.fn(); +const consumeGatewaySuspendHandoff = + vi.fn(); +const disarmGatewaySuspendHandoff = vi.fn(); +const rollbackGatewayRestartSignalAdmission = vi.fn(); +const requestGatewayRestartWithSignalAdmission = vi.fn(() => ({ status: "emitted" as const })); +const writeGatewayRestartHandoffSync = vi.fn( + ( + _opts: unknown, + ): { + kind: "gateway-supervisor-restart-handoff"; + version: 1; + intentId: string; + pid: number; + createdAt: number; + expiresAt: number; + source: "unknown"; + restartKind: "full-process"; + supervisorMode: "external"; + } | null => ({ + kind: "gateway-supervisor-restart-handoff", + version: 1, + intentId: "test-intent", + pid: process.pid, + createdAt: Date.now(), + expiresAt: Date.now() + 60_000, + source: "unknown", + restartKind: "full-process", + supervisorMode: "external", + }), +); +const scheduleGatewayRestart = vi.fn((_opts?: { delayMs?: number; reason?: string }) => ({ + ok: true, + pid: process.pid, + signal: "SIGUSR2" as const, + delayMs: 0, + mode: "emit" as const, + coalesced: false, + cooldownMsApplied: 0, +})); +const idleActiveWorkSnapshot = createActiveWorkSnapshot(); +const createGatewayActiveWorkSnapshot = vi.fn(() => idleActiveWorkSnapshot); +const waitForGatewayActiveWork = vi.fn( + async ( + _timeoutMs?: number, + options?: { onSnapshot?: (snapshot: GatewayActiveWorkSnapshot) => void }, + ) => { + const snapshot = createGatewayActiveWorkSnapshot(); + options?.onSnapshot?.(snapshot); + return { drained: snapshot.idle, snapshot }; + }, +); +const advanceCronActiveJobGeneration = vi.fn(); +const resetCronActiveJobs = vi.fn(); +const abortActiveCronTaskRuns = vi.fn((_reason?: string) => 0); +const retireActiveCronTaskRunTracking = vi.fn(); +const waitForActiveCronTaskRuns = vi.fn(async (_timeoutMs?: number) => ({ + drained: true, + active: 0, +})); +const waitForActiveCronJobs = vi.fn(async (_timeoutMs?: number) => ({ + drained: true, + active: 0, +})); +const reloadTaskRuntimeStateFromStore = vi.fn(); +const clearRuntimeConfigSnapshot = vi.fn(); +const restartGatewayProcessWithFreshPid = vi.fn< + (_opts?: { env?: NodeJS.ProcessEnv }) => { + mode: "supervised" | "disabled" | "failed"; + detail?: string; + exitCode?: number; + handoffSpawned?: Promise; + } +>(() => ({ mode: "disabled" })); +const respawnGatewayProcessForUpdate = vi.fn< + (_opts?: { env?: NodeJS.ProcessEnv }) => UpdateRespawnResultFixture +>(() => ({ mode: "disabled", detail: "OPENCLAW_NO_RESPAWN" })); +const { killProcessTree } = vi.hoisted(() => ({ killProcessTree: vi.fn() })); +vi.mock("../../process/kill-tree.js", async (importOriginal) => ({ + ...(await importOriginal()), + killProcessTree, +})); +const markUpdateRestartSentinelFailure = vi.fn<(reason: string) => Promise>(async () => null); +const writeRestartSentinelIfUnchanged = vi.fn< + typeof import("../../infra/restart-sentinel.js").writeRestartSentinelIfUnchanged +>(async () => null); +const readRestartSentinelReadOnly = + vi.fn(); +const waitForGatewayHealthyRestart = + vi.fn(); +vi.mock("../daemon-cli/restart-health.js", async (importOriginal) => ({ + ...(await importOriginal()), + waitForGatewayHealthyRestart: (...args: Parameters) => + waitForGatewayHealthyRestart(...args), +})); +const respawnHealth = ( + overrides: Partial = {}, +): GatewayRestartSnapshot => ({ + runtime: { status: "running", pid: 7777 }, + portUsage: { port: 18789, status: "busy", listeners: [{ pid: 7777 }], hints: [] }, + healthy: true, + waitOutcome: "healthy", + staleGatewayPids: [], + ...overrides, +}); +const abortPendingChannelReloads = vi.fn(); +const abortEmbeddedAgentRun = vi.fn( + (_sessionId?: string, _opts?: { mode?: "all" | "compacting"; reason?: "restart" }) => false, +); +const gatewayLog = { + debug: vi.fn(), + info: vi.fn(), + warn: vi.fn(), + error: vi.fn(), +}; +const flushLogger = vi.fn(async () => {}); +const writeDiagnosticStabilityBundleForFailureSync = vi.fn(() => ({ + message: "stability bundle recorded", +})); +const hasManagedProviderLocalServices = vi.fn(() => false); +const stopManagedProviderLocalServices = vi.fn(async () => {}); +const cancelShutdownHardExitWatchdog = vi.fn(); +const armShutdownHardExitWatchdog = vi.fn( + (_params: { delayMs: number; onError: (error: unknown) => void }) => ({ + cancel: cancelShutdownHardExitWatchdog, + }), +); + +vi.mock("../../infra/gateway-lock.js", async (original) => ({ + ...(await original()), + acquireGatewayLock: (opts?: { port?: number }) => acquireGatewayLock(opts), +})); + +vi.mock("../../infra/restart.js", async (importOriginal) => { + const actual = await importOriginal(); + return { + ...actual, + consumeGatewayRestartIntent: () => consumeGatewayRestartIntent(), + consumeGatewayRestartAuthorization: () => consumeGatewayRestartAuthorization(), + isGatewayRestartExternallyAllowed: () => isGatewayRestartExternallyAllowed(), + markGatewayRestartHandled: () => markGatewayRestartHandled(), + peekGatewayRestartReason: () => peekGatewayRestartReason(), + resetGatewayRestartStateForInProcessRestart: () => + resetGatewayRestartStateForInProcessRestart(), + rollbackGatewayRestartSignalAdmission: () => rollbackGatewayRestartSignalAdmission(), + requestGatewayRestartWithSignalAdmission, + scheduleGatewayRestart: (opts?: { delayMs?: number; reason?: string }) => + scheduleGatewayRestart(opts), + }; +}); + +vi.mock("../../infra/restart-intent.js", () => ({ + consumeGatewayRestartIntentPayloadSync: () => consumeGatewayRestartIntentPayloadSync(), + consumeGatewayRestartIntentSync: () => consumeGatewayRestartIntentSync(), +})); + +vi.mock("../../infra/update-managed-service-handoff.js", () => ({ + waitForSystemServiceUpdateHandoffs: () => waitForSystemServiceUpdateHandoffs(), + captureForegroundUpdateHandoffStop: ( + params: Parameters[0], + ) => captureForegroundUpdateHandoffStop(params), + isForegroundUpdateHandoff: (identity: ManagedUpdateOwner) => isForegroundUpdateHandoff(identity), + completeForegroundUpdateHandoffAfterClose: (identity: ManagedUpdateOwner) => + completeForegroundUpdateHandoffAfterClose(identity), + cancelManagedServiceUpdateHandoff: (identity: ManagedUpdateOwner) => + cancelManagedServiceUpdateHandoff(identity), + claimManagedServiceUpdateHandoff: (identity: ManagedUpdateOwner) => + claimManagedServiceUpdateHandoff(identity), + requestManagedServiceUpdateHandoffPark: (identity: ManagedUpdateOwner) => + requestManagedServiceUpdateHandoffPark(identity), + commitManagedServiceUpdateHandoff: ( + identity: ManagedUpdateOwner, + outcome?: "update" | "restore", + ) => commitManagedServiceUpdateHandoff(identity, outcome), +})); + +vi.mock("../../infra/gateway-suspend-coordinator.js", () => ({ + consumeGatewaySuspendHandoff: (...args: Parameters) => + consumeGatewaySuspendHandoff(...args), + disarmGatewaySuspendHandoff: (...args: unknown[]) => disarmGatewaySuspendHandoff(...args), + resetGatewaySuspendCoordinatorForLifecycleRestart: () => + resetGatewaySuspendCoordinatorForLifecycleRestart(), +})); + +vi.mock("../../infra/process-respawn.js", async (importOriginal) => ({ + ...(await importOriginal()), + respawnGatewayProcessForUpdate: (opts?: { env?: NodeJS.ProcessEnv }) => + respawnGatewayProcessForUpdate(opts), + restartGatewayProcessWithFreshPid: (opts?: { env?: NodeJS.ProcessEnv }) => + restartGatewayProcessWithFreshPid(opts), +})); + +vi.mock("../../infra/restart-sentinel.js", () => ({ + readRestartSentinelReadOnly: () => readRestartSentinelReadOnly(), + markUpdateRestartSentinelFailure: (reason: string) => markUpdateRestartSentinelFailure(reason), + writeRestartSentinelIfUnchanged: (...args: Parameters) => + writeRestartSentinelIfUnchanged(...args), +})); + +vi.mock("../../infra/restart-handoff.js", () => ({ + writeGatewayRestartHandoffSync: (opts: unknown) => writeGatewayRestartHandoffSync(opts), +})); + +vi.mock("../../infra/gateway-active-work.js", () => ({ + createGatewayActiveWorkSnapshot: () => createGatewayActiveWorkSnapshot(), + waitForGatewayActiveWork: ( + timeoutMs?: number, + options?: { onSnapshot?: (snapshot: GatewayActiveWorkSnapshot) => void }, + ) => waitForGatewayActiveWork(timeoutMs, options), +})); + +vi.mock("../../cron/active-jobs.js", () => ({ + advanceCronActiveJobGeneration: () => advanceCronActiveJobGeneration(), + resetCronActiveJobs: () => resetCronActiveJobs(), + waitForActiveCronJobs: (timeoutMs: number) => waitForActiveCronJobs(timeoutMs), +})); + +vi.mock("../../cron/service/active-run-cancellation.js", () => ({ + abortActiveCronTaskRuns: (reason?: string) => abortActiveCronTaskRuns(reason), + retireActiveCronTaskRunTracking: () => retireActiveCronTaskRunTracking(), + waitForActiveCronTaskRuns: (timeoutMs: number) => waitForActiveCronTaskRuns(timeoutMs), +})); + +vi.mock("../../tasks/runtime-internal.js", () => ({ + reloadTaskRuntimeStateFromStore: () => reloadTaskRuntimeStateFromStore(), +})); + +vi.mock("../../config/runtime-snapshot.js", () => ({ + clearRuntimeConfigSnapshot: () => clearRuntimeConfigSnapshot(), + getRuntimeConfigSourceSnapshot: () => null, + registerRuntimeConfigSnapshotPreparer: vi.fn(), +})); + +vi.mock("../../agents/embedded-agent-runner/runs.js", () => ({ + abortEmbeddedAgentRun: ( + sessionId?: string, + opts?: { mode?: "all" | "compacting"; reason?: "restart" }, + ) => abortEmbeddedAgentRun(sessionId, opts), +})); + +vi.mock("../../logging/subsystem.js", () => ({ + createSubsystemLogger: () => gatewayLog, +})); + +vi.mock("../../logging/logger.js", () => ({ + flushLogger: () => flushLogger(), +})); + +vi.mock("../../logging/diagnostic-stability-bundle.js", () => ({ + writeDiagnosticStabilityBundleForFailureSync, +})); + +vi.mock("../../agents/provider-runtime-lifecycle.js", () => ({ + hasManagedProviderLocalServices: () => hasManagedProviderLocalServices(), +})); + +vi.mock("../../agents/provider-local-service.js", () => ({ + stopManagedProviderLocalServices: () => stopManagedProviderLocalServices(), +})); + +vi.mock("../../gateway/server-reload-generation.js", () => ({ + abortPendingChannelReloads: () => abortPendingChannelReloads(), +})); + +vi.mock("./shutdown-hard-exit.js", () => ({ + armShutdownHardExitWatchdog: (params: { delayMs: number; onError: (error: unknown) => void }) => + armShutdownHardExitWatchdog(params), +})); + +async function runLoopWithStart(params: { + start: ReturnType; + runtime: RuntimeEnv; + ownsProcessLifecycle?: boolean; + lockPort?: number; + healthHost?: string; + beginBoot?: (startedAtMs: number) => void | Promise; + completeBoot?: (completion: GatewayBootLifecycleCompletion) => void; +}) { + vi.resetModules(); + const { runGatewayLoop } = await import("./run-loop.js"); + const loopPromise = runGatewayLoop({ + start: params.start as unknown as Parameters[0]["start"], + runtime: params.runtime, + ownsProcessLifecycle: params.ownsProcessLifecycle, + lockPort: params.lockPort, + healthHost: params.healthHost, + beginBoot: params.beginBoot, + completeBoot: params.completeBoot, + }); + return { loopPromise }; +} + +async function createSignaledLoopHarness(exitCallOrder?: string[], ownsProcessLifecycle = false) { + const close = createCloseMock(); + const { start, started } = createSignaledStart(close); + const { runtime, exited } = createRuntimeWithExitSignal(exitCallOrder); + const { loopPromise } = await runLoopWithStart({ start, runtime, ownsProcessLifecycle }); + await waitForStart(started); + return { close, start, runtime, exited, loopPromise }; +} + +function expectRestartHandoffCall(expected: { + restartKind: "full-process" | "update-process"; + reason: string | undefined; + supervisorMode: "external" | "launchd"; +}) { + expect(writeGatewayRestartHandoffSync).toHaveBeenCalledTimes(1); + const [handoff] = writeGatewayRestartHandoffSync.mock.calls[0] ?? []; + if (!handoff || typeof handoff !== "object" || Array.isArray(handoff)) { + throw new Error("expected restart handoff options object"); + } + const processInstanceId = (handoff as { processInstanceId?: unknown }).processInstanceId; + expect(typeof processInstanceId).toBe("string"); + if (typeof processInstanceId !== "string") { + throw new Error("expected restart handoff processInstanceId string"); + } + expect(processInstanceId).toMatch( + /^[0-9a-f]{8}-[0-9a-f]{4}-[1-5][0-9a-f]{3}-[89ab][0-9a-f]{3}-[0-9a-f]{12}$/, + ); + expect(handoff).toEqual({ + ...expected, + processInstanceId, + }); +} + +let gatewayWorkAdmissionActual: typeof import("../../process/gateway-work-admission.js"); +let supervisorEnvSnapshot: ReturnType | undefined; + +beforeEach(async () => { + vi.useRealTimers(); + vi.clearAllMocks(); + spawnProcess + .mockReset() + .mockImplementation( + (await vi.importActual("node:child_process")).spawn, + ); + acquireGatewayLock.mockReset().mockImplementation(async () => ({ + release: vi.fn(async () => {}), + })); + setPlatform("linux"); + readCgroup.mockReset().mockResolvedValue("0::/\n"); + systemctl.mockReset().mockResolvedValue({ + code: 0, + stdout: "LoadState=loaded\nTimeoutStopUSec=5min 30s", + stderr: "", + }); + // mockReset also drops any one-shot launchd deadline a previous case left queued. + readLaunchdStopTimeout.mockReset().mockResolvedValue({ stop: null }); + hostedStopExecute.mockReset().mockResolvedValue({ outcome: "accepted" }); + hostedStopDispose.mockReset().mockResolvedValue(undefined); + hostedStopPrepare.mockReset().mockImplementation(async (_owner, assertCurrent) => { + assertCurrent(); + return { execute: hostedStopExecute, dispose: hostedStopDispose }; + }); + supervisorEnvSnapshot = captureEnv([...SUPERVISOR_HINT_ENV_VARS, "OPENCLAW_NO_RESPAWN"]); + for (const key of [...SUPERVISOR_HINT_ENV_VARS, "OPENCLAW_NO_RESPAWN"]) { + deleteTestEnvValue(key); + } + + // clearAllMocks preserves queued one-shot results. A skipped lifecycle branch + // must not shift a stale supervisor or respawn decision into the next case. + consumeGatewayRestartIntent.mockReset(); + consumeGatewayRestartIntentPayloadSync.mockReset().mockReturnValue(null); + consumeGatewaySuspendHandoff.mockReset().mockReturnValue({ ok: true, value: false }); + disarmGatewaySuspendHandoff.mockClear(); + consumeGatewayRestartIntent.mockReturnValue(null); + peekGatewayRestartReason.mockReset(); + peekGatewayRestartReason.mockReturnValue(undefined); + restartGatewayProcessWithFreshPid.mockReset(); + restartGatewayProcessWithFreshPid.mockReturnValue({ mode: "disabled" }); + respawnGatewayProcessForUpdate.mockReset(); + waitForGatewayHealthyRestart.mockReset().mockResolvedValue(respawnHealth()); + writeRestartSentinelIfUnchanged.mockReset().mockResolvedValue(null); + readRestartSentinelReadOnly.mockReset().mockResolvedValue(null); + respawnGatewayProcessForUpdate.mockReturnValue({ + mode: "disabled", + detail: "OPENCLAW_NO_RESPAWN", + }); + hasManagedProviderLocalServices.mockReset(); + hasManagedProviderLocalServices.mockReturnValue(false); + stopManagedProviderLocalServices.mockReset(); + stopManagedProviderLocalServices.mockResolvedValue(undefined); + + gatewayWorkAdmissionActual = await vi.importActual("../../process/gateway-work-admission.js"); + gatewayWorkAdmissionActual.resetGatewayWorkAdmission(); + createGatewayActiveWorkSnapshot.mockReset(); + createGatewayActiveWorkSnapshot.mockReturnValue(idleActiveWorkSnapshot); + waitForGatewayActiveWork.mockReset(); + waitForGatewayActiveWork.mockImplementation(async (_timeoutMs, options) => { + const snapshot = createGatewayActiveWorkSnapshot(); + options?.onSnapshot?.(snapshot); + return { drained: snapshot.idle, snapshot }; + }); + cancelManagedServiceUpdateHandoff.mockReset(); + cancelManagedServiceUpdateHandoff.mockResolvedValue("restored-in-process"); + claimManagedServiceUpdateHandoff.mockReset(); + claimManagedServiceUpdateHandoff.mockReturnValue(true); + isForegroundUpdateHandoff.mockReset().mockReturnValue(false); + completeForegroundUpdateHandoffAfterClose.mockReset().mockResolvedValue({ respawn: true }); + captureForegroundUpdateHandoffStop.mockReset().mockReturnValue(undefined); + requestManagedServiceUpdateHandoffPark.mockReset(); + requestManagedServiceUpdateHandoffPark.mockResolvedValue(true); + waitForSystemServiceUpdateHandoffs.mockReset().mockReturnValue(undefined); + commitManagedServiceUpdateHandoff.mockReset(); + commitManagedServiceUpdateHandoff.mockResolvedValue(true); +}); + +afterEach(() => { + supervisorEnvSnapshot?.restore(); + supervisorEnvSnapshot = undefined; + vi.useRealTimers(); + if (originalPlatformDescriptor) { + Object.defineProperty(process, "platform", originalPlatformDescriptor); + } +}); + +describe("runGatewayLoop darwin launchd supervision", () => { + it("waits briefly before exiting on launchd supervised restart", async () => { + vi.clearAllMocks(); + peekGatewayRestartReason.mockReturnValue(undefined); + try { + setPlatform("darwin"); + process.env.OPENCLAW_LAUNCHD_LABEL = "ai.openclaw.gateway"; + restartGatewayProcessWithFreshPid.mockReturnValueOnce({ + mode: "supervised", + handoffSpawned: Promise.resolve(true), + }); + + await withIsolatedSignals(async ({ captureSignal }) => { + const { runtime, exited } = await createSignaledLoopHarness(); + const restartSignal = captureSignal("SIGUSR2"); + + vi.useFakeTimers(); + restartSignal(); + await vi.advanceTimersByTimeAsync(1499); + expect(runtime.exit).not.toHaveBeenCalled(); + await vi.advanceTimersByTimeAsync(1); + + await expect(exited).resolves.toBe(0); + expect(runtime.exit).toHaveBeenCalledWith(0); + expectRestartHandoffCall({ + restartKind: "full-process", + reason: undefined, + supervisorMode: "launchd", + }); + }); + } finally { + vi.useRealTimers(); + delete process.env.OPENCLAW_LAUNCHD_LABEL; + if (originalPlatformDescriptor) { + Object.defineProperty(process, "platform", originalPlatformDescriptor); + } + } + }); + + it("falls back in-process when the launchd restart handoff fails to spawn", async () => { + vi.clearAllMocks(); + peekGatewayRestartReason.mockReturnValue(undefined); + try { + setPlatform("darwin"); + process.env.OPENCLAW_LAUNCHD_LABEL = "ai.openclaw.gateway"; + restartGatewayProcessWithFreshPid.mockReturnValueOnce({ + mode: "supervised", + handoffSpawned: Promise.resolve(false), + }); + + await withIsolatedSignals(async ({ captureSignal }) => { + const { start, runtime, exited } = await createSignaledLoopHarness(); + const restartSignal = captureSignal("SIGUSR2"); + const sigint = captureSignal("SIGINT"); + + vi.useFakeTimers(); + restartSignal(); + await vi.advanceTimersByTimeAsync(1500); + + expect(start).toHaveBeenCalledTimes(2); + expect(runtime.exit).not.toHaveBeenCalled(); + expect(acquireGatewayLock).toHaveBeenCalledTimes(2); + expect(gatewayLog.warn).toHaveBeenCalledWith( + "launchd restart handoff failed to spawn; falling back to in-process restart", + ); + + sigint(); + await expect(exited).resolves.toBe(0); + }); + } finally { + vi.useRealTimers(); + delete process.env.OPENCLAW_LAUNCHD_LABEL; + if (originalPlatformDescriptor) { + Object.defineProperty(process, "platform", originalPlatformDescriptor); + } + } + }); + + it("leaves the successor to launchd after a SIGTERM restart intent", async () => { + vi.clearAllMocks(); + consumeGatewayRestartIntentPayloadSync.mockReturnValueOnce({ reason: "gateway.restart" }); + setPlatform("darwin"); + process.env.OPENCLAW_LAUNCHD_LABEL = "ai.openclaw.gateway"; + restartGatewayProcessWithFreshPid.mockReturnValueOnce({ + mode: "supervised", + handoffSpawned: Promise.resolve(true), + }); + + await withIsolatedSignals(async ({ captureSignal }) => { + const { start, exited } = await createSignaledLoopHarness(); + captureSignal("SIGTERM")(); + await expect(exited).resolves.toBe(0); + expect(start).toHaveBeenCalledOnce(); + expect(restartGatewayProcessWithFreshPid).not.toHaveBeenCalled(); + expect(respawnGatewayProcessForUpdate).not.toHaveBeenCalled(); + }); + }); + + // The stop budget refresh is the reason darwin reads the launchd job at all, so prove + // the job deadline it reports is what bounds the stop and arms the force exit. + it("bounds a launchd-supervised stop on the deadline the printed job reports", async () => { + vi.clearAllMocks(); + try { + setPlatform("darwin"); + process.env.OPENCLAW_LAUNCHD_LABEL = "ai.openclaw.gateway"; + readLaunchdStopTimeout.mockResolvedValue({ + stop: { timeoutMs: 30_000, source: "launchd system/ai.openclaw.gateway exit timeout" }, + }); + hasManagedProviderLocalServices.mockReturnValue(true); + stopManagedProviderLocalServices.mockReturnValue(new Promise(() => {})); + + await withIsolatedSignals(async ({ captureSignal }) => { + const { close, runtime } = await createSignaledLoopHarness(); + + vi.useFakeTimers(); + const clock = vi.spyOn(performance, "now").mockImplementation(() => Date.now()); + try { + captureSignal("SIGTERM")(); + await vi.advanceTimersByTimeAsync(24_999); + + expect(close).toHaveBeenCalledOnce(); + expect(stopManagedProviderLocalServices).toHaveBeenCalledOnce(); + expect(runtime.exit).not.toHaveBeenCalled(); + await vi.advanceTimersByTimeAsync(1); + + // launchd owns the successor, so an abandoned drain still exits 0 for the job. + expect(runtime.exit).toHaveBeenCalledExactlyOnceWith(0); + expect(writeDiagnosticStabilityBundleForFailureSync).toHaveBeenCalledWith( + "gateway.stop_shutdown_timeout", + undefined, + ); + expect(gatewayLog.info).toHaveBeenCalledWith( + "shutdown budget at shutdown: drain=15000ms shutdown=25000ms reserve=10000ms exitMargin=5000ms; source=launchd system/ai.openclaw.gateway exit timeout=30000ms", + ); + } finally { + clock.mockRestore(); + vi.clearAllTimers(); + vi.useRealTimers(); + } + }); + } finally { + delete process.env.OPENCLAW_LAUNCHD_LABEL; + if (originalPlatformDescriptor) { + Object.defineProperty(process, "platform", originalPlatformDescriptor); + } + } + }); +}); diff --git a/src/cli/gateway-cli/run-loop.model-acquisition.process.test.ts b/src/cli/gateway-cli/run-loop.model-acquisition.process.test.ts index 1151d63c4689..244aac1f1f30 100644 --- a/src/cli/gateway-cli/run-loop.model-acquisition.process.test.ts +++ b/src/cli/gateway-cli/run-loop.model-acquisition.process.test.ts @@ -79,9 +79,23 @@ it expect(exit, output).toEqual([0, null]); expect(elapsed, output).toBeLessThan(stopTimeoutMs); expect(output).toContain("process proof: acquisition-cancelled"); + // The fixture's launchd label is synthetic, so the per-stop re-inspection cannot + // find a job to read and the startup budget is retained instead. Both + // attributions prove the same thing this case is guarding: the shutdown budget + // came from the 20 second native stop timeout and not from the platform-neutral + // policy, which would report a 330000ms source and a far longer deadline. expect(output).toMatch( - new RegExp(`shutdown budget at shutdown:.*source=.*=${stopTimeoutMs}ms`), + new RegExp( + `shutdown budget at shutdown:.*source=(?:.*=${stopTimeoutMs}ms|startup shutdown budget=${shutdownTimeoutMs}ms)`, + ), ); + if (process.platform === "darwin") { + // The label resolves to no real job, so this cannot assert a successful read. + // What it does assert is that the darwin probe ran inside a real spawned + // Gateway on a real stop: only the launchd reader emits this, and reverting the + // darwin dispatch removes it. + expect(output).toContain("Unable to inspect the launchd job"); + } if (mode === "cooperative") { expect(output).toContain("process proof: acquisition-joined"); expect(output).not.toContain("shutdown deadline reached"); diff --git a/src/cli/gateway-cli/run-loop.test.ts b/src/cli/gateway-cli/run-loop.test.ts index 64b40d4439f0..905efa847311 100644 --- a/src/cli/gateway-cli/run-loop.test.ts +++ b/src/cli/gateway-cli/run-loop.test.ts @@ -78,6 +78,19 @@ vi.mock("../../daemon/systemd-exec.js", () => ({ execSystemctlUser: () => systemctl(), })); +// darwin now refreshes the stop budget on every stop/restart request, and the real +// reader spawns launchctl print against three domains through the mocked +// node:child_process module. Keep launchctl out of the loop and let each case state +// the deadline launchd is enforcing. The reader reports stop: null when launchd is not +// running the stop, which leaves the platform-neutral policy in force. +const readLaunchdStopTimeout = vi.fn< + typeof import("../../infra/launchd-stop-timeout.js").readLaunchdStopTimeout +>(async () => ({ stop: null })); +vi.mock("../../infra/launchd-stop-timeout.js", () => ({ + readLaunchdStopTimeout: (...args: Parameters) => + readLaunchdStopTimeout(...args), +})); + const acquireGatewayLock = vi.fn(async (_opts?: { port?: number }) => ({ release: vi.fn(async () => {}), })); @@ -468,6 +481,8 @@ beforeEach(async () => { stdout: "LoadState=loaded\nTimeoutStopUSec=5min 30s", stderr: "", }); + // mockReset also drops any one-shot launchd deadline a previous case left queued. + readLaunchdStopTimeout.mockReset().mockResolvedValue({ stop: null }); hostedStopExecute.mockReset().mockResolvedValue({ outcome: "accepted" }); hostedStopDispose.mockReset().mockResolvedValue(undefined); hostedStopPrepare.mockReset().mockImplementation(async (_owner, assertCurrent) => { @@ -2794,103 +2809,6 @@ describe("runGatewayLoop", () => { } }); - it("waits briefly before exiting on launchd supervised restart", async () => { - vi.clearAllMocks(); - peekGatewayRestartReason.mockReturnValue(undefined); - try { - setPlatform("darwin"); - process.env.OPENCLAW_LAUNCHD_LABEL = "ai.openclaw.gateway"; - restartGatewayProcessWithFreshPid.mockReturnValueOnce({ - mode: "supervised", - handoffSpawned: Promise.resolve(true), - }); - - await withIsolatedSignals(async ({ captureSignal }) => { - const { runtime, exited } = await createSignaledLoopHarness(); - const restartSignal = captureSignal("SIGUSR2"); - - vi.useFakeTimers(); - restartSignal(); - await vi.advanceTimersByTimeAsync(1499); - expect(runtime.exit).not.toHaveBeenCalled(); - await vi.advanceTimersByTimeAsync(1); - - await expect(exited).resolves.toBe(0); - expect(runtime.exit).toHaveBeenCalledWith(0); - expectRestartHandoffCall({ - restartKind: "full-process", - reason: undefined, - supervisorMode: "launchd", - }); - }); - } finally { - vi.useRealTimers(); - delete process.env.OPENCLAW_LAUNCHD_LABEL; - if (originalPlatformDescriptor) { - Object.defineProperty(process, "platform", originalPlatformDescriptor); - } - } - }); - - it("falls back in-process when the launchd restart handoff fails to spawn", async () => { - vi.clearAllMocks(); - peekGatewayRestartReason.mockReturnValue(undefined); - try { - setPlatform("darwin"); - process.env.OPENCLAW_LAUNCHD_LABEL = "ai.openclaw.gateway"; - restartGatewayProcessWithFreshPid.mockReturnValueOnce({ - mode: "supervised", - handoffSpawned: Promise.resolve(false), - }); - - await withIsolatedSignals(async ({ captureSignal }) => { - const { start, runtime, exited } = await createSignaledLoopHarness(); - const restartSignal = captureSignal("SIGUSR2"); - const sigint = captureSignal("SIGINT"); - - vi.useFakeTimers(); - restartSignal(); - await vi.advanceTimersByTimeAsync(1500); - - expect(start).toHaveBeenCalledTimes(2); - expect(runtime.exit).not.toHaveBeenCalled(); - expect(acquireGatewayLock).toHaveBeenCalledTimes(2); - expect(gatewayLog.warn).toHaveBeenCalledWith( - "launchd restart handoff failed to spawn; falling back to in-process restart", - ); - - sigint(); - await expect(exited).resolves.toBe(0); - }); - } finally { - vi.useRealTimers(); - delete process.env.OPENCLAW_LAUNCHD_LABEL; - if (originalPlatformDescriptor) { - Object.defineProperty(process, "platform", originalPlatformDescriptor); - } - } - }); - - it("leaves the successor to launchd after a SIGTERM restart intent", async () => { - vi.clearAllMocks(); - consumeGatewayRestartIntentPayloadSync.mockReturnValueOnce({ reason: "gateway.restart" }); - setPlatform("darwin"); - process.env.OPENCLAW_LAUNCHD_LABEL = "ai.openclaw.gateway"; - restartGatewayProcessWithFreshPid.mockReturnValueOnce({ - mode: "supervised", - handoffSpawned: Promise.resolve(true), - }); - - await withIsolatedSignals(async ({ captureSignal }) => { - const { start, exited } = await createSignaledLoopHarness(); - captureSignal("SIGTERM")(); - await expect(exited).resolves.toBe(0); - expect(start).toHaveBeenCalledOnce(); - expect(restartGatewayProcessWithFreshPid).not.toHaveBeenCalled(); - expect(respawnGatewayProcessForUpdate).not.toHaveBeenCalled(); - }); - }); - it("records external ownership even when native supervisor markers are inherited", async () => { vi.clearAllMocks(); peekGatewayRestartReason.mockReturnValue(undefined); diff --git a/src/cli/gateway-cli/run-loop.ts b/src/cli/gateway-cli/run-loop.ts index 6966b7f272e3..65aef954bc3a 100644 --- a/src/cli/gateway-cli/run-loop.ts +++ b/src/cli/gateway-cli/run-loop.ts @@ -609,7 +609,7 @@ export async function runGatewayLoop(params: { return timer; }; const timer = arm(startupBudget.timeoutMs); - if (process.platform === "linux") { + if (process.platform === "linux" || process.platform === "darwin") { void resolveGatewayShutdownBudget(supervisorMode, gatewayLog, { previous: startupBudget, acceptedAtMs: pendingRequest.acceptedAtMs, @@ -761,7 +761,7 @@ export async function runGatewayLoop(params: { } const completion = (async () => { - if (process.platform === "linux") { + if (process.platform === "linux" || process.platform === "darwin") { if (budget.nativeStopBudget && !getManagedUpdateOwner()) { armForceExitTimer(budget.timeoutMs); } diff --git a/src/infra/gateway-shutdown-budget.ts b/src/infra/gateway-shutdown-budget.ts index c07f05d3aadd..ddfe93798e34 100644 --- a/src/infra/gateway-shutdown-budget.ts +++ b/src/infra/gateway-shutdown-budget.ts @@ -4,5 +4,9 @@ export { GATEWAY_SUPERVISOR_EXIT_MARGIN_MS, GATEWAY_SHUTDOWN_TIMEOUT_MS, GATEWAY_SERVICE_STOP_TIMEOUT_MS, + isRespawnedByLauncher, LAUNCH_AGENT_EXIT_TIMEOUT_SECONDS, + resolveLauncherStopTimeoutMs, + resolveShutdownReserveMs, + resolveSupervisorExitMarginMs, } from "../../gateway-shutdown-budget.mjs"; diff --git a/src/infra/launchd-stop-timeout.test.ts b/src/infra/launchd-stop-timeout.test.ts new file mode 100644 index 000000000000..ee56e9a0145a --- /dev/null +++ b/src/infra/launchd-stop-timeout.test.ts @@ -0,0 +1,376 @@ +import { afterEach, beforeEach, describe, expect, it, vi } from "vitest"; +import { isRespawnedByLauncher } from "./gateway-shutdown-budget.js"; +import { readLaunchdStopTimeout } from "./launchd-stop-timeout.js"; + +const { execLaunchctl } = vi.hoisted(() => ({ execLaunchctl: vi.fn() })); +vi.mock("../daemon/launchd-exec.js", async (importOriginal) => ({ + ...(await importOriginal()), + execLaunchctl, +})); + +// The authoritative list is module-local beside the deadline it authorises in +// gateway-shutdown-budget.mjs, so it is restated here and tied back to the +// implementation by the recognition cases at the end of this file. +const RESPAWN_MARKERS = [ + "OPENCLAW_NODE_UPDATE_RESPAWNED", + "OPENCLAW_COMPILE_CACHE_DISABLED_RESPAWNED", + "OPENCLAW_PACKAGED_COMPILE_CACHE_RESPAWNED", +] as const; +const LAUNCHD_ENV = { XPC_SERVICE_NAME: "ai.openclaw.gateway" }; +// The service layout the recovery launcher bounds by the LaunchAgent exit timeout: +// launchd names the job in XPC_SERVICE_NAME and the handoff carries the label, which +// is the pair the launcher branches on and a respawned child inherits unchanged. +const SERVICE_ENV = { ...LAUNCHD_ENV, OPENCLAW_LAUNCHD_LABEL: "ai.openclaw.gateway" }; +// Adds one of the markers a respawning launcher stamps on its child, which is what +// separates a Gateway that launcher started from an unrelated parent. +const RESPAWNED_SERVICE_ENV = { ...SERVICE_ENV, OPENCLAW_NODE_UPDATE_RESPAWNED: "1" }; +const result = (stdout: string) => ({ code: 0, stdout, stderr: "", termination: "exit" }); + +/** + * Shaped like real `launchctl print` output rather than a bare field list: the + * job's own `state` at one tab, then coalition blocks carrying their own + * `state = active` at two tabs, then `job state`. A live Gateway LaunchDaemon + * prints `state` three times in exactly this arrangement. + */ +const printed = (state: string, fields: string) => + result( + `system/ai.openclaw.gateway = {\n\tactive count = 1\n\ttype = LaunchDaemon\n\tstate = ${state}\n\n${fields}\tresource coalition = {\n\t\tID = 18110\n\t\tstate = active\n\t}\n\n\tjetsam coalition = {\n\t\tID = 18111\n\t\tstate = active\n\t}\n\n\tjob state = running\n}\n`, + ); +const stopping = (fields: string) => printed("SIGTERMed", fields); + +beforeEach(() => { + vi.stubGlobal("process", { + ...process, + platform: "darwin", + pid: 4242, + ppid: 4241, + getuid: () => 501, + // The reader only runs inside a serving Gateway, and reconstructing an + // undeclared launcher's timer replays the argv branch the launcher took. + argv: ["/usr/local/bin/node", "/usr/local/bin/openclaw", "gateway", "run"], + }); + execLaunchctl.mockReset(); +}); +afterEach(() => vi.unstubAllGlobals()); + +describe("launchd stop timeout reads the job launchd is stopping", () => { + it("uses the operator job's effective exit timeout, not the template constant", async () => { + execLaunchctl.mockResolvedValue( + stopping("\tminimum runtime = 10\n\texit timeout = 5\n\tpid = 4242\n"), + ); + await expect(readLaunchdStopTimeout(LAUNCHD_ENV)).resolves.toEqual({ + stop: { timeoutMs: 5_000, source: "launchd system/ai.openclaw.gateway exit timeout" }, + }); + expect(execLaunchctl).toHaveBeenCalledExactlyOnceWith( + ["print", "system/ai.openclaw.gateway"], + 2_000, + ); + }); + + // The shared key-value parser keeps the LAST occurrence of a repeated key, and + // `launchctl print` repeats `state` inside coalition blocks, so asking it for + // `state` answers `active` and never sees the job at all. This case fails if + // the reader ever goes back to that parser for the job's state. + it("reads the job's own state, not a nested coalition's", async () => { + execLaunchctl.mockResolvedValue(stopping("\texit timeout = 47\n\tpid = 4242\n")); + await expect(readLaunchdStopTimeout(LAUNCHD_ENV)).resolves.toEqual({ + stop: { timeoutMs: 47_000, source: "launchd system/ai.openclaw.gateway exit timeout" }, + }); + }); + + // Measured on macOS 27 with ExitTimeOut 47: a plain `kill -TERM` left the job + // printing `state = running` while the process handled the signal, and it was + // still alive 85 seconds later. launchd never started its clock, so its + // deadline bounds nothing and the caller keeps the budget it already had. + it("declines the deadline when launchd is not the one stopping the job", async () => { + execLaunchctl.mockResolvedValue(printed("running", "\texit timeout = 5\n\tpid = 4242\n")); + // No deadline and nothing to warn about: this is the ordinary shape of a stop + // that some other sender delivered. + await expect(readLaunchdStopTimeout(LAUNCHD_ENV)).resolves.toEqual({ stop: null }); + // Our job was found in the first domain, so there is nothing left to search. + expect(execLaunchctl).toHaveBeenCalledTimes(1); + }); + + it("preserves direct-child drain when a marked launcher has not started its timer", async () => { + execLaunchctl.mockResolvedValue(printed("running", "\texit timeout = 60\n\tpid = 4241\n")); + await expect(readLaunchdStopTimeout(RESPAWNED_SERVICE_ENV)).resolves.toEqual({ stop: null }); + }); + + it("keeps launchd's unlimited exit timeout distinct from a missing one", async () => { + execLaunchctl.mockResolvedValue(stopping("\texit timeout = 0\n\tpid = 4242\n")); + await expect(readLaunchdStopTimeout(LAUNCHD_ENV)).resolves.toEqual({ + stop: { + timeoutMs: Infinity, + source: "launchd system/ai.openclaw.gateway unlimited exit timeout", + }, + }); + }); + + it("caps an unlimited launchd stop at the launcher's finite timer", async () => { + execLaunchctl.mockResolvedValue(stopping("\texit timeout = 0\n\tpid = 4241\n")); + await expect(readLaunchdStopTimeout(RESPAWNED_SERVICE_ENV)).resolves.toEqual({ + stop: { + timeoutMs: 19_000, + source: + "launchd system/ai.openclaw.gateway unlimited exit timeout capped at the launcher's 19000ms stop timer", + }, + }); + }); + + // An empty value is deliberately not in this table: there is no state to fail to + // recognise, so it belongs with the missing-state warning below. + it.each(["waiting", "exited", "not running", "SIGTERM", "sigtermed"])( + "treats the unrecognised state %j as not stopping", + async (state) => { + execLaunchctl.mockResolvedValue(printed(state, "\texit timeout = 5\n\tpid = 4242\n")); + await expect(readLaunchdStopTimeout(LAUNCHD_ENV)).resolves.toEqual({ stop: null }); + }, + ); + + it("accepts any signal launchd reports having delivered", async () => { + execLaunchctl.mockResolvedValue(printed("SIGKILLed", "\texit timeout = 9\n\tpid = 4242\n")); + await expect(readLaunchdStopTimeout(LAUNCHD_ENV)).resolves.toEqual({ + stop: { timeoutMs: 9_000, source: "launchd system/ai.openclaw.gateway exit timeout" }, + }); + }); + + // An absent `state` is indistinguishable from a job launchd is not stopping, so the + // budget would silently revert to the platform-neutral policy on a macOS that printed + // this block differently. The deadline parsed from the same block is what makes the + // case reportable, and the warning is what keeps it from being silent. An empty value + // reaches the same place as a missing line, because neither yields a state to read. + it.each([ + [ + "no state line", + `system/ai.openclaw.gateway = {\n\tactive count = 1\n\ttype = LaunchDaemon\n\n\texit timeout = 47\n\tpid = 4242\n\tjob state = running\n}\n`, + ], + ["an empty state value", printed("", "\texit timeout = 47\n\tpid = 4242\n").stdout], + ])("warns when the job printed a deadline but %s", async (_label, stdout) => { + execLaunchctl.mockResolvedValue(result(stdout)); + const read = await readLaunchdStopTimeout(LAUNCHD_ENV); + expect(read.stop).toBeNull(); + expect(read.warning).toBe( + "launchd system/ai.openclaw.gateway printed an exit timeout but no job state, so it is treated as not stopping and the Gateway stop policy is kept. Check the running job with launchctl print.", + ); + }); + + // A job with neither field is an ordinary not-stopping answer and must stay quiet, + // otherwise every in-process restart of a launchd-owned Gateway would warn. + it("stays silent when the job is simply not stopping", async () => { + execLaunchctl.mockResolvedValue(printed("running", "\texit timeout = 47\n\tpid = 4242\n")); + await expect(readLaunchdStopTimeout(LAUNCHD_ENV)).resolves.toEqual({ stop: null }); + }); + + it("falls back to the gui domain when the job is not a LaunchDaemon", async () => { + execLaunchctl + .mockResolvedValueOnce({ + code: 113, + stdout: "", + stderr: "Could not find service", + termination: "exit", + }) + .mockResolvedValueOnce(stopping("\texit timeout = 20\n\tpid = 4242\n")); + await expect(readLaunchdStopTimeout(LAUNCHD_ENV)).resolves.toEqual({ + stop: { timeoutMs: 20_000, source: "launchd gui/501/ai.openclaw.gateway exit timeout" }, + }); + }); + + // A service account with no logged-in session has no gui domain: launchctl + // answers 125 "Domain does not support specified action" for gui/ while + // user/ prints normally. Verified on a headless macOS service account. + it("reaches the user domain when the account has no gui session", async () => { + execLaunchctl + .mockResolvedValueOnce({ + code: 113, + stdout: "", + stderr: "Could not find service", + termination: "exit", + }) + .mockResolvedValueOnce({ + code: 125, + stdout: "", + stderr: "Domain does not support specified action", + termination: "exit", + }) + .mockResolvedValueOnce(stopping("\texit timeout = 30\n\tpid = 4242\n")); + await expect(readLaunchdStopTimeout(LAUNCHD_ENV)).resolves.toEqual({ + stop: { timeoutMs: 30_000, source: "launchd user/501/ai.openclaw.gateway exit timeout" }, + }); + expect(execLaunchctl).toHaveBeenCalledTimes(3); + }); + + it("refuses a same-named job in every other domain and says why", async () => { + execLaunchctl.mockResolvedValue(stopping("\texit timeout = 300\n\tpid = 99\n")); + const read = await readLaunchdStopTimeout(LAUNCHD_ENV); + // Nothing was established, so no deadline is reported at all and the caller + // keeps the platform-neutral policy it already resolved. + expect(read.stop).toBeNull(); + for (const target of [ + "system/ai.openclaw.gateway", + "gui/501/ai.openclaw.gateway", + "user/501/ai.openclaw.gateway", + ]) { + expect(read.warning).toContain(`${target}: pid 99 is neither this process nor its launcher`); + } + }); + + // The installed service can keep a launcher parent while the serving Gateway + // runs as its child, so the job prints the launcher's pid. Requiring + // pid === process.pid there would reject the job that enforces the deadline. + it("accepts the job when it is this process's launcher parent", async () => { + execLaunchctl.mockResolvedValue(stopping("\texit timeout = 12\n\tpid = 4241\n")); + await expect(readLaunchdStopTimeout(LAUNCHD_ENV)).resolves.toEqual({ + stop: { timeoutMs: 12_000, source: "launchd system/ai.openclaw.gateway exit timeout" }, + }); + }); + + // The launcher declares nothing, so its reap timer is derived from the expression + // the launcher itself arms from. A longer operator deadline cannot be spent under + // it: the parent force-kills this process first. Deriving rather than being told is + // what makes this hold when an already-running older launcher started this Gateway, + // which is the only shape an upgrade can take. + // + // Every marker is covered because all three respawn call sites reach the same + // launcher through the same function and arm the same timer. Gating on the + // Node-recovery marker alone left the two compile-cache respawns uncapped, and the + // packaged one can wrap a foreground `gateway run` on an installed service. + it.each(RESPAWN_MARKERS)( + "caps the job deadline at the launcher's derived reap timer for %s", + async (marker) => { + execLaunchctl.mockResolvedValue(stopping("\texit timeout = 55\n\tpid = 4241\n")); + await expect(readLaunchdStopTimeout({ ...SERVICE_ENV, [marker]: "1" })).resolves.toEqual({ + stop: { + timeoutMs: 19_000, + source: + "launchd system/ai.openclaw.gateway exit timeout capped at the launcher's 19000ms stop timer", + }, + }); + }, + ); + + // Holding the parent slot is not evidence of a reap timer. An operator wrapper can + // keep the job's pid and start the Gateway itself while running none, and capping + // its deadline would cut a drain nothing was going to interrupt. + it("leaves a parent that did not respawn this process capping nothing", async () => { + execLaunchctl.mockResolvedValue(stopping("\texit timeout = 55\n\tpid = 4241\n")); + await expect(readLaunchdStopTimeout(SERVICE_ENV)).resolves.toEqual({ + stop: { timeoutMs: 55_000, source: "launchd system/ai.openclaw.gateway exit timeout" }, + }); + }); + + // Outside the service layout the launcher bounds its child by the platform-neutral + // stop policy, which outlasts any job deadline launchd will enforce, so nothing is + // cut even though the marker says a launcher is there. + it("derives the policy deadline when the respawning job is not the service", async () => { + execLaunchctl.mockResolvedValue(stopping("\texit timeout = 55\n\tpid = 4241\n")); + await expect( + readLaunchdStopTimeout({ ...LAUNCHD_ENV, OPENCLAW_NODE_UPDATE_RESPAWNED: "1" }), + ).resolves.toEqual({ + stop: { timeoutMs: 55_000, source: "launchd system/ai.openclaw.gateway exit timeout" }, + }); + }); + + // launchd reaps the whole job first, so a shorter job deadline still wins. + it("keeps a job deadline shorter than the launcher's reap timer", async () => { + execLaunchctl.mockResolvedValue(stopping("\texit timeout = 5\n\tpid = 4241\n")); + await expect(readLaunchdStopTimeout(RESPAWNED_SERVICE_ENV)).resolves.toEqual({ + stop: { timeoutMs: 5_000, source: "launchd system/ai.openclaw.gateway exit timeout" }, + }); + }); + + // The launcher timer belongs to a parent, so a Gateway that is the job itself is + // never shortened by one even while carrying an inherited respawn marker. + it("ignores the launcher timer when this process is the job", async () => { + execLaunchctl.mockResolvedValue(stopping("\texit timeout = 55\n\tpid = 4242\n")); + await expect(readLaunchdStopTimeout(RESPAWNED_SERVICE_ENV)).resolves.toEqual({ + stop: { timeoutMs: 55_000, source: "launchd system/ai.openclaw.gateway exit timeout" }, + }); + }); + + // launchd is stopping the job, so a deadline is running even though its value + // is unreadable. Guess short here: guessing long is what lets the drain die. + it.each(["\tpid = 4242\n", "\texit timeout = not-a-number\n\tpid = 4242\n"])( + "uses the conservative default when a running stop has no readable deadline", + async (fields) => { + execLaunchctl.mockResolvedValue(stopping(fields)); + const read = await readLaunchdStopTimeout(LAUNCHD_ENV); + // A deadline IS running here, so this one is a real native budget even + // though its value had to be defaulted. + expect(read.stop?.timeoutMs).toBe(20_000); + expect(read.stop?.source).toBe( + "launchd system/ai.openclaw.gateway exit timeout unavailable; default ExitTimeOut", + ); + expect(read.warning).toContain( + "launchd is stopping system/ai.openclaw.gateway but its exit timeout is missing or invalid", + ); + }, + ); + + // THE REGRESSION GUARD for the failed-inspection path. A probe that established + // nothing must report no deadline: handing back the Gateway's own stop policy + // here is what let the caller classify an unverified number as a native stop + // budget, capping a longer requested restart drain and arming a forced exit. + it("reports no deadline when the job cannot be inspected", async () => { + execLaunchctl.mockResolvedValue({ + code: 1, + stdout: "", + stderr: "permission denied", + termination: "exit", + }); + const read = await readLaunchdStopTimeout(LAUNCHD_ENV); + expect(read.stop).toBeNull(); + expect(read.warning).toContain( + "system/ai.openclaw.gateway: launchctl print exited 1: permission denied", + ); + expect(read.warning).toContain("keeping the Gateway stop policy"); + expect(read.warning).toContain("Check the running job with launchctl print."); + }); + + it("survives launchctl throwing rather than exiting nonzero", async () => { + execLaunchctl.mockRejectedValue(new Error("spawn ENOENT")); + const read = await readLaunchdStopTimeout(LAUNCHD_ENV); + expect(read.stop).toBeNull(); + expect(read.warning).toContain("launchctl print threw"); + }); + + it("honours an explicit label override", async () => { + execLaunchctl.mockResolvedValue(stopping("\texit timeout = 45\n\tpid = 4242\n")); + await expect( + readLaunchdStopTimeout({ ...LAUNCHD_ENV, OPENCLAW_LAUNCHD_LABEL: "com.example.gw" }), + ).resolves.toEqual({ + stop: { timeoutMs: 45_000, source: "launchd system/com.example.gw exit timeout" }, + }); + }); + + it("warns instead of throwing when the configured label is invalid", async () => { + const read = await readLaunchdStopTimeout({ + ...LAUNCHD_ENV, + OPENCLAW_LAUNCHD_LABEL: "bad label/../etc", + }); + expect(read.stop).toBeNull(); + expect(read.warning).toContain("label could not be resolved"); + expect(execLaunchctl).not.toHaveBeenCalled(); + }); + + it("stays out of the way when this process is not a launchd job", async () => { + await expect(readLaunchdStopTimeout({})).resolves.toEqual({ stop: null }); + expect(execLaunchctl).not.toHaveBeenCalled(); + }); +}); + +// Ties the names this file caps on back to the implementation's own list. Without +// this, dropping a marker from that list would leave the cap cases above still +// passing against the two that remained. +describe("launcher respawn markers", () => { + it.each(RESPAWN_MARKERS)("recognises %s as a respawning launcher", (marker) => { + expect(isRespawnedByLauncher({ [marker]: "1" })).toBe(true); + }); + + it("recognises nothing else", () => { + expect(isRespawnedByLauncher({})).toBe(false); + expect(isRespawnedByLauncher({ OPENCLAW_SUPERVISOR_MODE: "external" })).toBe(false); + // Only the set value counts, so an emptied or disabled marker caps nothing. + expect(isRespawnedByLauncher({ OPENCLAW_NODE_UPDATE_RESPAWNED: "0" })).toBe(false); + expect(isRespawnedByLauncher({ OPENCLAW_NODE_UPDATE_RESPAWNED: "" })).toBe(false); + }); +}); diff --git a/src/infra/launchd-stop-timeout.ts b/src/infra/launchd-stop-timeout.ts new file mode 100644 index 000000000000..22934b0b35b9 --- /dev/null +++ b/src/infra/launchd-stop-timeout.ts @@ -0,0 +1,187 @@ +import { parseStrictPositiveInteger } from "@openclaw/normalization-core/number-coercion"; +import { truncateUtf16Safe } from "@openclaw/normalization-core/utf16-slice"; +import { isForegroundGatewayRunArgv } from "../cli/gateway-run-argv.js"; +import { execLaunchctl, formatLaunchctlResultDetail } from "../daemon/launchd-exec.js"; +import { resolveLaunchAgentLabel } from "../daemon/launchd-label.js"; +import { parseKeyValueOutput } from "../daemon/runtime-parse.js"; +import { formatErrorMessage } from "./errors.js"; +import { + isRespawnedByLauncher, + LAUNCH_AGENT_EXIT_TIMEOUT_SECONDS, + resolveLauncherStopTimeoutMs, +} from "./gateway-shutdown-budget.js"; +import { detectRespawnSupervisor } from "./supervisor-markers.js"; + +type LaunchdStopTimeout = { timeoutMs: number; source: string }; + +// A warning without a deadline means inspection was inconclusive. Only a +// confirmed launchd stop or our launcher's independent reap timer yields a +// native budget; the caller owns the fallback policy. +export type LaunchdStopRead = { stop: LaunchdStopTimeout | null; warning?: string }; + +const LAUNCHCTL_PRINT_TIMEOUT_MS = 2_000; + +// Check all domains: a service account may lack a GUI session, and a matching +// label in a different domain must not supply this process's deadline. +function resolveLaunchdDomains(label: string): string[] { + const uid = typeof process.getuid === "function" ? process.getuid() : 501; + // A LaunchDaemon and a LaunchAgent can carry the same label in different + // domains, so each is a candidate and the pid decides which one is ours. + // A service account with no logged-in session has no gui domain at all, and + // its per-user jobs live in `user/`, so both user domains are checked. + return [`system/${label}`, `gui/${uid}/${label}`, `user/${uid}/${label}`]; +} + +// launchctl can report the Gateway or a launcher parent that holds the job PID. +function resolveJobRelation(pid: number | undefined): "self" | "launcher" | null { + if (pid === undefined) { + return null; + } + if (pid === process.pid) { + return "self"; + } + return pid === process.ppid ? "launcher" : null; +} + +// Nested coalition blocks also contain state: use the single-tab job field, +// not the last state selected by the generic key-value parser. +function readJobState(printed: string): string | undefined { + return /^\tstate = (?.+)$/mu.exec(printed)?.groups?.state?.trim(); +} + +// launchd reports SIGTERMed during bootout/kickstart, but keeps running on an +// externally delivered SIGTERM. Only the former starts ExitTimeOut's clock. +function isLaunchdStoppingJob(state: string | undefined): boolean { + return state !== undefined && /^SIG[A-Z0-9]+ed$/u.test(state); +} + +// A respawn marker, not parenthood alone, proves the launcher runs a reap timer. +// Derive that timer from the same shared expression: an already-running parent +// cannot be taught a new announced deadline by upgrading the child. +function resolveParentLauncherStopTimeoutMs(env: NodeJS.ProcessEnv): number | undefined { + // One of these is set on every child the launcher respawns, and all of them predate + // this deadline being derived, so a marker is present for a Gateway that launcher + // started and absent for any other parent. An operator wrapper that keeps the job's + // pid and starts the Gateway itself runs no such timer, and capping its deadline + // would cut a drain nothing was going to interrupt. + if (!isRespawnedByLauncher(env)) { + return undefined; + } + return resolveLauncherStopTimeoutMs({ + env, + platform: process.platform, + // The launcher branched on its own argv, and it respawns the child with the same + // user arguments, so testing ours reproduces the branch it took. + foreground: isForegroundGatewayRunArgv(process.argv), + }); +} + +// The job is stopping but its value is missing; use launchd's 20s default. +function defaultStopDeadline(target: string, reason: string): LaunchdStopRead { + const timeoutMs = LAUNCH_AGENT_EXIT_TIMEOUT_SECONDS * 1_000; + return { + stop: { timeoutMs, source: `launchd ${target} exit timeout unavailable; default ExitTimeOut` }, + warning: `launchd is stopping ${target} but ${reason}; using ${timeoutMs}ms default. Check the running job with launchctl print.`, + }; +} + +// Unknown ownership of the stop cannot justify shortening a potentially +// unconstrained drain. Warn, but leave the caller's policy authoritative. +function unresolved(failures: string[]): LaunchdStopRead { + return { + stop: null, + // Each failure already names the target it came from, and the label-resolution + // case has no label to name, so the prefix deliberately carries neither. + warning: `Unable to inspect the launchd job; ${failures + .map((failure) => truncateUtf16Safe(failure.replaceAll(/\s+/g, " "), 500)) + .join("; ")}; keeping the Gateway stop policy. Check the running job with launchctl print.`, + }; +} + +// Read the loaded job only on shutdown. Its ExitTimeOut applies only while +// launchd stops it; a respawn launcher can enforce an independent deadline. +export async function readLaunchdStopTimeout( + env: NodeJS.ProcessEnv = process.env, +): Promise { + if (detectRespawnSupervisor(env, "darwin") !== "launchd") { + return { stop: null }; + } + const failures: string[] = []; + let label: string; + try { + label = resolveLaunchAgentLabel(env); + } catch (error: unknown) { + return unresolved([`label could not be resolved: ${formatErrorMessage(error)}`]); + } + for (const target of resolveLaunchdDomains(label)) { + const failed = (reason: string) => failures.push(`${target}: ${reason}`); + const result = await execLaunchctl(["print", target], LAUNCHCTL_PRINT_TIMEOUT_MS).catch( + (error: unknown) => { + failed(`launchctl print threw: ${formatErrorMessage(error)}`); + return undefined; + }, + ); + if (!result) { + continue; + } + if (result.code !== 0) { + failed(`launchctl print exited ${result.code}: ${formatLaunchctlResultDetail(result)}`); + continue; + } + const printed = result.stdout || result.stderr || ""; + const entries = parseKeyValueOutput(printed, "="); + // Adopting a deadline from a same-named job in the other domain would be + // worse than the fallback, so the printed job must be ours. + const pid = parseStrictPositiveInteger(entries.pid ?? ""); + const relation = resolveJobRelation(pid); + if (!relation) { + failed(`pid ${pid ?? "missing"} is neither this process nor its launcher`); + continue; + } + // This is our job, so stop searching. Whether its deadline binds this stop is + // a separate question from whether the job was found. + const launcherMs = + relation === "launcher" ? resolveParentLauncherStopTimeoutMs(env) : undefined; + const state = readJobState(printed); + if (!isLaunchdStoppingJob(state)) { + // A direct signal to the Gateway starts neither launchd's clock nor its + // parent launcher's timer. Node cannot identify the sender, so capping on + // the parent marker would truncate that supported long drain. A signal + // forwarded by the launcher is indistinguishable here (existing limitation). + return !state && entries["exit timeout"] !== undefined + ? { + stop: null, + warning: `launchd ${target} printed an exit timeout but no job state, so it is treated as not stopping and the Gateway stop policy is kept. Check the running job with launchctl print.`, + } + : { stop: null }; + } + const rawSeconds = entries["exit timeout"]?.trim(); + if (rawSeconds === "0") { + return launcherMs !== undefined + ? { + stop: { + timeoutMs: launcherMs, + source: `launchd ${target} unlimited exit timeout capped at the launcher's ${launcherMs}ms stop timer`, + }, + } + : { stop: { timeoutMs: Infinity, source: `launchd ${target} unlimited exit timeout` } }; + } + const seconds = parseStrictPositiveInteger(rawSeconds ?? ""); + if (seconds === undefined) { + return defaultStopDeadline(target, "its exit timeout is missing or invalid"); + } + const jobMs = seconds * 1_000; + // A parent that reaps this process on its own timer binds before the job's + // ExitTimeOut, and spending the longer deadline would only get the drain + // force-killed. + return launcherMs !== undefined && launcherMs < jobMs + ? { + stop: { + timeoutMs: launcherMs, + source: `launchd ${target} exit timeout capped at the launcher's ${launcherMs}ms stop timer`, + }, + } + : { stop: { timeoutMs: jobMs, source: `launchd ${target} exit timeout` } }; + } + return unresolved(failures); +} diff --git a/src/infra/node-runtime-recovery.shutdown.test.ts b/src/infra/node-runtime-recovery.shutdown.test.ts index 7bb2b8b6cf68..45c32f2242c0 100644 --- a/src/infra/node-runtime-recovery.shutdown.test.ts +++ b/src/infra/node-runtime-recovery.shutdown.test.ts @@ -18,6 +18,9 @@ beforeEach(() => { vi.useFakeTimers(); child = new ChildProcess(); kill = vi.spyOn(child, "kill").mockReturnValue(true); + // `spawn` is hoisted once for the file, so its call log survives across cases + // and `toHaveBeenCalledExactlyOnceWith` would only ever hold for the first one. + spawn.mockClear(); spawn.mockReturnValue(child); exit = vi.spyOn(process, "exit").mockImplementation(vi.fn()); vi.spyOn(process, "kill").mockReturnValue(true); @@ -50,6 +53,21 @@ it.each([ XPC_SERVICE_NAME: "ai.openclaw.fixture", }); detach = () => child.emit("exit", 0, null); + // The launcher tells the child nothing about the timer it armed: the serving + // Gateway derives the same deadline from the same shared expression, which is what + // lets a Gateway started by an already-running older launcher bound itself + // correctly. So the env must reach the child unchanged, and the escalation + // asserted below is what that derivation has to land on. + expect(spawn).toHaveBeenCalledExactlyOnceWith( + "node", + ["child.mjs"], + expect.objectContaining({ + env: { + OPENCLAW_LAUNCHD_LABEL: "ai.openclaw.fixture", + XPC_SERVICE_NAME: "ai.openclaw.fixture", + }, + }), + ); const signal = process.listeners("SIGTERM").find((listener) => !previous.has(listener)); expect(signal).toBeDefined(); signal!("SIGTERM"); diff --git a/test/scripts/package-acceptance-workflow.test.ts b/test/scripts/package-acceptance-workflow.test.ts index 1da9bb9aa72a..6563801db0d5 100644 --- a/test/scripts/package-acceptance-workflow.test.ts +++ b/test/scripts/package-acceptance-workflow.test.ts @@ -1291,6 +1291,7 @@ describe("frozen admission workflow barriers", () => { for (const [index, child] of children.entries()) { const toolingPaths = child.sources.tooling.map(({ path }) => path); const selectedPaths = child.sources.selected.map(({ path }) => path); + expect(toolingPaths).toContain("scripts/lib/record-shared.mjs"); if (index === 1) { expect(toolingPaths).toContain(extra); expect(selectedPaths).toContain("extensions/codex/package.json"); diff --git a/test/scripts/preflight-frozen-target-contracts.test.ts b/test/scripts/preflight-frozen-target-contracts.test.ts index 9044ebbd8bc4..9151646162d6 100644 --- a/test/scripts/preflight-frozen-target-contracts.test.ts +++ b/test/scripts/preflight-frozen-target-contracts.test.ts @@ -92,7 +92,6 @@ function fixture( recursive: true, }); for (const file of [ - "record-shared.mjs", "update-compat-contract.mjs", "openclaw-e2e-instance.sh", "docker-e2e-watchdog.mjs", @@ -554,6 +553,7 @@ describe("frozen admission bootstrap repairs", () => { it.each([ reader, "scripts/lib/docker-e2e-scenarios.mts", + "scripts/lib/record-shared.mjs", shell, "scripts/lib/trusted-native-typescript.mjs", "scripts/lib/native-typescript.mts",