diff --git a/docs/automation/cron-jobs/how-it-works.md b/docs/automation/cron-jobs/how-it-works.md index df6a54321057..518bd55eeed2 100644 --- a/docs/automation/cron-jobs/how-it-works.md +++ b/docs/automation/cron-jobs/how-it-works.md @@ -19,6 +19,7 @@ How the Gateway scheduler runs a job, what it keeps between runs, and how a repe - One-shot jobs (`--at`) auto-delete after successful completion: delivery is confirmed, not requested, intentionally suppressed, or explicitly best-effort. Failed or unknown required delivery retains the job disabled for inspection without replaying the payload. Pass `--keep-after-run` to keep successful jobs too. - Per-run wall-clock budget: `--timeout-seconds` when set. Otherwise, isolated/detached agent-turn jobs are bounded by the scheduler's own 60-minute watchdog before the underlying agent-turn timeout (`agents.defaults.timeoutSeconds`, default 48 hours) would ever apply; command jobs default to 10 minutes, and script payloads default to 5 minutes. - On Gateway startup, overdue agent-turn jobs and jobs that await a heartbeat are rescheduled instead of replayed immediately, keeping model/tool execution out of scheduler startup. This includes heartbeat monitors, migrated heartbeat tasks, and main-session system events with immediate wake. Startup catch-up delays survive label or payload-content reconciliation and another restart; changing the schedule starts a new scheduling decision. +- When a running Gateway wakes after a deadline, the scheduler coalesces missed timer ticks. Cron rechecks eligible jobs using its stored deadlines and run receipts. The shared [Gateway scheduler](/concepts/architecture#timed-work-and-shutdown) owns timer arming and shutdown joining; cron owns execution, catch-up policy, and durable results. - If you drive `openclaw agent` from system cron or another external scheduler, wrap it with a hard-kill escalation even though the CLI already handles `SIGTERM`/`SIGINT`. Gateway-backed runs ask the Gateway to abort accepted runs; `--local` runs get the same abort signal. For GNU `timeout`, prefer `timeout -k 60 600 openclaw agent ...` over plain `timeout 600 ...` — the `-k` value is the backstop if the process cannot drain in time. For systemd units, use a `SIGTERM` stop signal with a grace window (`TimeoutStopSec`) before the final kill. Reusing a `--run-id` while the original Gateway run is still active reports the duplicate as in-flight instead of starting a second run. diff --git a/docs/concepts/architecture.md b/docs/concepts/architecture.md index 9360b34528fc..f43a7c56b7f0 100644 --- a/docs/concepts/architecture.md +++ b/docs/concepts/architecture.md @@ -141,6 +141,27 @@ Details: [Gateway protocol](/gateway/protocol), [Pairing](/channels/pairing), - Health: `health` over WS (also included in `hello-ok`). - Supervision: launchd/systemd for auto-restart. +### Timed work and shutdown + +The Gateway kernel owns one `GatewayScheduler` for registered maintenance and cron +wakeups. Owners receive that instance and register jobs; one host timer drives the +next wake. For durable work, stores retain deadlines and their owners reconstruct +schedules at startup rather than persisting a second scheduler state. + +After sleep, a late wake dispatches each runnable, due registration once for that +wake. A periodic registration waits for its callback and tracked work to finish +before starting its next interval; missed ticks are coalesced. Wall time catches +sleep, while elapsed time keeps relative delays and cadences moving through a +backward clock correction. Rescheduling replaces a waiting job by default; +`mode: "earliest"` preserves earlier wall and elapsed deadlines for the same job ID +so a stale read cannot postpone an already promised wake. + +`beginClose()` closes scheduling admission, cancels pending wakes, and signals +shutdown. `stop()` joins callbacks already running and work tracked by their async +scope; the Gateway lifecycle owns the outer shutdown budget and resource teardown. +Request deadlines, stream-local timers, and child-process cleanup stay with their +operation owners. SQLite WAL checkpoint timers stay with the storage owner. + ## Invariants - Exactly one Gateway controls a single Baileys session per host.