mirror of
https://github.com/razzant/ouroboros.git
synced 2026-10-02 19:58:46 +00:00
fix: keep an ordinary close inside the launcher budget (#1142)
The launcher's graceful stop SIGTERMed the whole server process group, so the multiprocessing Manager and pooled workers died before the server's own lifespan teardown began; the supervisor loop met BrokenPipe on EVENT_Q three times and raised the false "Supervisor loop died" owner alarm, while uvicorn's unbounded graceful drain parked the terminal-custody write past the launcher's 10 s wait, ending three window closes in SIGKILL. Server half (self-sufficient against an old immutable launcher): - server._SignalStopServer.handle_exit sets _supervisor_stop at the signal, so the loop leaves its tick and a torn-Manager error is never a crash; - uvicorn.Config carries timeout_graceful_shutdown=SERVER_GRACEFUL_SHUTDOWN_TIMEOUT_SEC (runtime_limits, paired with LAUNCHER_STOP_GRACE_SEC as one budget). Launcher half (lands with the next release): the POSIX graceful phase signals only the server PID; the group SIGKILL fallback and post-exit sweep stay. Tests: unit coverage of the handler, the constant wiring and stop_agent, plus a serial real-process test that reproduces the old group SIGTERM with an in-flight request open and asserts the teardown lands and no supervisor_failure is written (it fails on the unfixed server). Docs 01/05/09 updated; version-neutral. Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
This commit is contained in:
parent
87fd00f400
commit
78780e0b7d
13 changed files with 254 additions and 19 deletions
|
|
@ -546,7 +546,7 @@ System self-modification, external workspace, and genesis remain distinct task c
|
|||
|
||||
Two continuity roles: `launcher.py` owns the PID lock, bundle bootstrap, the server process, presentation, the restart signal, and cleanup (desktop launchers run outside the managed repo; Android source-host mode retains a separate immutable seed); `server.py` is the self-editable inner runtime. Native packages ship an opt-in systemd user unit as an alternate ingress, not a third role — deliberately without a restart policy, because the launcher owns managed restart, the crash fuse, and panic-to-complete-stop (`KillMode=control-group`).
|
||||
|
||||
Spawn custody: POSIX children start in a new session/process group; Windows creates the server suspended, assigns a kill-on-close Job, then resumes — failure to establish Job custody refuses to run. The launcher Job permits explicit breakaway; only the shared daemon requests it, while ordinary generation children remain covered by Job close. Lazy worker and direct-server starts use the same platform helper. An old packaged or external Job that cannot confirm breakaway retains the earlier working spawn with a disclosed survive-close limitation; desktop managed source updates cannot replace their immutable launcher. The launcher records `data/state/server_process.json` (PID, pgid, server/repo paths, requested and actual ports, argv, creation time) and re-proves identity before cleanup. Forced tree/group cleanup excludes the shared daemon subtree; existing listener sweeps retain their separate scope.
|
||||
Spawn custody: POSIX children start in a new session/process group; Windows creates the server suspended, assigns a kill-on-close Job, then resumes — failure to establish Job custody refuses to run. The launcher Job permits explicit breakaway; only the shared daemon requests it, while ordinary generation children remain covered by Job close. Lazy worker and direct-server starts use the same platform helper. An old packaged or external Job that cannot confirm breakaway retains the earlier working spawn with a disclosed survive-close limitation; desktop managed source updates cannot replace their immutable launcher. The launcher records `data/state/server_process.json` (PID, pgid, server/repo paths, requested and actual ports, argv, creation time) and re-proves identity before cleanup. Forced tree/group cleanup excludes the shared daemon subtree; a graceful stop signals the server PID alone (§9); sweeps keep their scope.
|
||||
|
||||
Same-install reaper (`launcher_server_reaper.py`): holding the PID lock licenses the reap, which runs at `main()` preflight and at the top of every launcher generation. A PID is proven only on three live facts — the exact `<REPO_DIR>/server.py` argv token, `OUROBOROS_DATA_DIR`, and `OUROBOROS_MANAGED_BY_LAUNCHER=1` — revalidated immediately before the signal, with descendants captured before the root signal and the whole pass bounded to three rounds. Kills require the byte-exact /proc environment: `ps -E` output never authorizes a kill, because argv is indistinguishable from an env assignment there, so non-/proc hosts stay report-only. Server selection does not require a custody row — missing server records are the defect being repaired. Caller-supplied retained daemon roots filter descendants before final root revalidation. POSIX-only; Windows generation children die with the launcher Job while the shared daemon breaks away; never on panic or window-close. The startup stray check is report-only and annotates `same_install`/`foreign`.
|
||||
|
||||
|
|
|
|||
|
|
@ -2,7 +2,7 @@
|
|||
|
||||
This chapter owns the single scheduler for pooled work: what a healthy tick does, what the queue holds and what its durable snapshot may restore, how a task is addressed and named at admission, how owner waits lend capacity, and the intent-then-custody skeleton every cancellation follows. It exists because these invariants decide whether a stopped task ends honestly or leaves a ghost, and none of them can be reconstructed from any single module's code.
|
||||
|
||||
`server.py::_run_supervisor()` is the single scheduler for pooled tasks. A healthy tick publishes liveness, rotates runtime logs, checks worker health, drains worker and direct-chat events (a wake-up's frames are a direct turn's), accepts owner bridge input, enforces deadlines and schedules, runs throttled reconciliation and evolution admission, assigns eligible work, persists `state/queue_snapshot.json`, and last ticks the consciousness alarm clock, which owns no thread — this pass is the only thing that can start a wake-up. Bridge intake precedes timeout, maintenance, evolution and assignment work so a slow control-plane step cannot hide a new owner message. Three consecutive loop failures clear supervisor readiness, stop its watchdog generation and notify the owner, rather than leave a healthy-looking server that no longer assigns work; a failure while a shutdown or restart is in progress (the teardown sets a process-local stop event first) is not a crash and never holds the shutdown.
|
||||
`server.py::_run_supervisor()` is the single scheduler for pooled tasks. A healthy tick publishes liveness, rotates runtime logs, checks worker health, drains worker and direct-chat events (a wake-up's frames are a direct turn's), accepts owner bridge input, enforces deadlines and schedules, runs throttled reconciliation and evolution admission, assigns eligible work, persists `state/queue_snapshot.json`, and last ticks the consciousness alarm clock, which owns no thread — this pass is the only thing that can start a wake-up. Bridge intake precedes timeout, maintenance, evolution and assignment work so a slow control-plane step cannot hide a new owner message. Three consecutive loop failures clear supervisor readiness, stop its watchdog generation and notify the owner, rather than leave a healthy-looking server that no longer assigns work; a failure while a shutdown or restart is in progress is not a crash and never holds the shutdown — the process-local stop event is set at the SIGTERM/SIGINT instant by `server._SignalStopServer.handle_exit` and again first thing in the lifespan teardown, because a launcher that signals the whole group kills the event Manager before the teardown can run (issue #1142; §9).
|
||||
|
||||
The legacy `state.json` budget projection reads the validated usage ledger before taking `STATE_LOCK`, so a multi-megabyte replay cannot hold the short reader lock while waiting on the monetary lock. The same read returns a ledger provenance marker `(compaction_epoch, seq)`: the epoch comes from the lock-free leading baseline header and `seq` is the live file's validated high-water sequence. A compaction increments the epoch while renumbering live rows, so the pair remains ordered even when the file gets shorter; neither timestamps nor totals prove order. Inside `STATE_LOCK`, any marker strictly lower than the saved one — a lower epoch, or a lower sequence within the same epoch — is proven stale and leaves the compatibility projection untouched. An equal or higher marker is written. Restore to an older ledger or a malformed saved marker freezes this projection until explicit repair; automatic recovery is absent. Quarantine-file presence likewise keeps `integrity_degraded` and freezes this compatibility projection indefinitely. This read is display (invariant 28); admission is exact. An unavailable or malformed marker leaves the prior projection untouched; this fail-safe also applies when lock acquisition times out and the writer deliberately proceeds without the lock. On that no-lock path the marker still rejects a demonstrably stale snapshot, but the compare-and-save sequence remains non-atomic for concurrent no-lock writers, as it was before this protection.
|
||||
|
||||
|
|
|
|||
|
|
@ -2,7 +2,7 @@
|
|||
|
||||
This chapter owns the ways this installation stops: ordinary window close, which preserves the shared daemon; the owner's manual Restart, which is a clean stop of everything this server generation owns; and Panic, which is a complete explicit stop. It exists because each has a different custody contract — what is signalled, what is only disclosed as unconfirmed, and what the next generation is expected to reconcile — and confusing them is how a live paid run gets killed or a dead one gets reported as stopped.
|
||||
|
||||
Ordinary window close preserves this installation's shared Claudexor daemon and work submitted by other clients, on graceful and forced launcher cleanup paths alike. It stops the server generation's workers and services, waits for its process group or Job, performs owned cleanup and releases the PID lock; the Windows daemon lives outside the launcher Job, and an old immutable launcher needs a later package update for this guarantee. A same-pin planned restart retains the daemon; a planned restart after a changed engine pin stops it in lifespan teardown (`server_restart._stop_owned_daemon_for_new_pin`: the landed `load_runtime_pin` compared with the serving engine through attach-only `read_owned_gateway`, then the explicit stop below), and an unpublished or unreadable pin or an unreachable, foreign or stopped daemon leaves the handoff unchanged. An unconfirmed stop retains custody and diagnostics while the restart proceeds; the next generation reconciles old runs without inventing spend. In lifespan teardown the ONE irreversible durable write goes first — the stop flag, the bounded supervisor join, then `kill_workers`, which terminalizes every interrupted task — ahead of the best-effort extension-reconcile and host-service waits and the kill-all sweeps, because that order keeps the write reachable inside the launcher's force-exit budget (an external SIGKILL can still cut it short, and the later bridge and event-bus teardown can lose a late frame). A stop that never reaches that write leaves it to the next generation's snapshot restore; queue init hands the direct-chat roster `state/direct_roots.json` to it in the step that clears it, so listed roots are fenced like any surviving RUNNING row and end Cancelled, not as infra-failed orphans. Ordinary server shutdown closes its own Host Service listener; reserved-port sweeps remain separate launcher/recovery/Panic behaviour: they are not PID-identity proof and can terminate an unrelated listener on a reserved port.
|
||||
Ordinary window close preserves this installation's shared Claudexor daemon and work submitted by other clients, on graceful and forced launcher cleanup paths alike. It stops the server generation's workers and services, waits for its process group or Job, performs owned cleanup and releases the PID lock; the Windows daemon lives outside the launcher Job, and an old immutable launcher needs a later package update for this guarantee. The graceful phase of `launcher.stop_agent` signals ONLY the server PID (`proc.terminate()`, on every platform) and waits `LAUNCHER_STOP_GRACE_SEC`; the group SIGKILL fallback after that wait and the post-exit group sweep are unchanged. The server is the custody owner of its multiprocessing Manager and pooled workers, and a group-wide SIGTERM at t=0 killed them before the server's own teardown began (issue #1142: the supervisor loop then met `BrokenPipe` on `EVENT_Q` three times and raised the false `supervisor_failure` owner alarm). The server half is self-sufficient against an old launcher that still SIGTERMs the group: `server._SignalStopServer.handle_exit` sets `_supervisor_stop` at the signal, so the loop leaves its tick and a torn-Manager error is never counted as a crash, and `uvicorn.Config(timeout_graceful_shutdown=SERVER_GRACEFUL_SHUTDOWN_TIMEOUT_SEC)` bounds the open HTTP/WS drain that otherwise parked the lifespan teardown past the launcher budget (the pair is one budget in `runtime_limits.py`; `tests/test_shutdown_signal_1142.py` drives a real server through the old group SIGTERM with an in-flight request open). A same-pin planned restart retains the daemon; a planned restart after a changed engine pin stops it in lifespan teardown (`server_restart._stop_owned_daemon_for_new_pin`: the landed `load_runtime_pin` compared with the serving engine through attach-only `read_owned_gateway`, then the explicit stop below), and an unpublished or unreadable pin or an unreachable, foreign or stopped daemon leaves the handoff unchanged. An unconfirmed stop retains custody and diagnostics while the restart proceeds; the next generation reconciles old runs without inventing spend. In lifespan teardown the ONE irreversible durable write goes first — the stop flag, the bounded supervisor join, then `kill_workers`, which terminalizes every interrupted task — ahead of the best-effort extension-reconcile and host-service waits and the kill-all sweeps, because that order keeps the write reachable inside the launcher's force-exit budget (an external SIGKILL can still cut it short, and the later bridge and event-bus teardown can lose a late frame). A stop that never reaches that write leaves it to the next generation's snapshot restore; queue init hands the direct-chat roster `state/direct_roots.json` to it in the step that clears it, so listed roots are fenced like any surviving RUNNING row and end Cancelled, not as infra-failed orphans. Ordinary server shutdown closes its own Host Service listener; reserved-port sweeps remain separate launcher/recovery/Panic behaviour: they are not PID-identity proof and can terminate an unrelated listener on a reserved port.
|
||||
|
||||
The owner's manual Restart (`/restart`, the chat Restart button) is a clean stop of everything the current server generation owns, then the re-exec. The bridge delegates to `server_restart._perform_owner_restart` (re-exported by `server`) with its transport notice callback. Before a supervisor publishes its bridge, HTTP admission binds these same Restart and Panic operations directly; deferred execution hands off only to a live supervisor with a published bridge. A thread still initializing, or a stale bridge left by a dead thread, cannot consume a command, so the direct owner executes it. That common operation is checkout-first: `_safe_restart_serialized` (update lock, strict managed-update transaction gate, then `safe_restart`'s checkout, dependency sync, import test and stable fallback) runs before anything is stopped, so every refusal — an assisted-update merge mid-resolution, a failed checkout, an unwritable no-resume flag — leaves the server intact and is answered with one "Restart cancelled" line. Past that point the restart always follows: the durable no-resume intent (`state/owner_restart_no_resume.flag` plus the `panic_stop.flag` compatibility pair, consumed by `auto_resume_after_restart`), then `server_restart._stop_owned_work` — one immediate cancel intent per RUNNING task, live direct activity and in-flight post-task synthesis (`cancel_intents.request_cancel`, source `owner_restart`), `kill_workers`, delegated-run cancellation through the owner-gone seam `reconcile_orphaned_runs(running_task_ids=set())` over `read_owned_gateway`, and the attested owned-daemon `stop_outcome()` exactly as Panic makes it. Everything between the cancel intents and the daemon stop is attach-only: an `ensure_owned_gateway` there would start a dead daemon. Nothing is a veto or a deferral: an unconfirmed or raising worker shutdown leaves a critical log, an unconfirmed or raising daemon stop the `process_stop_unconfirmed` supervisor row; custody stays retained and the restart proceeds. The next generation does its ordinary startup work: owner restart is a no-resume cause (`delegate_recovery.NO_RESUME_CAUSES`), so nothing is adopted; the startup custody sweep reconciles every open delegated run as owner-gone (a run the stopped daemon took down answers absent → `close_absent_run`, no invented spend; a journal-recovered terminal settles; a pending invocation replays under its own key), and an in-flight owned model operation stays a `dispatched`/`unresolved` physical-attempt row that already counts as spend upper bound. After an unconfirmed stop the new generation attaches to the still-live daemon rather than spawning a second one, and the lifespan's one background `warm_owned_daemon()` (provisioned homes only) makes the first delegation after a Restart find the daemon serving. The explicit stop is installation-wide — manual Restart ends runs served by the owned daemon, another client's included — which differs intentionally from ordinary window close; planned self-restart, managed-update handoff and Panic keep their separate contracts.
|
||||
|
||||
|
|
|
|||
|
|
@ -2,7 +2,7 @@
|
|||
|
||||
Machine extraction of the `docs/ARCHITECTURE.md` "Data layout (`~/Ouroboros/`)" tree — the durable-file orientation carrier (this tree's counterpart of the reference PERSISTENCE_OWNERS derivation checklist) — regenerated by `python scripts/regenerate_inventories.py`. Do not edit. Every entry is probed against reality: repo entries must exist as tracked paths; data-plane entries must appear as a literal in the runtime sources that construct them. A durable file renamed or removed in code while its tree row survives = red (`tests/test_generated_inventories.py`).
|
||||
|
||||
Source: `docs/architecture/01-high-level-architecture.md`, physical LF lines 597-686; UTF-8 SHA-256 `87781e6db841d1d036d6c6f09911e7ce88a393cc8923714e49230ae7f9cad4fc`.
|
||||
Source: `docs/architecture/01-high-level-architecture.md`, physical LF lines 597-686; UTF-8 SHA-256 `bc3a49f91db6548b978f8ab451b3afb05b6fc7f101deafcc08fe7162ae9203e5`.
|
||||
|
||||
- entries: **79** (code-ref: 72, repo-dir: 6, repo-path: 1)
|
||||
|
||||
|
|
|
|||
|
|
@ -2,14 +2,14 @@
|
|||
|
||||
AST-derived inventory of compatibility facades, regenerated by `python scripts/regenerate_inventories.py`. Do not edit. A facade row is any runtime module whose top-level `from <population module> import ...` statements carry the `noqa: F401` re-export marker — the codebase's declared "this binding exists for its binding, not for this module's own use" convention (reference FACADE_CONSUMERS method). Leaf domains come from `ouroboros/domains.toml`; a leaf outside the facade's domain is marked ✗ (that edge also appears in the manifest's pinned direction matrix). `tests/test_generated_inventories.py` pins byte-identity, so any re-export surface change must regenerate this file.
|
||||
|
||||
- facade modules: **61**; marked re-export bindings: **2321**; cross-domain facade→leaf pairs: **132**
|
||||
- facade modules: **61**; marked re-export bindings: **2323**; cross-domain facade→leaf pairs: **132**
|
||||
|
||||
| facade | domain | bindings | leaves |
|
||||
|---|---|---:|---|
|
||||
| `launcher.py` | D18 | 2 | `ouroboros/launcher_windows_runtime.py` (2) |
|
||||
| `ouroboros/agent.py` | D01 | 32 | `ouroboros/agent_dispatch.py` (15)<br>`ouroboros/agent_startup_checks.py` (4)<br>`ouroboros/config.py` (2 ✗D12)<br>`ouroboros/subagent_dispatch_notes.py` (4 ✗D07)<br>`ouroboros/subagents.py` (7 ✗D07) |
|
||||
| `ouroboros/agent_task_pipeline.py` | D01 | 28 | `ouroboros/dialogue_provenance.py` (2 ✗D15)<br>`ouroboros/post_task_synthesis.py` (11)<br>`ouroboros/synthesis_cost_text.py` (5)<br>`ouroboros/task_finalization.py` (10) |
|
||||
| `ouroboros/config.py` | D12 | 133 | `ouroboros/model_slots.py` (17)<br>`ouroboros/provider_models.py` (6 ✗D02)<br>`ouroboros/review_model_routes.py` (10)<br>`ouroboros/runtime_limits.py` (62)<br>`ouroboros/settings_defaults.py` (19)<br>`ouroboros/settings_integrity.py` (4)<br>`ouroboros/settings_scales.py` (13)<br>`ouroboros/update_channels.py` (2) |
|
||||
| `ouroboros/config.py` | D12 | 135 | `ouroboros/model_slots.py` (17)<br>`ouroboros/provider_models.py` (6 ✗D02)<br>`ouroboros/review_model_routes.py` (10)<br>`ouroboros/runtime_limits.py` (64)<br>`ouroboros/settings_defaults.py` (19)<br>`ouroboros/settings_integrity.py` (4)<br>`ouroboros/settings_scales.py` (13)<br>`ouroboros/update_channels.py` (2) |
|
||||
| `ouroboros/context.py` | D03 | 4 | `ouroboros/context_runtime_facts.py` (4) |
|
||||
| `ouroboros/delegate_custody.py` | D07 | 10 | `ouroboros/delegate_custody_reconcile.py` (9)<br>`ouroboros/delegate_evidence.py` (1) |
|
||||
| `ouroboros/extension_loader.py` | D14 | 101 | `ouroboros/contracts/plugin_api.py` (7 ✗D19)<br>`ouroboros/extension_child_catalog.py` (8)<br>`ouroboros/extension_companion.py` (3)<br>`ouroboros/extension_import_staging.py` (6)<br>`ouroboros/extension_isolated_deps.py` (4)<br>`ouroboros/extension_liveness.py` (8)<br>`ouroboros/extension_plugin_api.py` (6)<br>`ouroboros/extension_registry_state.py` (20)<br>`ouroboros/extension_surface_names.py` (12)<br>`ouroboros/extension_ui_validation.py` (5)<br>`ouroboros/gateway/host_service.py` (1 ✗D11)<br>`ouroboros/provider_models.py` (1 ✗D02)<br>`ouroboros/skill_loader.py` (15)<br>`ouroboros/skill_token.py` (1)<br>`ouroboros/tools/skill_exec.py` (1)<br>`ouroboros/utils.py` (3 ✗D18) |
|
||||
|
|
|
|||
11
launcher.py
11
launcher.py
|
|
@ -32,6 +32,7 @@ os.environ.setdefault("PYTHONDONTWRITEBYTECODE", "1")
|
|||
from ouroboros.config import (
|
||||
AGENT_SERVER_PORT,
|
||||
DATA_DIR,
|
||||
LAUNCHER_STOP_GRACE_SEC,
|
||||
PANIC_EXIT_CODE,
|
||||
PORT_FILE,
|
||||
REPO_DIR,
|
||||
|
|
@ -95,7 +96,7 @@ from ouroboros.platform_layer import (
|
|||
subprocess_new_group_kwargs,
|
||||
terminate_job,
|
||||
terminate_process_group_id,
|
||||
terminate_process_tree, request_native_attention,
|
||||
request_native_attention,
|
||||
)
|
||||
from ouroboros.utils import atomic_write_json, utc_now_iso
|
||||
|
||||
|
|
@ -532,11 +533,9 @@ def stop_agent() -> None:
|
|||
|
||||
log.info("Stopping agent (pid=%s)...", proc.pid)
|
||||
try:
|
||||
if IS_WINDOWS:
|
||||
proc.terminate()
|
||||
else:
|
||||
terminate_process_tree(proc)
|
||||
proc.wait(timeout=10)
|
||||
# Graceful phase signals only the server: it owns its Manager and workers (#1142).
|
||||
proc.terminate()
|
||||
proc.wait(timeout=LAUNCHER_STOP_GRACE_SEC)
|
||||
except subprocess.TimeoutExpired:
|
||||
if IS_WINDOWS and job is not None:
|
||||
terminate_job(job)
|
||||
|
|
|
|||
|
|
@ -98,7 +98,7 @@ from ouroboros.runtime_limits import (
|
|||
WORKER_READY_WINDOW_SEC, # noqa: F401
|
||||
WORKER_READY_MAX_ATTEMPTS, WORKER_READY_CEILING_SEC, # noqa: F401
|
||||
EXTENSION_STREAM_CHUNK_BYTES, # noqa: F401
|
||||
EXTENSION_CHILD_CLEANUP_GRACE_SEC, # noqa: F401
|
||||
EXTENSION_CHILD_CLEANUP_GRACE_SEC, LAUNCHER_STOP_GRACE_SEC, SERVER_GRACEFUL_SHUTDOWN_TIMEOUT_SEC, # noqa: F401
|
||||
NESTED_SETTLEMENT_MARGIN_SEC, # noqa: F401
|
||||
NETWORK_WAIT_NOTE_INTERVAL_SEC, # noqa: F401
|
||||
NETWORK_WAIT_BACKOFF_START_SEC, # noqa: F401
|
||||
|
|
|
|||
|
|
@ -32,6 +32,12 @@ EXTERNAL_PLATFORM_UPDATE_TIMEOUT_SEC = 3600.0
|
|||
EXTENSION_STREAM_CHUNK_BYTES = 64 * 1024
|
||||
# Exit/pipe-drain grace after a response ends; never a response lifetime timer.
|
||||
EXTENSION_CHILD_CLEANUP_GRACE_SEC = 2
|
||||
# Ordinary close (#1142): the launcher waits this long for the server to exit before the group SIGKILL
|
||||
# fallback; uvicorn's graceful drain of open HTTP/WS tasks is bounded to the second value so the lifespan
|
||||
# teardown — the terminal-custody write `kill_workers` — starts well inside the first. The pair is one
|
||||
# budget, not two knobs: raise the drain only with the launcher wait (which ships with a release).
|
||||
LAUNCHER_STOP_GRACE_SEC = 10.0
|
||||
SERVER_GRACEFUL_SHUTDOWN_TIMEOUT_SEC = 3.0
|
||||
NESTED_SETTLEMENT_MARGIN_SEC = 30 # Structural ordering margin, not a cognition timeout.
|
||||
# Owner-note cadence while a task waits out a provider-connection outage; the effective interval is min(this, idle_timeout/2) so the notes also keep the idle rail alive.
|
||||
NETWORK_WAIT_NOTE_INTERVAL_SEC = 300
|
||||
|
|
|
|||
23
server.py
23
server.py
|
|
@ -199,6 +199,7 @@ def _restart_current_process(host: str, port: int) -> None:
|
|||
)
|
||||
|
||||
from ouroboros.config import (
|
||||
SERVER_GRACEFUL_SHUTDOWN_TIMEOUT_SEC,
|
||||
SETTINGS_DEFAULTS,
|
||||
SettingsIntegrityError,
|
||||
load_settings, save_settings, verify_settings_integrity,
|
||||
|
|
@ -1596,6 +1597,23 @@ def _emergency_process_cleanup(*, port_sweep: bool = True) -> None:
|
|||
except Exception:
|
||||
pass
|
||||
|
||||
class _SignalStopServer(uvicorn.Server):
|
||||
"""uvicorn.Server whose SIGTERM/SIGINT handler stops the supervisor loop AT THE SIGNAL.
|
||||
|
||||
The launcher's stop may SIGTERM the whole server process group, so the multiprocessing
|
||||
Manager and pooled workers can be gone before uvicorn's drain reaches the lifespan
|
||||
teardown (#1142). Setting the stop event here, not only in the lifespan ``finally``,
|
||||
makes the loop leave its tick and keeps a torn-Manager BrokenPipe from counting as a
|
||||
supervisor crash — the false "Supervisor loop died" owner alarm. A restart request
|
||||
(``should_exit`` without a signal) is unchanged: ``_restart_requested`` already ends
|
||||
the loop, and the lifespan ``finally`` still sets the event on every path.
|
||||
"""
|
||||
|
||||
def handle_exit(self, sig: int, frame) -> None:
|
||||
_supervisor_stop.set()
|
||||
super().handle_exit(sig, frame)
|
||||
|
||||
|
||||
def main() -> int:
|
||||
# A benchmark-owned child may receive an integrity pin from its parent.
|
||||
# Verify the exact bytes before even resolving the saved bind host; a
|
||||
|
|
@ -1635,8 +1653,11 @@ def main() -> int:
|
|||
log_level="warning",
|
||||
ws_ping_interval=20,
|
||||
ws_ping_timeout=20,
|
||||
# Bound the open HTTP/WS drain so the lifespan teardown (terminal custody) starts inside
|
||||
# the launcher's stop budget instead of leaving terminalization to the next boot (#1142).
|
||||
timeout_graceful_shutdown=SERVER_GRACEFUL_SHUTDOWN_TIMEOUT_SEC,
|
||||
)
|
||||
server = uvicorn.Server(config)
|
||||
server = _SignalStopServer(config)
|
||||
_uvicorn_exited = threading.Event()
|
||||
|
||||
def _check_restart():
|
||||
|
|
|
|||
|
|
@ -77,7 +77,9 @@ CHAPTER_BYTE_BUDGETS: dict[str, int] = {
|
|||
# displace; two neighbouring sentences were compressed by 168 bytes first.
|
||||
# +300: the queue snapshot and the supervisor focus event carry the root's
|
||||
# bounded authored focus (cross-focus awareness).
|
||||
"docs/architecture/05-supervisor-loop.md": 30900,
|
||||
# 30900 -> 31150 (issue #1142): the crash counter's shutdown exemption names WHERE the stop
|
||||
# event is set (the uvicorn signal handler, then the lifespan teardown) and why both are needed.
|
||||
"docs/architecture/05-supervisor-loop.md": 31150,
|
||||
# 286850 -> 287600: "an answer that has not arrived is a gap" is a new invariant of
|
||||
# plan review and task acceptance (the slot census vocabulary, the `awaiting`
|
||||
# projection, the only-awaited task outcome); the in-flight sentence it grew from is
|
||||
|
|
@ -130,7 +132,10 @@ CHAPTER_BYTE_BUDGETS: dict[str, int] = {
|
|||
"docs/architecture/07-configuration.md": 36991,
|
||||
# 18947 -> 19287: CI failure collection now documents diagnostic desktop builds while release remains gated.
|
||||
"docs/architecture/08-git-branching-ci-and-build.md": 19287,
|
||||
"docs/architecture/09-shutdown-and-process-cleanup.md": 12405,
|
||||
# 12405 -> 13600 (issue #1142): the ordinary-close paragraph gains the mechanism the chapter had
|
||||
# no text for — graceful stop signals the server PID only, the server half (stop event at the
|
||||
# signal, bounded uvicorn drain) is self-sufficient against an old group-SIGTERM launcher.
|
||||
"docs/architecture/09-shutdown-and-process-cleanup.md": 13600,
|
||||
# 17655 -> 20400: the supervisor-reliability sprint adds eight invariants the chapter lacked
|
||||
# (typed permanent engine refusal, interrupted parent, stalled-loop facts, source-ack
|
||||
# pre-check, host-owed round, reviewer tool bound, off-thread custody, fence transport) —
|
||||
|
|
|
|||
|
|
@ -397,7 +397,7 @@ def test_main_normal_exit_does_not_run_emergency_cleanup(monkeypatch, tmp_path):
|
|||
monkeypatch.setattr(server, "find_free_port", lambda _host, port: port)
|
||||
monkeypatch.setattr(server, "write_port_file", lambda *_a, **_k: None)
|
||||
monkeypatch.setattr(server.uvicorn, "Config", lambda *a, **k: object())
|
||||
monkeypatch.setattr(server.uvicorn, "Server", FakeServer)
|
||||
monkeypatch.setattr(server, "_SignalStopServer", FakeServer) # the main() server seam (#1142)
|
||||
monkeypatch.setattr(server, "_emergency_process_cleanup", lambda: cleanup_calls.append("cleanup"))
|
||||
monkeypatch.setattr(server, "_event_loop", None) # the watcher's close_all_ws hop needs no loop here
|
||||
server._restart_requested.clear()
|
||||
|
|
@ -435,7 +435,7 @@ def test_main_graceful_restart_cleanup_avoids_port_sweep(monkeypatch, tmp_path):
|
|||
monkeypatch.setattr(server, "find_free_port", lambda _host, port: port)
|
||||
monkeypatch.setattr(server, "write_port_file", lambda *_a, **_k: None)
|
||||
monkeypatch.setattr(server.uvicorn, "Config", lambda *a, **k: object())
|
||||
monkeypatch.setattr(server.uvicorn, "Server", FakeServer)
|
||||
monkeypatch.setattr(server, "_SignalStopServer", FakeServer) # the main() server seam (#1142)
|
||||
monkeypatch.setattr(server, "_LAUNCHER_MANAGED", True)
|
||||
monkeypatch.setattr(server, "_emergency_process_cleanup", lambda **kw: cleanup_calls.append(kw))
|
||||
monkeypatch.setattr(server.os, "_exit", lambda code: (_ for _ in ()).throw(ExitCalled(code)))
|
||||
|
|
|
|||
|
|
@ -427,6 +427,8 @@ def test_launcher_forced_stop_passes_retained_subtree(monkeypatch):
|
|||
class Process:
|
||||
pid = 9990011
|
||||
waits = 0
|
||||
def terminate(self):
|
||||
calls.append("graceful")
|
||||
def wait(self, timeout):
|
||||
self.waits += 1
|
||||
if self.waits == 1:
|
||||
|
|
@ -435,7 +437,6 @@ def test_launcher_forced_stop_passes_retained_subtree(monkeypatch):
|
|||
monkeypatch.setattr(launcher, "_agent_proc", Process())
|
||||
monkeypatch.setattr(launcher, "_agent_job", None)
|
||||
monkeypatch.setattr(launcher, "_retained_shared_daemon_pids", lambda: {9990022})
|
||||
monkeypatch.setattr(launcher, "terminate_process_tree", lambda proc: calls.append("graceful"))
|
||||
monkeypatch.setattr(launcher, "kill_process_tree", lambda proc, **kw: calls.append((proc.pid, kw)))
|
||||
monkeypatch.setattr(launcher, "_cleanup_recorded_server_group_for_pid", lambda *a: None)
|
||||
monkeypatch.setattr(launcher, "_kill_stale_on_port", lambda *a: None)
|
||||
|
|
|
|||
203
tests/test_shutdown_signal_1142.py
Normal file
203
tests/test_shutdown_signal_1142.py
Normal file
|
|
@ -0,0 +1,203 @@
|
|||
"""Ordinary close must not end as a SIGKILL or a false "supervisor died" alarm (#1142).
|
||||
|
||||
Two mechanisms, two halves:
|
||||
|
||||
* server half — ``server._SignalStopServer.handle_exit`` sets ``_supervisor_stop`` at the
|
||||
signal, and ``uvicorn.Config`` carries ``timeout_graceful_shutdown`` so the lifespan
|
||||
teardown (the terminal-custody write) starts inside the launcher's stop budget;
|
||||
* launcher half — the POSIX graceful phase signals only the server PID, leaving the group
|
||||
SIGKILL fallback and the post-exit group sweep untouched.
|
||||
|
||||
The real-process test deliberately reproduces the OLD launcher behaviour (SIGTERM to the
|
||||
whole group), because packaged launchers are immutable until the next release: the server
|
||||
half must be self-sufficient against it.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import json
|
||||
import os
|
||||
import signal
|
||||
import socket
|
||||
import subprocess
|
||||
import sys
|
||||
import time
|
||||
import urllib.request
|
||||
|
||||
import pytest
|
||||
|
||||
REPO_ROOT = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
|
||||
|
||||
|
||||
def _read_rows(path):
|
||||
if not os.path.exists(path):
|
||||
return []
|
||||
with open(path, encoding="utf-8") as handle:
|
||||
return [json.loads(line) for line in handle if line.strip()]
|
||||
|
||||
|
||||
def test_signal_handler_stops_the_supervisor_loop_at_the_signal():
|
||||
import uvicorn
|
||||
|
||||
import server
|
||||
|
||||
server._supervisor_stop.clear()
|
||||
try:
|
||||
instance = server._SignalStopServer(uvicorn.Config(lambda scope, receive, send: None))
|
||||
instance.handle_exit(signal.SIGTERM, None)
|
||||
assert server._supervisor_stop.is_set(), "the loop must learn about the teardown at the signal"
|
||||
assert instance.should_exit is True, "uvicorn's own graceful exit must still be requested"
|
||||
assert instance.force_exit is False
|
||||
finally:
|
||||
server._supervisor_stop.clear()
|
||||
|
||||
|
||||
def test_main_server_bounds_the_graceful_drain_from_the_shared_constant():
|
||||
import inspect
|
||||
|
||||
import server
|
||||
from ouroboros.runtime_limits import LAUNCHER_STOP_GRACE_SEC, SERVER_GRACEFUL_SHUTDOWN_TIMEOUT_SEC
|
||||
|
||||
source = inspect.getsource(server.main)
|
||||
assert "timeout_graceful_shutdown=SERVER_GRACEFUL_SHUTDOWN_TIMEOUT_SEC" in source
|
||||
assert "_SignalStopServer(config)" in source
|
||||
# The drain plus the lifespan's bounded supervisor join must leave the launcher budget room
|
||||
# for the terminal-custody write itself.
|
||||
assert SERVER_GRACEFUL_SHUTDOWN_TIMEOUT_SEC + 2 < LAUNCHER_STOP_GRACE_SEC
|
||||
|
||||
|
||||
def test_launcher_graceful_phase_signals_only_the_server_pid(monkeypatch):
|
||||
import launcher
|
||||
|
||||
calls: list = []
|
||||
|
||||
class Process:
|
||||
pid = 4242
|
||||
|
||||
def terminate(self):
|
||||
calls.append("terminate")
|
||||
|
||||
def wait(self, timeout):
|
||||
calls.append(("wait", timeout))
|
||||
|
||||
monkeypatch.setattr(launcher, "IS_WINDOWS", False)
|
||||
monkeypatch.setattr(launcher, "_agent_proc", Process())
|
||||
monkeypatch.setattr(launcher, "_agent_job", None)
|
||||
monkeypatch.setattr(launcher.os, "killpg", lambda *a, **k: calls.append("killpg"))
|
||||
monkeypatch.setattr(launcher, "kill_process_tree", lambda proc, **kw: calls.append("kill_tree"))
|
||||
monkeypatch.setattr(launcher, "_cleanup_recorded_server_group_for_pid", lambda *a: None)
|
||||
launcher.stop_agent()
|
||||
assert calls == ["terminate", ("wait", launcher.LAUNCHER_STOP_GRACE_SEC)]
|
||||
assert not hasattr(launcher, "terminate_process_tree"), "the group SIGTERM helper left the launcher"
|
||||
|
||||
|
||||
def _free_port() -> int:
|
||||
with socket.socket() as sock:
|
||||
sock.bind(("127.0.0.1", 0))
|
||||
return int(sock.getsockname()[1])
|
||||
|
||||
|
||||
def _wait_json(url: str, key: str, timeout_sec: float) -> None:
|
||||
deadline = time.time() + timeout_sec
|
||||
last = ""
|
||||
while time.time() < deadline:
|
||||
try:
|
||||
with urllib.request.urlopen(url, timeout=2) as resp: # noqa: S310 - local test server
|
||||
if json.loads(resp.read().decode("utf-8")).get(key) is True:
|
||||
return
|
||||
except Exception as exc: # pragma: no cover - diagnostic only
|
||||
last = str(exc)
|
||||
time.sleep(0.25)
|
||||
raise RuntimeError(f"{url} never reported {key}=true: {last}")
|
||||
|
||||
|
||||
@pytest.mark.serial
|
||||
@pytest.mark.skipif(sys.platform == "win32", reason="process groups and SIGTERM are POSIX")
|
||||
def test_group_sigterm_reaches_terminal_custody_without_a_false_supervisor_alarm(tmp_path):
|
||||
"""A real server in an isolated data root, SIGTERMed the way the OLD launcher does (whole
|
||||
group, so its Manager dies first) while a browser WebSocket is open: the lifespan teardown
|
||||
must still run inside the launcher budget and the owner chat must get no supervisor_failure."""
|
||||
from ouroboros.process_containment import ProcessContainer
|
||||
from ouroboros.runtime_limits import LAUNCHER_STOP_GRACE_SEC
|
||||
|
||||
port = _free_port()
|
||||
data_dir = tmp_path / "data"
|
||||
(data_dir / "state").mkdir(parents=True)
|
||||
home = tmp_path / "home"
|
||||
home.mkdir()
|
||||
model = "openai-compatible::never-called"
|
||||
(data_dir / "settings.json").write_text(json.dumps({
|
||||
"OPENAI_COMPATIBLE_API_KEY": "shutdown-test-key",
|
||||
"OPENAI_COMPATIBLE_BASE_URL": f"http://127.0.0.1:{port + 2}/v1",
|
||||
"OUROBOROS_MODEL": model, "OUROBOROS_MODEL_LIGHT": model, "OUROBOROS_MODEL_FALLBACKS": model,
|
||||
"OUROBOROS_MAX_WORKERS": 1,
|
||||
"OUROBOROS_RUNTIME_MODE": "light",
|
||||
}), encoding="utf-8")
|
||||
# An owner chat is bound, so the false alarm WOULD be written if the crash counter fired.
|
||||
(data_dir / "state" / "state.json").write_text(json.dumps({"owner_chat_id": 1}), encoding="utf-8")
|
||||
env = {
|
||||
**os.environ,
|
||||
"HOME": str(home),
|
||||
"OUROBOROS_APP_ROOT": str(tmp_path),
|
||||
"OUROBOROS_DATA_DIR": str(data_dir),
|
||||
"OUROBOROS_SETTINGS_PATH": str(data_dir / "settings.json"),
|
||||
"OUROBOROS_REPO_DIR": REPO_ROOT,
|
||||
"OUROBOROS_SERVER_HOST": "127.0.0.1",
|
||||
"OUROBOROS_SERVER_PORT": str(port),
|
||||
"OUROBOROS_HOST_SERVICE_PORT": str(port + 1),
|
||||
"OUROBOROS_MANAGED_BY_LAUNCHER": "1",
|
||||
}
|
||||
url = f"http://127.0.0.1:{port}"
|
||||
container = ProcessContainer()
|
||||
proc = container.spawn(
|
||||
[sys.executable, "server.py"], cwd=REPO_ROOT, env=env,
|
||||
stdout=subprocess.DEVNULL, stderr=subprocess.DEVNULL,
|
||||
)
|
||||
ws = None
|
||||
stuck = None
|
||||
try:
|
||||
_wait_json(f"{url}/api/state", "supervisor_ready", 90)
|
||||
time.sleep(1.0) # let the loop run a few ticks against the live Manager
|
||||
from websockets.sync.client import connect
|
||||
|
||||
ws = connect(f"ws://127.0.0.1:{port}/ws", open_timeout=10) # the desktop window's socket
|
||||
# An in-flight request whose body never completes: uvicorn's graceful drain waits for
|
||||
# it, which is what parked the lifespan teardown behind the launcher budget.
|
||||
stuck = socket.create_connection(("127.0.0.1", port), timeout=5)
|
||||
stuck.sendall(
|
||||
b"POST /api/settings HTTP/1.1\r\nHost: 127.0.0.1\r\n"
|
||||
b"Content-Type: application/json\r\nContent-Length: 100000\r\n\r\n{\"OUROBOROS_MODEL\": \""
|
||||
)
|
||||
time.sleep(0.5)
|
||||
started = time.monotonic()
|
||||
os.killpg(os.getpgid(proc.pid), signal.SIGTERM) # exactly what the pre-#1142 launcher does
|
||||
try:
|
||||
code = proc.wait(timeout=LAUNCHER_STOP_GRACE_SEC)
|
||||
except subprocess.TimeoutExpired:
|
||||
pytest.fail("server did not exit within the launcher stop budget — the next launcher would SIGKILL it")
|
||||
elapsed = time.monotonic() - started
|
||||
finally:
|
||||
if stuck is not None:
|
||||
stuck.close()
|
||||
if ws is not None:
|
||||
try:
|
||||
ws.close()
|
||||
except Exception:
|
||||
pass
|
||||
if proc.poll() is None:
|
||||
proc.kill()
|
||||
error = container.reap()
|
||||
container.close()
|
||||
assert not error, error
|
||||
|
||||
# uvicorn re-raises the captured SIGTERM once the lifespan has completed, so a clean
|
||||
# signalled exit reads -SIGTERM (or 0); -SIGKILL is the launcher fallback this test forbids.
|
||||
assert code in (0, -signal.SIGTERM), f"graceful exit expected, got {code}"
|
||||
chat_rows = _read_rows(data_dir / "logs" / "chat.jsonl")
|
||||
alarms = [row for row in chat_rows if row.get("system_type") == "supervisor_failure"
|
||||
or "Supervisor loop died" in str(row.get("text") or "")]
|
||||
assert alarms == [], alarms
|
||||
shutdown_rows = [row for row in _read_rows(data_dir / "logs" / "supervisor.jsonl")
|
||||
if row.get("type") == "server_shutdown"]
|
||||
assert shutdown_rows and shutdown_rows[-1].get("cause") == "external_signal", shutdown_rows
|
||||
assert elapsed < LAUNCHER_STOP_GRACE_SEC, elapsed
|
||||
Loading…
Add table
Add a link
Reference in a new issue