mirror of
https://github.com/razzant/ouroboros.git
synced 2026-10-03 04:07:04 +00:00
Centralize post-admission drive settlement, preserve captured identities and complete input closures, make metadata reads pure, serve confined nested files and directory archives, and keep maintenance off the supervisor loop. Preserve generation fences at actual mutation boundaries and truthful queued forwarding receipts.
68 lines
33 KiB
Markdown
68 lines
33 KiB
Markdown
# 5. Supervisor Loop
|
|
|
|
This chapter owns the single scheduler for pooled work: what a healthy tick does, what the queue holds and what its durable snapshot may restore, how a task is addressed and named at admission, how owner waits lend capacity, and the intent-then-custody skeleton every cancellation follows. It exists because these invariants decide whether a stopped task ends honestly or leaves a ghost, and none of them can be reconstructed from any single module's code.
|
|
|
|
`server.py::_run_supervisor()` is the single scheduler for pooled tasks. A healthy tick publishes liveness, rotates runtime logs, checks worker health, drains one BOUNDED batch of worker and direct-chat events (at most `SUPERVISOR_EVENT_BATCH_MAX_EVENTS`; the `SUPERVISOR_EVENT_BATCH_MAX_SEC` budget is checked between handlers, so one slow handler can overrun it; FIFO, the remainder next turn; a wake-up's frames are a direct turn's), accepts owner bridge input, writes the compatibility budget projection once when `llm_usage` events marked it dirty, enforces deadlines and schedules, runs throttled reconciliation and evolution admission, assigns eligible work, persists `state/queue_snapshot.json`, and last ticks the consciousness alarm clock, which owns no thread — this pass is the only thing that can start a wake-up. Bridge intake precedes the projection write, timeout, maintenance, evolution and assignment work so a slow control-plane step cannot hide a new owner message, and the events bound keeps a producer that never lets the queue empty from hiding one either; a turn that hit its bound skips the idle sleep so a backlog still drains at full speed. Three consecutive loop failures clear supervisor readiness, stop its watchdog generation and notify the owner, rather than leave a healthy-looking server that no longer assigns work; a failure while a shutdown or restart is in progress is not a crash and never holds the shutdown — the process-local stop event is set at the SIGTERM/SIGINT instant by `server_process._SignalStopServer.handle_exit` and again first thing in the lifespan teardown, because a launcher that signals the whole group kills the event Manager before the teardown can run (issue #1142; §9).
|
|
|
|
The legacy `state.json` budget projection reads the validated usage ledger before taking `STATE_LOCK`, so a multi-megabyte replay cannot hold the short reader lock while waiting on the monetary lock. The same read returns a ledger provenance marker `(compaction_epoch, seq)`: the epoch comes from the lock-free leading baseline header and `seq` is the live file's validated high-water sequence. A compaction increments the epoch while renumbering live rows, so the pair remains ordered even when the file gets shorter; neither timestamps nor totals prove order. Inside `STATE_LOCK`, any marker strictly lower than the saved one — a lower epoch, or a lower sequence within the same epoch — is proven stale and leaves the compatibility projection untouched. An equal or higher marker is written. Restore to an older ledger or a malformed saved marker freezes this projection until explicit repair; automatic recovery is absent. Quarantine-file presence likewise keeps `integrity_degraded` and freezes this compatibility projection indefinitely. This read is display (invariant 28); admission is exact. An unavailable or malformed marker leaves the prior projection untouched; this fail-safe also applies when lock acquisition times out and the writer deliberately proceeds without the lock. On that no-lock path the marker still rejects a demonstrably stale snapshot, but the compare-and-save sequence remains non-atomic for concurrent no-lock writers, as it was before this protection. The persisted projection carries totals only; per-root money is a ledger render (`usage_projection`), never a `state.json` key. The writer renders only what it persists (`usage_writer_snapshot`: totals, marker, the OpenRouter bucket its drift check compares, the totals-only projection), once per loop turn after bridge intake; a refused or failed write keeps the projection dirty and is retried no more often than `BUDGET_PROJECTION_RETRY_SEC`, and the OpenRouter ground-truth check fires on crossing each multiple of 50 physical calls. `/api/state` and the budget line read the same slim shapes.
|
|
|
|
`PENDING` and `RUNNING`, guarded by `supervisor.queue._queue_lock`, are the live task-lifecycle authority. Admission reserves identity before project, workspace, attachment or routing side effects can create a duplicate; refuses a disabled pool, duplicate task, project deletion, accepted or sealed root, exhausted root budget, or an unusable or unprovisionable workspace (typed, with the repair in `detail`); attaches the task contract; and preserves stable priority order. Assignment runs on the same locked state and skips reaping slots, budget-paused work, closed project roots, conflicting project writers, and tasks exceeding the root's subagent capacity or depth reservation. Evolution alone is fenced by runtime mode — blocked in Light — three times (`supervisor/evolution_lifecycle.py`): entry points refuse a campaign start, `enqueue_evolution_task_if_needed()` pauses and disables a carried campaign, and assignment drops what slipped through while `evolution_block_reason()` is set; generic `supervisor.queue.enqueue_task()` has no runtime-mode predicate. Configured worker count is therefore not available capacity: the truthful value is the assignable idle count after custody, reaping and admission fences.
|
|
|
|
A headless task is ADDRESSED when it is admitted, not when it is displayed (`log_addressing.ingress_chat_id`). A registered project's run has exactly ONE destination: an explicit `chat_id` may only agree with that thread, and any other value — the hidden partition included — is refused with a typed 400 rather than honoured or silently overridden, because a run addressed away from its room puts a card in Main whose project holds none of its work. Without a Project, ordinary API tasks default to `HIDDEN_CHAT_ID` (0); the confirmed browser Publish flow carries `source="web"` and `WEB_UI_CHAT_ID` to request Main (caller-declared addressing, not an authentication proof), and other non-Project conversation addresses stay refused. A run scoped to a REGISTERED, active project is admitted into that project's thread, and Main receives the one host-stamped completion row only when the work is actually in that room: addressed there at admission or BOUND to the project. Registration alone does not qualify. Every other run stays in the hidden partition, silent in every chat, read back through the terminal, `--result-json-out`, the chat-blind Logs panel and `GET /api/tasks/<id>`. A reserved but inactive project keeps its chat acceptable so the queue's lifecycle fence refuses with its own typed reason.
|
|
|
|
The run is also NAMED at admission and chat promotion, without a new model call: a caller-supplied `title` (`ouroboros run --title` or the top-level contract field; `metadata.title` is refused with a 400 like `metadata.project_id`) is authorship and fills both `title` and `suggested_name`; otherwise the request's first line fills `suggested_name` ALONE, so a truncated prompt never outranks a real name coined later. A `task_named` frame is broadcast on admission so the live card is never born showing its status phrase as a title.
|
|
|
|
`queue_snapshot.json` projects pending/running rows atomically, acceptance/root-budget fences, worker counts, assignable capacity, pool-disabled reason and bounded root focus; it is not another scheduler. Startup restores recent snapshots into empty PENDING. Ordinary surviving RUNNING rows receive durable `server_shutdown` cancel intents (`terminalized_running`); cancellation custody later terminalizes them, expires quizzes and closes owner waits. Boot `kill_workers` settles their PENDING children as `pending_parent_interrupted`, preventing ghosts/orphans. Only acknowledged planned-restart owner handoffs and exact budget pauses outlive snapshot age.
|
|
|
|
Exact pauses restore `_budget_pause`/checkpoint and `_is_direct_chat` without wake; unreadable sources or invalid/acceptance fences retain typed `_budget_pause_hold`. Complete RUNNING pauses park, not crash-fence. `_budget_pause_resume` grants are revoked if unconsumed at restart or money disappears before dispatch. Assignment carries original `started_at` and cumulative `budget_paused_sec`; child selection needs the current root grant/fence. Details: [§6 Exact budget pause and Resume](06-agent-core.md#exact-budget-pause-and-resume).
|
|
|
|
Terminal tasks stay terminal; active cancellation belongs to custody; descendants of accepted/sealed roots finalize cancelled except held exact pauses above. Malformed fences fail closed. Assigned tasks mirror RUNNING into its durable result: both orphan healers read STORED status, so an unmirrored root would escape snapshot-less cleanup. Root orphan settlement closes quiz/wait through the same terminal owner. Focus merges under the queue lock via worker events, with no new timer/ledger/wake. `direct_roots.json` projects direct roots and aggregate gap/freshness.
|
|
|
|
`supervisor/queue_schedules.py` owns the existing `state/scheduled_tasks.json` table and is its only writer. Every read-modify-write — the tick, the skill resync, the gateway upsert, `schedule_followup`'s cap-and-write, `manage_schedules` — enters through `schedule_transaction`, which takes BOTH locks itself, queue lock then the table's sidecar file lock. The order lives in the transaction, not in its callers: a caller-composed order is what inverted, when the follow-up tool's one-transaction cap-read-and-write reached for the queue lock inside that hold against the tick's queue-then-table. It is reentrant PER TABLE (keyed by the resolved lock path), so a nested call on another drive root still takes that root's lock rather than riding a depth counter. `load_schedule_store` is the strict read every writer and owner surface uses: only an ABSENT path is an empty table it may create, while a present non-regular file, bytes that do not parse and a row that is not an object each raise `ScheduleStoreUnreadable` naming which it was. A write is refused on it — the next atomic write would replace real rows with an empty document or drop the rows it could not parse — and a READ answers unavailable, because "no schedules" is a claim while an unparseable table means the state is unknown.
|
|
|
|
A mutation announces its INTENT to `logs/events.jsonl` before changing anything and its OUTCOME after, as bounded `schedule_mutation` facts sharing one `operation_id`; the task template never enters the event. A failed intent write refuses the mutation; a failed outcome write is disclosed as `changed_audit_incomplete`, never rolled back (an undo would be a second unaudited mutation). `manage_schedules` and `POST /api/schedules/{id}/action` are the same seam: the caller NAMES the action and gives a reason, and nothing infers a command from that text. Restore is owner/root authority; Observe's argument narrowing does not redefine it. Lifecycle actions govern FUTURE dispatch only and report `running_or_queued` rather than implying a stop — `null` when the answering process cannot see the live queue, because PENDING/RUNNING are the supervisor process's own dicts and a worker-side answer of `false` would claim nothing is in flight; a worker reads the durable queue snapshot, rewritten on exactly those transitions (a snapshot older than 300 s also answers unknown, idle periods included), and the tick treats unknown as still-in-flight. A skill row is retained as a suppressed record (keyed by schedule id, so it survives a manifest edit) that resync may not re-arm; restore re-evaluates the skill's presence, manifest declaration, `supervised_task` permission and readiness instead of enabling it, reporting `restored_not_ready` with the blocker when it cannot; an ABSENT skill or manifest entry keeps the suppression and refuses typed (`manifest_absent`, `changed=false`) — a marker lifted with nothing to restore into is forgotten by the next resync (`supervisor/schedule_lifecycle.py`); a one-shot with `completed_at` is consumed history only a fresh `run_at` replaces, and a timezone-only edit re-reads the same instant rather than clearing that receipt. A skill row's `enabled` is reconciled from readiness, so the upsert REFUSES an edit that would flip it and names the lifecycle action instead — the owner's durable statement is the suppression marker, not a flag the next resync overwrites. Retention stays the unified GC.
|
|
|
|
At actual execution start, `agent._persist_running_record` mirrors a split root's running timestamp and execution-drive address into its canonical result through the existing terminal-preserving writer.
|
|
|
|
Pooled completion separates a finished file-save attempt from publishable terminal truth: the worker prepares its terminal task files (`headless.prepare_terminal_task_files`) before its buffered `task_done`, and `events_task_done` publishes only when `headless.terminal_task_files_ready` confirms CURRENT — on a split drive the child-bound copyback, with no workspace artifact finalization pending. Legacy or faulted completions go through `enqueue_terminal_file_recovery` (`worker_health` prepares and recovers, `task_reaper` runs the queues): unreadable or missing CURRENT keeps RUNNING ownership and retries on the health cadence — never a false Done, never replayed model work — and only CONFIRMED absence of a terminal source reaches the lifecycle-fault owner, which marks execution `infra_failed` while an early sticky completed status, the authored answer, review and cost survive.
|
|
|
|
A required owner wait keeps the task RUNNING, its worker and browser after a completed-tool checkpoint. Capacity is lent; a capped grant resumes the same attempt/start without replay, consuming continuation authority before dispatch. Any mail wakes, only an owner answer answers; Stop, deadline and ceiling still bind. A settled Project result is no longer steerable even if post-work holds RUNNING: fast mail and typed steer refuse it and a quiz answer takes the late-answer path (one actor-drive settlement predicate, `owner_mailbox.mailbox_drain_ended`), so cleanup cannot eat unread input. Ordinary Main/Project waits stay in-process (`owner_wait.direct_owner_wait`), lend no pooled slot and resume the live stack/browser on addressed text or quiz; a saved source grants no cold restart.
|
|
|
|
A crash or automatic timeout retry keeps the current-attempt checkpoint as evidence and never blindly replays the task; cold continuation exists only for a confirmed planned restart, which cannot preserve an OS browser session. An owner-requested MANAGED UPDATE is such a restart: its writer fence prepares the handoff before stopping the pool and passes the parked ids as `preserve_running_task_ids`, so a wait is requeued instead of interrupted, and a preparation failure blocks the update rather than terminalizing the wait; the re-exec seam arms the transaction only in the phases `_safe_restart_serialized` allows (launcher mode observes exit code 42), no-resume flags suppress it, and an aborted update's leftover record never authorizes a later manual Restart or rollback. A manual ROLLBACK returns to an older runtime, so it deliberately parks nothing. Cold continuation resumes the saved economics — original CostCeiling, hard clocks, the saved model's ContextFit, TaskModelWait choices, the pending budget decision of the saved round — and calendar deadlines and owner-wait time are unchanged.
|
|
|
|
Cancellation is intent-then-custody; intent and outcome are separate fields, because one field carrying both wedges a task forever. The skeleton:
|
|
|
|
1. **Intent.** Every cancel ingress — the agent `cancel_task` tool, the HTTP single and cascade endpoints, evolution stop, project deletion, a cascade sweep's per-descendant mints, the boot migration of legacy `cancel_requested` files — writes one durable row through `ouroboros/cancel_intents.request_cancel` into the locked projection `state/cancel_intents.json` (active intents only; every transition also appends a forensic `cancel_intent` supervisor-ledger row). Every ingress fails closed — tool `CANCEL_INTENT_WRITE_FAILED`, HTTP 503, `CANCEL_INTENT_PROJECTION_CORRUPT` for a corrupt projection — so no teardown runs without a durable, watchdog-replayable fence; evolution stop keeps any task whose intent write failed and reports the stop INCOMPLETE with typed per-task outcomes — the campaign stays OPEN under the durable `evolution_owner_stopped` flag, which only an OWNER start ingress clears (the agent's `toggle_evolution` against it is refused: the owner's stop is sticky).
|
|
2. **Scope.** A cascade ingress mints `scope: cascade`; recorded scope is widen-only (single→cascade, never narrowed, so Stop-now cannot shrink a cascade). A cascade over an already-settled root with live descendants still mints the coordination intent (`allow_settled_target`), the watchdog's replay trigger for the subtree. Timeout reaping is deliberately NOT a cancel ingress: the reaper keeps its own custody over the `reaping` slot marker, because cancellation is reserved for explicit intent.
|
|
3. **Claim.** `supervisor.task_lifecycle.cancel_task_custody` is the ONE settle owner: it claims the intent (owner + generation) before any custody mutation, and a refused claim exits having touched nothing, so racing custodies never double-settle; the pre-assignment pending drop holds the same fence, and a budget-exhausted queued task PAUSES, so there is no batch terminalizer beside custody. A LIVE direct-chat turn is stopped through the same custody (`supervisor/worker_chat_lane.py`): the chat lane writes a typed `finalize_now` control into the owner mailbox once per turn, and the loop ends at its next round boundary with zero further model calls, paid post-task synthesis included. A Stop-now landing after the loop returned, while that synthesis runs, is still addressable: the in-flight synthesis counts as live ownership, custody keeps the immediate intent open, and the synthesis worker checks it before each paid stage, disclosing the skipped ones as `post_task_stop_reason` `owner_stopped:skipped=<stages>` on a `degraded` checkpoint. Custody waits `OUROBOROS_DIRECT_TURN_STOP_WAIT_SEC`; the typed outcome is `gone` / `ended` / `live`, and `live` releases the claim for the sweep rather than publishing a fabricated `cancelled` row over a turn that is still running.
|
|
4. **Kill and re-check.** Custody confirms process death, then re-reads the child's real settled result. Natural completion WINS: a child that finished before the kill keeps its result, artifacts and cost, and the cancel settles as already-settled.
|
|
5. **Reconcile and capture.** The task's open delegated runs are reconciled from durable custody rows and always re-audited and disclosed. Workspace artifacts are captured from the real tree; a failed or owed-but-unrunnable capture is `failed`, never `missing`, and a shared-tree capture carries `attribution: shared_unproven`.
|
|
6. **Settle.** The settled result carries reconstructed-or-honestly-unknown cost, never a fabricated final `$0`; `parent_decision` is stamped only at this outcome. `cancel_publication._intent_outcome_fields` preserves recorded cancellation provenance as `cancel_origin`: source, scope, reason, `request_id`/`requested_at`, `requested_by` when present, and the typed observation `request_origin`. These facts survive removal of the active intent and travel through the terminal event, task detail, history and result-tool reads, including conditional reads whose answer body is unchanged. HTTP proves transport, never a personal owner; absent actor evidence remains absent. The existing `requested_by` condition for parent-decision semantics is unchanged.
|
|
7. **Owe, then publish.** The owner's terminal answer is registered as OWED in the durable outbox (or a typed no-chat handoff row) BEFORE the intent settles and before `task_done` publishes, so a crash between settle and send replays the answer instead of losing it. A cascade delivers one root message with a children digest under the deterministic delivery id `cascade:<root_tid>:<request_id>`, each child's line rebuilt from its current durable status.
|
|
8. **Watchdog.** The supervisor tick runs the cancel/delivery/usage sweep off drain (`server_maintenance._run_cancel_delivery_ref_sweep`, ~20 s cadence); `sweep_cancel_intents` re-feeds unclaimed or abandoned-claim intents into custody — a cascade replayed as a cascade — so a lost control event or a custody attempt that died mid-teardown cannot wedge a cancellation. Only the physical no-live check settles a cascade's coordination intent, after the tree's summary is registered as owed.
|
|
|
|
Readers see the typed projection `cancel_state: "pending"` (with `cancel_reason`) on effective results until the settle; the UI shows "Cancelling…" and restores the Cancel button only when a fetched live non-pending task detail proves the intent is gone. Steering writes — `steer_task`, mailbox follow-ups, `forward_to_worker` — are refused typed while a cancel is pending (that fence is what makes the owner-stop single-turn rail safe), and queue restore and pre-assignment consult the projection under the queue lock, so a cancelled pending task never starts. `task_done` asserts a SETTLED outcome and is validated against the DURABLE result for every event: a non-settled event status, or a settled or blank status over a non-settled or absent durable row, is a lifecycle fault — left to custody when a cancellation is pending, otherwise published with a typed infrastructure-failure axis that preserves an existing sticky terminal status.
|
|
|
|
Terminal answers ride one durable delivery seam (`supervisor/terminal_delivery.py`): a bounded PENDING outbox `state/terminal_deliveries.json` (owed before enqueue, replayed on boot and on the tick; eviction past capacity is the typed `terminal_delivery_exhausted`, never a silent pop) with restart-surviving `delivery_id` dedupe shared with the natural final-answer path; a loud UNREVIEWED salvage message (bounded preview, exact omitted count, full-copy receipt) for cancelled and non-retry-reaped tasks; one root message for a cascade; nothing for a retryable reap; routing follows the task's lineage chat. The already-settled and finalize-on-miss paths run the same delegated-run audit as the kill path, so a cancel over a dead task with live delegated runs never reads as a clean completion. An agent-requested cancel publishes nothing of its own — the custody seam's terminal rows and the typed tool result are the truth — except a FAILED settle, which speaks as the typed `cancellation_fault` progress incident. The custody, completion-wins and owed-before-published invariants are restated in §10.
|
|
|
|
Stop POLICY is an axis on the same durable intent, independent of cascade scope. An omitted or empty-body cancellation is the synchronous IMMEDIATE teardown, keeping programmatic callers' bounded budgets. An explicit `stop_policy=finalize_then_cancel` answers 202 with the intent OPEN and runs one bounded owner-stop finalization episode (`supervisor/owner_stop.py`): live descendants settle first and feed a bounded child-result projection into the root's final turn; the root receives a `finalize_now` control whose typed first line (`owner_requested_finalization`) routes to its own loop rail — zero or one tool-less model turn, terminalizing completed/best-effort under the honest owner reason rather than a false deadline reason. The grace budget starts at the durable `control_drained_at` (first drain wins, so a task inside a long tool call still gets its final turn when the hard bounds allow) under the request-anchored `OWNER_STOP_OUTER_CAP_SEC` cap; a held task bypasses only the generic idle/finalization-grace rails — its explicit deadline and absolute ceiling remain independent hard axes and are never widened — and expiry, a hard-bound hit, a pending root or an already-settled root feeds ordinary custody. Policy transitions are monotonic: an immediate request HARDENS a pending graceful intent (preserving any cascade scope); graceful can never soften an accepted immediate. A successful graceful root suppresses the redundant cascade summary; Panic bypasses both. The UI projects the soft stop through `cancel_state`+`stop_policy` ("Finalizing…").
|
|
|
|
Beside stopping sits the owner "hurry" control: a typed task-local `kind=hurry` owner-mailbox control (`ouroboros/owner_hurry.py`, `gateway/task_hurry.py`) that skips the next otherwise-eligible acceptance panel with a typed reason, zeroes remaining improvement passes, and makes force-plan projection task-locally advisory — never a chat message, never a settings mutation, never a P3/commit/review-gate weakening. Its effect is attempt-scoped (`task["_attempt"]`): a shared `retry_reset` strips it on every same-id requeue. These invariants hold for every install configuration class.
|
|
|
|
The event bus is process-lifetime rather than worker-generation-lifetime: respawns reuse one manager-backed queue shared by workers and direct chat, because a force-killed producer can corrupt a raw multiprocessing feeder frame and a queue rebuilt on pool rotation strands surviving producers on the old endpoint. Live-frame publication of persisted rows is exactly-once and process-symmetric: `ouroboros/utils.py::append_jsonl` streams only runtime `logs/*.jsonl` rows into the process log sink (never `chat.jsonl`, never state/memory/receipt stores), and each process suppresses the types whose live delivery has a dedicated owner (`WORKER_LOG_SINK_SUPPRESSED_TYPES`, the server superset `SERVER_LOG_SINK_SUPPRESSED_TYPES`). One persisted event produces exactly one live frame (`tests/test_log_forwarding.py`); an LLM call failure is one durable `llm_api_error` row and nothing else.
|
|
|
|
Heartbeat and progress are different evidence: a heartbeat proves a process or loop is alive; owner-visible progress and model-usage events prove the task advanced. Progress keeps an orchestrator alive, but deadline, Stop and budget remain hard. The absolute ceiling ends solve/owner-stop finalization, not settled post-work; Stop, deadline, money and per-call bounds still apply (§6). After the typed finalization episode (§6), timeout handling freezes its decision under the queue lock, marks the worker `reaping`, and hands kill, join, salvage, retry and respawn to the single off-loop reaper; an orchestrator with live descendants is not blindly retried, because a retry would replay its plan and spawn a competing tree. No retry or new assignment may occupy a timed-out slot until the original process is provably dead: if kill and join cannot establish death, the reaper keeps a low-rank RUNNING result and the `reaping` slot, emits a visible `task_reaper_wedged` receipt and restart hint, and writes no terminal, `task_done`, retry or respawn — one slot is sacrificed rather than letting a still-running process race a replacement and overwrite its result; the next supervisor generation reconciles the record after old-generation process custody. Typed timeout codes (`queue_timeouts.TIMEOUT_TERMINAL_REASONS`) retain their `task_incident` identity; owner-facing grace, kill and salvage text use `project_dialogue.TASK_CAUSE_PHRASES`, with unknown codes still raw.
|
|
|
|
A spawned or respawned slot is not assignable until its child's PID-bound `worker_ready` row arrives (`supervisor/worker_pool_lifecycle.py`). A live child's own `worker_starting` row, emitted before extension loading and agent construction, permits one extension of `WORKER_READY_WINDOW_SEC` to `WORKER_READY_CEILING_SEC` (300 seconds from birth, both in `runtime_limits.py`); foreign or pre-spawn rows cannot extend another slot. `worker_ready_window_extended` records that decision. A silent child keeps the original window, and logging failure cannot block startup. After `WORKER_READY_MAX_ATTEMPTS` failed attempts, `Worker.readiness_exhausted` is final for that exact slot — late events cannot reopen it. Total exhaustion, distinguished from busy/booting/reaping capacity and from a live owner-wait stack, closes pooled ingress (owner `/review` included) without blocking direct chat/control or boot/update recovery; once RUNNING completion custody has settled, `disable_exhausted_worker_pool` fails unstarted PENDING work honestly with a Restart hint, and a new task cannot clear the latch. Readiness stays separate from liveness and task idle time; a watcher error releases only still-booting, non-exhausted slots to the crash detector (`worker_ready_released`). Linux workers use forkserver; macOS and Windows use spawn.
|
|
|
|
`worker_health.recover_confirmed_dead_worker` reserves exact custody under the queue lock, then enqueues reaper `confirmed_dead_worker`. Saved terminal bytes win even after signal death; unknown/incomplete publication retains `TerminalFileRecoveryPending`. Only confirmed absence reaches crash policy: signal means infrastructure failure; eligible non-signal crashes retry within `QUEUE_MAX_RETRIES`, preserving cost and owner-wait restrictions. Budget continuation forbids ordinary retry; unconsumed grants re-park, consumed grants remain terminal, and failed revocation retains source/grant with honest persistence status (§6).
|
|
|
|
Crash storms suppress respawn while terminal sources settle, then fence pooled admission; direct chat remains. Startup runs the same recovery after process custody, before `_startup_prune_sweeps`, also in no-provider lifespan without spawning. Unknown/live ownership defers to avoid racing writers. Unresolved/protected sources or ownership/read errors set `preserve_task_sources`, skipping drive deletion. `_recover_terminal_task_files` may restore an older canonical scheduled row's start binding only from a known non-direct child's positive running/started-at record, with fresh-queue/later-boot orphan proof, no pending owner or active cancel. Existing orphan/terminal owners remain authoritative; nothing resumes. The report records `rebound`.
|
|
|
|
Startup and throttled maintenance reconcile three residue classes; the ~600 s custody and ~300 s reconcile passes each run off the loop thread on their own latch (§10). Process custody checks strict PID, start-time, command, owner-task, session and generation evidence before it reaps. Delegated-run reconciliation applies the same owner-gone reasoning to harness rows (§6 Delegated subagents). Task, review and project reconciliation repair records whose producer no longer exists; task reconciliation decides and persists on a status-only read, fenced by the row's attempt basis and queue ownership, settles quiz/owner-wait only for the row it heals, then the drive-custody pass (child-ref retry, then bounded drive settlements under the queue interlock; startup copies no child store); the pass stamps its cadence when it ends; the generation fences each item, commit and write (`task_custody.publication_fence`). None of these are command-line-class kill sweeps, and one instance never reaps another. The dedicated watchdog, started with the startup phase, separately observes phase-stamped loop liveness (a stall row carries the loop thread's bounded stack) and every native actor; a wedged chat turn alerts with a `/restart` hint, a loop stall only journals, and neither kills a thread. Other owner conversations run on independent native actors, without a second scheduler.
|
|
|
|
Cooperative project checkpointing has two equivalent quiescence triggers: a host-minted genesis or cooperative tree is checked when its root settles with no live descendants, and again when the last child settles beneath an already-terminal root — a root-scope budget stop terminalizes the root before its children, so a root-only trigger would see a live tree once and never return. The bounded git chain runs on a daemon thread, revalidates quiescence under the queue lock immediately before mutation, and replays a trigger that arrives during an in-flight check. Only host-minted project roots are eligible: owner-attached folders are never auto-committed, credential-shaped files stay excluded and disclosed, and every material success, skip or error receives a durable receipt.
|
|
|
|
Bridge intake is batched: `LocalChatBridge.get_updates` blocks only for the first item, then drains a bounded ready snapshot. Every update keeps a monotonic id and its own reply transport (`activate_update_transport`); malformed items are logged and skipped individually. Web owner ingress durably writes the canonical `chat.jsonl` row with `log_chat(require_write=True, ensure_record_boundary=True)` (as the named ingress does, so a torn prior tail cannot swallow it) before queue handoff or `ingress_accepted` echo, passing that row as the dequeue witness. Each async ingress-lock caller (`gateway/ws.py` chat, Host Service delivery, late quiz answers) runs via `gateway._helpers.run_sync_to_completion` off the ASGI loop: unrelated HTTP and sockets stay responsive, a socket's later chats remain ordered, and cancellation settles row → queue → echo. A command on any socket remains an inline queue put. A failed handler requeues the unprocessed tail with its ids for the next read; accepted `/restart` does likewise, while `/panic` never delays the hard stop for handback. Requeue is memory-only and lost on process exit; accepted rows on disk survive without automatic replay.
|
|
|
|
The bridge recognizes `/panic`, `/restart`, `/review`, `/evolve [on|off]`, `/bg [start|stop|status]` and `/status`; all other text enters ordinary agent routing. External transports may invoke these commands only with positive owner identity and a transport-specific owner-chat binding, and the commands reuse runtime-mode, queue, cancellation and typed-result authority rather than implementing parallel control paths. Runtime logs rotate on the same supervisor tick and archive readers preserve their retained timelines; only explicitly isolated devtool roots may use the narrow rotation sentinel from §1.
|