One Ouroboros across Main and project rooms: each root may publish a short authored focus (update_focus) that rides the durable task result and the [INDEPENDENT_ROOTS] tail, so concurrent foci can see one another without a shared chat, a new wake or any widened authority; explicit cross-room journal/workpad reads are honoured instead of silently redirected. Dialogue consolidation summarizes each source room of a logical chunk from its own bytes (Light draft + source-grounded Light correction), assembles the typed room sections deterministically into one shared block and carries them through era compression; a failed room withholds its whole chunk while earlier complete chunks stay published; legacy mixed blocks keep unknown provenance. Room labels resolve against the canonical registry root even on a forked task drive. BIBLE P1 states the principle in three sentences. Version-neutral contribution: release carriers untouched. Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
27 KiB
5. Supervisor Loop
This chapter owns the single scheduler for pooled work: what a healthy tick does, what the queue holds and what its durable snapshot may restore, how a task is addressed and named at admission, how owner waits lend capacity, and the intent-then-custody skeleton every cancellation follows. It exists because these invariants decide whether a stopped task ends honestly or leaves a ghost, and none of them can be reconstructed from any single module's code.
server.py::_run_supervisor() is the single scheduler for pooled tasks. A healthy tick publishes liveness, rotates runtime logs, checks worker health, drains worker and direct-chat events (a wake-up's frames are a direct turn's), accepts owner bridge input, enforces deadlines and schedules, runs throttled reconciliation and evolution admission, assigns eligible work, persists state/queue_snapshot.json, and last ticks the consciousness alarm clock, which owns no thread — this pass is the only thing that can start a wake-up. Bridge intake precedes timeout, maintenance, evolution and assignment work so a slow control-plane step cannot hide a new owner message. Three consecutive loop failures clear supervisor readiness, stop its watchdog generation and notify the owner, rather than leave a healthy-looking server that no longer assigns work; a failure while a shutdown or restart is in progress (the teardown sets a process-local stop event first) is not a crash and never holds the shutdown.
The legacy state.json budget projection reads the validated usage ledger before taking STATE_LOCK, so a multi-megabyte replay cannot hold the short reader lock while waiting on the monetary lock. The same read returns a ledger provenance marker (compaction_epoch, seq): the epoch comes from the lock-free leading baseline header and seq is the live file's validated high-water sequence. A compaction increments the epoch while renumbering live rows, so the pair remains ordered even when the file gets shorter; neither timestamps nor totals prove order. Inside STATE_LOCK, any marker strictly lower than the saved one — a lower epoch, or a lower sequence within the same epoch — is proven stale and leaves the compatibility projection untouched. An equal or higher marker is written. Restore to an older ledger or a malformed saved marker freezes this projection until explicit repair; automatic recovery is absent. Quarantine-file presence likewise keeps integrity_degraded and freezes this compatibility projection indefinitely. This read is display (invariant 28); admission is exact. An unavailable or malformed marker leaves the prior projection untouched; this fail-safe also applies when lock acquisition times out and the writer deliberately proceeds without the lock. On that no-lock path the marker still rejects a demonstrably stale snapshot, but the compare-and-save sequence remains non-atomic for concurrent no-lock writers, as it was before this protection.
PENDING and RUNNING, guarded by supervisor.queue._queue_lock, are the live task-lifecycle authority. Admission reserves identity before project, workspace, attachment or routing side effects can create a duplicate; refuses a disabled pool, duplicate task, project deletion, accepted or sealed root, exhausted root budget, or an unusable or unprovisionable workspace (typed, with the repair in detail); attaches the task contract; and preserves stable priority order. Assignment runs on the same locked state and skips reaping slots, budget-paused work, closed project roots, conflicting project writers, and tasks exceeding the root's subagent capacity or depth reservation. Evolution alone is fenced by runtime mode — blocked in Light — three times (supervisor/evolution_lifecycle.py): entry points refuse a campaign start, enqueue_evolution_task_if_needed() pauses and disables a carried campaign, and assignment drops what slipped through while evolution_block_reason() is set; generic supervisor.queue.enqueue_task() has no runtime-mode predicate. Configured worker count is therefore not available capacity: the truthful value is the assignable idle count after custody, reaping and admission fences.
A headless task is ADDRESSED when it is admitted, not when it is displayed (log_addressing.ingress_chat_id). A registered project's run has exactly ONE destination: an explicit chat_id may only agree with that thread, and any other value — the hidden partition included — is refused with a typed 400 rather than honoured or silently overridden, because a run addressed away from its room puts a card in Main whose project holds none of its work. Without a Project, ordinary API tasks default to HIDDEN_CHAT_ID (0); the confirmed browser Publish flow carries source="web" and WEB_UI_CHAT_ID to request Main (caller-declared addressing, not an authentication proof), and other non-Project conversation addresses stay refused. A run scoped to a REGISTERED, active project is admitted into that project's thread, and Main receives the one host-stamped completion row only when the work is actually in that room: addressed there at admission or BOUND to the project. Registration alone does not qualify. Every other run stays in the hidden partition, silent in every chat, read back through the terminal, --result-json-out, the chat-blind Logs panel and GET /api/tasks/<id>. A reserved but inactive project keeps its chat acceptable so the queue's lifecycle fence refuses with its own typed reason.
The run is also NAMED at admission and chat promotion, without a new model call: a caller-supplied title (ouroboros run --title or the top-level contract field; metadata.title is refused with a 400 like metadata.project_id) is authorship and fills both title and suggested_name; otherwise the request's first line fills suggested_name ALONE, so a truncated prompt never outranks a real name coined later. A task_named frame is broadcast on admission so the live card is never born showing its status phrase as a title.
queue_snapshot.json is an atomic recovery and diagnostic projection, not a second scheduler: pending and running rows, acceptance and root-budget fences, worker counts, assignable capacity, any pool-disabled reason, and the latest bounded root focus. Startup restores a recent snapshot into an empty pending queue and never resurrects ordinary RUNNING work: it FENCES every surviving RUNNING row with a durable cancel intent (reason='server_shutdown', ledgered as terminalized_running), which cancellation custody terminalizes a watchdog window later, expiring its open quiz and closing the paired owner wait; a PENDING child below it is marked pending_parent_interrupted and settled by the boot's kill_workers, so a closed window leaves neither ghost nor orphan. Only an owner-wait handoff with an acknowledged planned-restart transaction outlives snapshot age. Terminal tasks stay terminal, a task with an active cancel intent is left to custody, descendants of an accepted or sealed root finalize as cancelled, and malformed fence evidence fails closed. Assignment mirrors RUNNING into the durable task result for EVERY assigned task, not only a subagent: both orphan healers read the STORED status, and an unmirrored root is a ghost no snapshot-less boot can settle, so orphan reconciliation is a terminal writer for roots too, closing the same quiz and wait the task-done seam closes. Focus updates merge through the existing worker event path under the queue lock; no awareness timer, ledger, or wake is created. direct_roots.json carries the symmetric direct-root projection and one aggregate gap/freshness fact.
At actual execution start, agent._persist_running_record mirrors a split root's running timestamp and execution-drive address into its canonical result through the existing terminal-preserving writer.
Pooled completion separates a finished file-save attempt from publishable terminal truth: the worker prepares its terminal task files (headless.prepare_terminal_task_files) before its buffered task_done, and events_task_done publishes only when headless.terminal_task_files_ready confirms CURRENT — on a split drive the child-bound copyback, with no workspace artifact finalization pending. Legacy or faulted completions go through enqueue_terminal_file_recovery (worker_health prepares and recovers, task_reaper runs the queues): unreadable or missing CURRENT keeps RUNNING ownership and retries on the health cadence — never a false Done, never replayed model work — and only CONFIRMED absence of a terminal source reaches the lifecycle-fault owner, which marks execution infra_failed while an early sticky completed status, the authored answer, review and cost survive.
A required owner wait keeps its task RUNNING and retains the same worker, command queue, live browser and services; the park is durable (completed-tool checkpoint, owner_wait projection) before the snapshot confirms it. The worker lends only its active_capacity, which the assignment tick replenishes; addressed input requests a wake, which waits for active capacity, and neither a failed grant nor maintenance reservations can buy an extra replacement beyond the configured cap. Resumption never replays: same attempt, same start timestamp, continuation authority consumed before dispatch. Waiting spares only the idle rail — stop, deadline and absolute ceiling keep their authority. Ordinary Main/Project roots wait in-process instead (owner_wait.direct_owner_wait): no pooled slot is held, lent or synthesized, and the saved source is evidence, never a cold-restart grant — an addressed text or quiz answer resumes the same live stack and browser.
A crash or automatic timeout retry keeps the current-attempt checkpoint as evidence and never blindly replays the task; cold continuation exists only for a confirmed planned restart, which cannot preserve an OS browser session. An owner-requested MANAGED UPDATE is such a restart: its writer fence prepares the handoff before stopping the pool and passes the parked ids as preserve_running_task_ids, so a wait is requeued instead of interrupted, and a preparation failure blocks the update rather than terminalizing the wait; the re-exec seam arms the transaction only in the phases _safe_restart_serialized allows (launcher mode observes exit code 42), no-resume flags suppress it, and an aborted update's leftover record never authorizes a later manual Restart or rollback. A manual ROLLBACK returns to an older runtime, so it deliberately parks nothing. Cold continuation resumes the saved economics — original CostCeiling, hard clocks, the saved model's ContextFit, TaskModelWait choices, the pending budget decision of the saved round — and calendar deadlines and owner-wait time are unchanged.
Cancellation is intent-then-custody; intent and outcome are separate fields, because one field carrying both wedges a task forever. The skeleton:
- Intent. Every cancel ingress — the agent
cancel_tasktool, the HTTP single and cascade endpoints, evolution stop, project deletion, a cascade sweep's per-descendant mints, the boot migration of legacycancel_requestedfiles — writes one durable row throughouroboros/cancel_intents.request_cancelinto the locked projectionstate/cancel_intents.json(active intents only; every transition also appends a forensiccancel_intentsupervisor-ledger row). Every ingress fails closed — toolCANCEL_INTENT_WRITE_FAILED, HTTP 503,CANCEL_INTENT_PROJECTION_CORRUPTfor a corrupt projection — so no teardown runs without a durable, watchdog-replayable fence; evolution stop keeps any task whose intent write failed and reports the stop INCOMPLETE with typed per-task outcomes — the campaign stays OPEN under the durableevolution_owner_stoppedflag, which only an OWNER start ingress clears (the agent'stoggle_evolutionagainst it is refused: the owner's stop is sticky). - Scope. A cascade ingress mints
scope: cascade; recorded scope is widen-only (single→cascade, never narrowed, so Stop-now cannot shrink a cascade). A cascade over an already-settled root with live descendants still mints the coordination intent (allow_settled_target), the watchdog's replay trigger for the subtree. Timeout reaping is deliberately NOT a cancel ingress: the reaper keeps its own custody over thereapingslot marker, because cancellation is reserved for explicit intent. - Claim.
supervisor.task_lifecycle.cancel_task_custodyis the ONE settle owner: it claims the intent (owner + generation) before any custody mutation, and a refused claim exits having touched nothing, so racing custodies never double-settle; the pre-assignment pending drop holds the same fence, and a budget-exhausted queued task PAUSES, so there is no batch terminalizer beside custody. A LIVE direct-chat turn is stopped through the same custody (supervisor/worker_chat_lane.py): the chat lane writes a typedfinalize_nowcontrol into the owner mailbox once per turn, and the loop ends at its next round boundary with zero further model calls, paid post-task synthesis included. A Stop-now landing after the loop returned, while that synthesis runs, is still addressable: the in-flight synthesis counts as live ownership, custody keeps the immediate intent open, and the synthesis worker checks it before each paid stage, disclosing the skipped ones aspost_task_stop_reasonowner_stopped:skipped=<stages>on adegradedcheckpoint. Custody waitsOUROBOROS_DIRECT_TURN_STOP_WAIT_SEC; the typed outcome isgone/ended/live, andlivereleases the claim for the sweep rather than publishing a fabricatedcancelledrow over a turn that is still running. - Kill and re-check. Custody confirms process death, then re-reads the child's real settled result. Natural completion WINS: a child that finished before the kill keeps its result, artifacts and cost, and the cancel settles as already-settled.
- Reconcile and capture. The task's open delegated runs are reconciled from durable custody rows and always re-audited and disclosed. Workspace artifacts are captured from the real tree; a failed or owed-but-unrunnable capture is
failed, nevermissing, and a shared-tree capture carriesattribution: shared_unproven. - Settle. The settled result carries reconstructed-or-honestly-unknown cost, never a fabricated final
$0;parent_decisionis stamped only at this outcome.cancel_publication._intent_outcome_fieldspreserves recorded cancellation provenance ascancel_origin: source, scope, reason,request_id/requested_at,requested_bywhen present, and the typed observationrequest_origin. These facts survive removal of the active intent and travel through the terminal event, task detail, history and result-tool reads, including conditional reads whose answer body is unchanged. HTTP proves transport, never a personal owner; absent actor evidence remains absent. The existingrequested_bycondition for parent-decision semantics is unchanged. - Owe, then publish. The owner's terminal answer is registered as OWED in the durable outbox (or a typed no-chat handoff row) BEFORE the intent settles and before
task_donepublishes, so a crash between settle and send replays the answer instead of losing it. A cascade delivers one root message with a children digest under the deterministic delivery idcascade:<root_tid>:<request_id>, each child's line rebuilt from its current durable status. - Watchdog. The supervisor tick runs the cancel/delivery/ref sweep off drain (
server_maintenance._run_cancel_delivery_ref_sweep, ~20 s cadence);sweep_cancel_intentsre-feeds unclaimed or abandoned-claim intents into custody — a cascade replayed as a cascade — so a lost control event or a custody attempt that died mid-teardown cannot wedge a cancellation. Only the physical no-live check settles a cascade's coordination intent, after the tree's summary is registered as owed.
Readers see the typed projection cancel_state: "pending" (with cancel_reason) on effective results until the settle; the UI shows "Cancelling…" and restores the Cancel button only when a fetched live non-pending task detail proves the intent is gone. Steering writes — steer_task, mailbox follow-ups, forward_to_worker — are refused typed while a cancel is pending (that fence is what makes the owner-stop single-turn rail safe), and queue restore and pre-assignment consult the projection under the queue lock, so a cancelled pending task never starts. task_done asserts a SETTLED outcome and is validated against the DURABLE result for every event: a non-settled event status, or a settled or blank status over a non-settled or absent durable row, is a lifecycle fault — left to custody when a cancellation is pending, otherwise published with a typed infrastructure-failure axis that preserves an existing sticky terminal status.
Terminal answers ride one durable delivery seam (supervisor/terminal_delivery.py): a bounded PENDING outbox state/terminal_deliveries.json (owed before enqueue, replayed on boot and on the tick; eviction past capacity is the typed terminal_delivery_exhausted, never a silent pop) with restart-surviving delivery_id dedupe shared with the natural final-answer path; a loud UNREVIEWED salvage message (bounded preview, exact omitted count, full-copy receipt) for cancelled and non-retry-reaped tasks; one root message for a cascade; nothing for a retryable reap; routing follows the task's lineage chat. The already-settled and finalize-on-miss paths run the same delegated-run audit as the kill path, so a cancel over a dead task with live delegated runs never reads as a clean completion. An agent-requested cancel publishes nothing of its own — the custody seam's terminal rows and the typed tool result are the truth — except a FAILED settle, which speaks as the typed cancellation_fault progress incident. The custody, completion-wins and owed-before-published invariants are restated in §10.
Stop POLICY is an axis on the same durable intent, independent of cascade scope. An omitted or empty-body cancellation is the synchronous IMMEDIATE teardown, keeping programmatic callers' bounded budgets. An explicit stop_policy=finalize_then_cancel answers 202 with the intent OPEN and runs one bounded owner-stop finalization episode (supervisor/owner_stop.py): live descendants settle first and feed a bounded child-result projection into the root's final turn; the root receives a finalize_now control whose typed first line (owner_requested_finalization) routes to its own loop rail — zero or one tool-less model turn, terminalizing completed/best-effort under the honest owner reason rather than a false deadline reason. The grace budget starts at the durable control_drained_at (first drain wins, so a task inside a long tool call still gets its final turn when the hard bounds allow) under the request-anchored OWNER_STOP_OUTER_CAP_SEC cap; a held task bypasses only the generic idle/finalization-grace rails — its explicit deadline and absolute ceiling remain independent hard axes and are never widened — and expiry, a hard-bound hit, a pending root or an already-settled root feeds ordinary custody. Policy transitions are monotonic: an immediate request HARDENS a pending graceful intent (preserving any cascade scope); graceful can never soften an accepted immediate. A successful graceful root suppresses the redundant cascade summary; Panic bypasses both. The UI projects the soft stop through cancel_state+stop_policy ("Finalizing…").
Beside stopping sits the owner "hurry" control: a typed task-local kind=hurry owner-mailbox control (ouroboros/owner_hurry.py, gateway/task_hurry.py) that skips the next otherwise-eligible acceptance panel with a typed reason, zeroes remaining improvement passes, and makes force-plan projection task-locally advisory — never a chat message, never a settings mutation, never a P3/commit/review-gate weakening. Its effect is attempt-scoped (task["_attempt"]): a shared retry_reset strips it on every same-id requeue. These invariants hold for every install configuration class.
The event bus is process-lifetime rather than worker-generation-lifetime: respawns reuse one manager-backed queue shared by workers and direct chat, because a force-killed producer can corrupt a raw multiprocessing feeder frame and a queue rebuilt on pool rotation strands surviving producers on the old endpoint. Live-frame publication of persisted rows is exactly-once and process-symmetric: ouroboros/utils.py::append_jsonl streams only runtime logs/*.jsonl rows into the process log sink (never chat.jsonl, never state/memory/receipt stores), and each process suppresses the types whose live delivery has a dedicated owner (WORKER_LOG_SINK_SUPPRESSED_TYPES, the server superset SERVER_LOG_SINK_SUPPRESSED_TYPES). One persisted event produces exactly one live frame (tests/test_log_forwarding.py); an LLM call failure is one durable llm_api_error row and nothing else.
Heartbeat and progress are different evidence: a heartbeat proves a process or loop is alive; owner-visible progress and model-usage events prove the task advanced. Fresh descendant progress or queued descendants can keep an orchestrator alive, while an explicit deadline, absolute ceiling, cancellation and budget stop remain hard. After the typed finalization episode (§6), timeout handling freezes its decision under the queue lock, marks the worker reaping, and hands kill, join, salvage, retry and respawn to the single off-loop reaper; an orchestrator with live descendants is not blindly retried, because a retry would replay its plan and spawn a competing tree. No retry or new assignment may occupy a timed-out slot until the original process is provably dead: if kill and join cannot establish death, the reaper keeps a low-rank RUNNING result and the reaping slot, emits a visible task_reaper_wedged receipt and restart hint, and writes no terminal, task_done, retry or respawn — one slot is sacrificed rather than letting a still-running process race a replacement and overwrite its result; the next supervisor generation reconciles the record after old-generation process custody.
A spawned or respawned slot is not assignable until its child's PID-bound worker_ready row arrives (supervisor/worker_pool_lifecycle.py). A live child's own worker_starting row, emitted before extension loading and agent construction, permits one extension of WORKER_READY_WINDOW_SEC to WORKER_READY_CEILING_SEC (300 seconds from birth, both in runtime_limits.py); foreign or pre-spawn rows cannot extend another slot. worker_ready_window_extended records that decision. A silent child keeps the original window, and logging failure cannot block startup. After WORKER_READY_MAX_ATTEMPTS failed attempts, Worker.readiness_exhausted is final for that exact slot — late events cannot reopen it. Total exhaustion, distinguished from busy/booting/reaping capacity and from a live owner-wait stack, closes pooled ingress (owner /review included) without blocking direct chat/control or boot/update recovery; once RUNNING completion custody has settled, disable_exhausted_worker_pool fails unstarted PENDING work honestly with a Restart hint, and a new task cannot clear the latch. Readiness stays separate from liveness and task idle time; a watcher error releases only still-booting, non-exhausted slots to the crash detector (worker_ready_released). Linux workers use forkserver; macOS and Windows use spawn.
Unexpected worker death reserves exact custody under the queue lock and enqueues confirmed_dead_worker on the reaper (worker_health.recover_confirmed_dead_worker). A saved terminal source wins even after signal death; unknown or incomplete file publication keeps the same job (TerminalFileRecoveryPending); only confirmed absence of one reaches the crash policy: a signal is an infrastructure failure, an otherwise eligible non-signal crash retries within QUEUE_MAX_RETRIES, preserving owner-wait replay restrictions and cost. A crash storm suppresses respawn while terminal sources settle, then its fence stops pooled admission; direct chat stays available. Startup runs the same terminal-file recovery in _run_supervisor after process custody and before _startup_prune_sweeps (the no-provider lifespan runs it too, spawning nothing); unknown or still-live ownership defers it rather than racing a writer, and any unresolved or protected source, or an ownership/read error, sets preserve_task_sources, skipping task-drive deletion for that pass. For older canonical scheduled rows, _recover_terminal_task_files restores that start binding only from a known non-direct child's positive running/started-at record when the existing fresh-queue and later-worker-boot checks prove it orphaned, with no pending queue owner or active cancel; the normal orphan reconciler and terminal guards retain authority, without resuming work. The recovery report includes rebound.
Startup and throttled maintenance reconcile three residue classes, the ~600 s pass off the loop thread (§10). Process custody checks strict PID, start-time, command, owner-task, session and generation evidence before it reaps. Delegated-run reconciliation applies the same owner-gone reasoning to harness rows (§6 Delegated subagents). Task, review and project reconciliation repair records whose producer no longer exists. None of these are command-line-class kill sweeps, and one instance never reaps another. The dedicated watchdog separately observes phase-stamped loop liveness and every native actor; a wedged chat turn alerts with a /restart hint, a loop stall only journals, and neither kills a thread. Other owner conversations run on independent native actors, without a second scheduler.
Cooperative project checkpointing has two equivalent quiescence triggers: a host-minted genesis or cooperative tree is checked when its root settles with no live descendants, and again when the last child settles beneath an already-terminal root — a root-scope budget stop terminalizes the root before its children, so a root-only trigger would see a live tree once and never return. The bounded git chain runs on a daemon thread, revalidates quiescence under the queue lock immediately before mutation, and replays a trigger that arrives during an in-flight check. Only host-minted project roots are eligible: owner-attached folders are never auto-committed, credential-shaped files stay excluded and disclosed, and every material success, skip or error receives a durable receipt.
The bridge recognizes /panic, /restart, /review, /evolve [on|off], /bg [start|stop|status] and /status; all other text enters ordinary agent routing. External transports may invoke these commands only with positive owner identity and a transport-specific owner-chat binding, and the commands reuse runtime-mode, queue, cancellation and typed-result authority rather than implementing parallel control paths. Runtime logs rotate on the same supervisor tick and archive readers preserve their retained timelines; only explicitly isolated devtool roots may use the narrow rotation sentinel from §1.