ouroboros/docs/ARCHITECTURE.md
Ouroboros bc7037bd71 P7-1: fence surviving RUNNING snapshot rows with a durable cancel intent
Restore read only snap['pending'], so a window closed on live work left both
roots stored as their pre-restart status forever: the startup re-persist then
overwrote the snapshot and the one owner line named the task that was NOT lost.

Restore is the last holder of the pre-restart running list, but it must not be
a second terminal writer racing a worker that outlived SIGTERM. It now walks
snap['running'] before the stale early return and mints one durable cancel
intent per surviving row (reason='server_shutdown', source='snapshot_restore'),
skipping rows that are terminal, cancel-requested, already owned by an active
intent, or about to be revived as pending; an unreadable cancel authority mints
nothing and is disclosed instead. The existing custody path then claims, kills,
reconciles and writes the terminal with its own text, and the task-done seam
expires the open quiz and closes the paired owner wait (owner Q11=A).

The fenced ids are counted as terminalized_running in the existing restore
ledger row and returned through a keyword-only out-parameter, so the int return
and its asserts are unchanged. The boot notice fires on restored OR fenced work
and states an INTENT: custody writes each terminal result at least a watchdog
window (10s) later. The live startup order (restore, then kill_workers) is
unchanged and the minting is idempotent across boots.
2026-09-12 17:14:46 +03:00

647 KiB
Raw Blame History

Ouroboros v7.0.0 — Architecture & Reference

This file is NOT a changelog. Version history lives in README.md, git tags, and commit log.

This is the present-tense operational map of Ouroboros (BIBLE P6), in three layers: structure (what exists and where), operation (files, env keys, state paths, endpoints, flows), and rationale. Every important WHY stays here at least briefly; mechanism detail lives in the module docstring the map points to by name. Rationale must be self-contained — future maintainers should not need old commits to understand why a guard, review gate, or lifecycle exists.


1. High-Level Architecture

User
  │
  ▼
launcher.py (PyWebView)       ← desktop window; immutable release-reviewed outer shell running the packaged copy outside managed hot-swap
  │
  │  spawns subprocess
  ▼
server.py (Starlette+uvicorn) ← HTTP + WebSocket on configurable host:port (default localhost:8765; Docker/non-loopback via OUROBOROS_SERVER_HOST=0.0.0.0); its lifespan APPLIES the boot provider normalization in-process and persists no provider decision — the route is re-derived by every consumer (the task-start projection, the settings GET and onboarding reads, the context-fit route resolver). Startup is a read, with one exception that lives inside the read seam: the lifespan's `load_settings()` runs `context_mode_compat.normalize_and_persist_context_mode_compat`, which rewrites the `OUROBOROS_CONTEXT_MODE`/`OUROBOROS_CONTEXT_MODE_AUTO_LOW` compat pair left by the RETIRED persistent auto-Low mechanism — and only that pair — when it actually changed and the settings lock is held. What retired is that mechanism, not the keys: `OUROBOROS_CONTEXT_MODE` is the live owner-selected context horizon in §7's settings table, and `OUROBOROS_CONTEXT_MODE_AUTO_LOW` survives only as its provenance tombstone
  │
  ├── web/                     ← Web UI (SPA with ES modules in web/modules/; §3)
  │   ├── ui.css + modules/ui_primitives.js ← One shared palette/control stylesheet for the SPA, onboarding and optional author pages; self-contained safe fields, collection, escaping and tone/status functions, re-exported by existing helpers
  │   ├── modules/page_header.js, ui_interactions.js, scroll_fade.js ← Header/tab markup and selected-state keyboard binder; dialog focus, menu behavior and independent popup geometry; scroll-edge decoration tied to actual overflow, all with owned teardown
  │   ├── modules/chat_decision.js, chat_render_batch.js, task_phase_chip.js, lifecycle_card.js ← Chat helpers: the typed decision (quiz) cards shown to the owner; pure render-batch helpers behind the `rebuildAll` replay; pure desired-phase chip projection where terminal truth wins; skill lifecycle card state with best-effort polling
  │   ├── modules/project_work_pointer.js ← Project-only navigation to the existing loaded root cards; pure target selection plus one button/coverage binding, with no execution or history authority
  │   ├── modules/model_wait.js ← Model-wait views inside existing chat cards: revision/attempt-aware snapshots and owner actions through the shared decision ingress; the host retains wait and execution ownership
  │   ├── modules/dashboard.js, logs.js, costs.js, files.js ← Dashboard tab-strip host with static guard markers; Logs page (backfill plus live-stream duplicate guard); Costs page (breakdown buckets — an open zero is never shown as free); Files page (file browser over `/api/files/*`, downloads through the host bridge)
  │   ├── modules/skills.js, marketplace.js, skill_review_card.js, skill_publish_flow.js ← Installed skills UI (review, grant, enable, repair, update, uninstall, delete); ClawHub marketplace inside the Skills page; Skill Review chat cards; typed detail rows for the publish dialog
  │   ├── modules/settings_ui.js, settings_catalog.js, settings_controls.js, settings_local_model.js, mcp_settings.js ← Settings page in reading order (Accounts → Secrets → Models → Agents); model-catalog refresh with a 25-second bound and a sequence guard; effort-segment and form-control binders; the local-model form; MCP settings cards that keep masked tokens until edited
  │   ├── modules/model_roles.js, model_chooser.js ← Shared Models editor for Settings and onboarding: per-role source, model, account and context drafts, plus ordered fallback rows; the editable chooser is shared with actor/reviewer route editors, and catalog arrival enriches suggestions without assigning a value or replacing the input
  │   ├── modules/subagents_settings.js, subagent_status_primitives.js, reviewer_slots.js, route_editor_primitives.js, harness_accounts.js, harness_login_cards.js, claudexor_status_store.js ← Agents surfaces: the Available subagents editor shared by Settings and first-run onboarding; pure status/meta projection of one subagent card; Review lanes rows; neutral route-editor primitives shared by both editors; Agent accounts; host-neutral agent login cards (controller + view); the ONE client-side store over `GET /api/claudexor/status` (`facetReadState`)
  │   ├── modules/onboarding_agents_step.js, onboarding_overlay.js, project_create.js, utils.js ← the first-run "Connect your accounts" step; the framed wizard's sandbox policy kept in one place because it is a security boundary; the New Project dialog and project row actions; shared frontend utilities (escaping, formatting)
  │   ├── modules/review_presentation.js, review_dom_patch.js, harness_presentation.js ← Review Checkpoint grouping/status, keyed DOM reconciliation, neutral harness identity presentation; read-side only; `executorIdentityMarkup` renders projected card identity/model facts at that existing presentation owner
  │   └── modules/widgets.js + widget_module.js + widget_frame.js + widget_card.js + widget_reorder.js + widget_chart.js + widget_list.js + masonry.js ← Widgets page host (`mountTab` dispatcher, card registry, declarative renderer) + the two framed mounts (extension-route iframe; module `srcdoc` iframe with its CSP/sandbox constants, parent fetch/resize bridge and disposer) + the child-side bootstrap served into the module frame (bridge grammar, `Response` rebuilt over a stream, resize reports, dispose acknowledgement) + the framed card chrome (launch policy incl. `retain`, Start/Stop, policy menu, facade) + key-order reorder handles + declarative chart/table helpers + pure list helpers (per-card and order-independent change signatures, keyed patch plan) + the key-ordered masonry that writes only `--masonry-*` custom properties
  │
  ├── supervisor/              ← Background thread inside server.py
  │   ├── active_activity.py   ← Process-local owner of in-flight native chat actors, including preparation and post-task delivery (`DirectActivityRegistry`); private actor handles support controls and writer drain, public snapshots feed `/api/state` `active_direct_turns` and WS typing frames (`activity_id`, `client_message_id`, `phase`, `kind`); no queue records
  │   ├── message_bus.py       ← Queue-based local message bus (Web UI + reviewed transport skills)
  │   ├── workers.py           ← Multiprocessing worker pool (forkserver on Linux, spawn on macOS/Windows; never fork from the multi-threaded supervisor)
  │   ├── worker_assignment.py, worker_chat_lane.py, worker_health.py, worker_pool_lifecycle.py, worker_process.py, worker_promotion.py ← The pool's leaves: handing a pending task to a free worker and refusing the ones that must not run; the direct and ephemeral chat lanes and their resume after a restart (conversation stays admitted while the authorized assisted resolver holds the repository, and the owner-control path is preloaded before conflict markers land); health-owned crash detection and terminal-file recovery through the reaper queue; lifecycle-owned execution admission and pool population (readiness, exhausted-slot disablement, pid records, reaping, respawn) and `kill_worker_tree`, the ONE worker tree-kill every teardown and backstop uses — the installation's daemon roots spared always, kept services only on one task's cancel or timeout; what runs INSIDE a worker child process from entry to crash record; and turning a chat turn — or a project scope — into a queued task
  │   ├── worker_owner_wait.py ← Queue-owned active-capacity transfer for required owner waiting: the original task stays RUNNING and its process stays custodied, while another worker receives the active slot; correlated grant consumes the checkpoint before another model round
  │   ├── state.py             ← Persistent state (state/state.json) with file locking
  │   ├── queue.py             ← Task queue (PENDING/RUNNING lists) + activity-based timeout enforcement; the ONE task-state authority — the lifecycle/publication/transition modules below extend it without becoming second authorities; exact-attempt main-LLM in-flight state spares only the idle rail
  │   ├── queue_schedules.py, queue_snapshot.py, queue_timeouts.py ← Queue leaves re-exported through `supervisor.queue`: recurring schedules (the durable file, the skill sync, what they enqueue); the durable queue snapshot a restart finds and what it may restore; and activity-based liveness — which running task has stopped being alive
  │   ├── cognitive_operations.py ← Typed in-memory LLM/review/VLM operation leases for the idle rail; no scheduler or durable timing ledger
  │   ├── task_model_wait.py   ← Live model-wait event projection and forwarding, with task-attempt/owner checks and quota-clock reads for queue liveness
  │   ├── task_admission.py    ← Token-owned admission reservations fence duplicate user-ingress ids before Project/workspace/attachment side effects; schedule-dispatch refusals project through the same boundary; queue.py stays the state authority
  │   ├── task_lifecycle.py    ← Cancellation custody — the ONE settle owner of durable cancel intents: claim → capture → confirmed death → natural-completion re-check → owed delivery registration → settle → delivery/cleanup, plus the `sweep_cancel_intents` watchdog and the queue-owned root-budget admission fence; every custody rule it enforces is stated once in §10 (cancellation custody)
  │   ├── cancel_publication.py ← Cancellation settlement publication, re-imported by `task_lifecycle.py`: typed CANCEL_* outcome vocabulary, artifact-honest cancelled result fields, physical-ledger cost reconstruction, salvage adapter, owed-before-settle outbox registration, publication of the STORED terminal truth, capture-miss terminalization/delivery adapter
  │   ├── queue_transitions.py ← Queue-owned lifecycle transitions that are not cancellation custody: acceptance-fence open/inspect/seal, explicit budget resume, typed `stop_evolution_tasks` (per-task typed outcomes through the durable-intent ingress, never an in-place prune; an incomplete stop leaves the campaign OPEN via the durable `evolution_owner_stopped` flag, the settle-time backstop in events.py closes it when the last live evolution task settles, and both start ingresses clear the flag BEFORE minting a fresh campaign), and fenced Project deletion (cascade only lineage ROOTS — descendants fall with their trees, one cascade and one summary per tree; tombstone only after provable quiescence; a settled-but-LIVE root still mints the coordination intent, and wind-down defers and RE-CHECKS bounded instead of re-running the cancel pass over a settled-lingering set, which would deliver duplicate owner summaries); imports nothing from task_lifecycle; `supervisor.queue` re-exports these names
  │   ├── terminal_delivery.py ← Durable terminal-answer delivery seam: restart-surviving `delivery_id` dedupe + bounded PENDING outbox `state/terminal_deliveries.json` (owed before enqueue, cleared in the delivering write, replayed on boot and on the supervisor tick), shared by natural final answers (every non-ephemeral root registers at durable-result persistence), cancel salvage, cascade digest, and non-retry reap; the cascade digest enumerates descendants by ANCESTRY (parent-chain walk, never `root_task_id` equality); eviction past outbox capacity is disclosed via typed `terminal_delivery_exhausted`, never a silent pop; salvage messages carry a bounded preview plus a full-copy receipt (path, size, full 64-hex sha256 or an explicit marker) and route by lineage chat — no resolvable chat records a typed `terminal_delivery_handoff` row; reads/mutations are row-strict per §10; the delivery id digests only the STABLE part (task id + status framing + core answer) so a replay whose rebuilt note shrank dedups instead of double-sending; the per-origin projection is `host_salvage` receipt / `host_notice` own text kept as a System row with its markdown / `model_final` assistant projection
  │   ├── task_reaper.py       ← Single-owner off-loop queue/pump for timeout teardown and health-prepared terminal-file/crash jobs, with same-job deferred replay bound to the captured worker/attempt/root, so old file recovery cannot replace a newer execution; keeps supervisor intake responsive. An unconfirmed death holds the slot reaping and the task RUNNING with task_reaper_wedged; confirmed death precedes delegated-custody reconciliation and retry. No cancel intents minted (§5 Supervisor Loop).
  │   ├── owner_stop.py        ← Owner graceful stop: `finalize_then_cancel` policy as an axis on the SAME durable cancel intent (monotonic — immediate HARDENS a pending graceful, never softens back; hardening revokes an unread control via mailbox revocation, and the loop revalidates durable policy at drain); one deterministic typed `finalize_now` control whose first line is the `owner_requested_finalization` literal, routed by the loop to its own rail (zero or one tool-less turn); descendants settle first with a bounded child projection fed to the root's final turn; the grace budget starts at the durable control DRAIN (first drain wins), bounded by `request + OWNER_STOP_OUTER_CAP_SEC`, and neither anchor is ever progress-extended; `running_owner_stop_tasks` bypasses only the generic idle/finalization-grace rails; a COMPLETED finalize root suppresses the redundant cascade summary
  │   ├── schedule_time.py     ← Cron/timezone schedule time parsing helpers
  │   ├── evolution_lifecycle.py ← Evolution campaign state + transaction lifecycle: campaign file IO, start/pause, begin/update transaction, cycle-outcome recording, deterministic worktree cleanup, owner cycle reports, idle dispatch over queue-owned state, supervisor auto-restart request
  │   ├── events.py            ← Worker→supervisor event dispatcher with exact attempt/execution/round/call correlation for the process-local active main-LLM row; composes the frozen subagent task text (`_compose_subagent_text`; the acting `[WRITE SURFACE]` block states only the write-root authority boundary — actor identity comes from the immutable configured snapshot, startup/wake facts from bootstrap); an event type absent from `EVENT_HANDLERS` is DROPPED into a truncated `unknown_worker_event` row, and `tests/test_worker_event_registry.py` pins the registry by AST scan — an unexplained allowlist entry is exactly the silent blessing the scan exists to end; the scan is shape-bounded, and outside its reach the discipline is code review; holds shrink-only byte debt above the module byte ceiling
  │   ├── subagent_task_truth.py ← Delegation-truth enrichment of the subagent `task_done` transport frame (`enrich_task_done_event`)
  │   ├── event_taxonomy.py, events_budget.py, events_chat_delivery.py, events_coop_checkpoint.py, events_evolution_done.py, events_project_routing.py, events_runtime_controls.py, events_schedule_task.py, events_subagent_admission.py, events_task_done.py, events_worker_reports.py ← The handler leaves `events.py` merges into `EVENT_HANDLERS`, one owner per family: the declared disposition of every event kind the runtime puts on `EVENT_Q` (`event_taxonomy.py`, the registry the AST scan reads); usage-accounting and budget-pause reports; owner-facing chat delivery (text, media, typing); cooperative repository checkpoints when a task tree goes quiescent; terminal handling of an evolution task and its campaign; where a chat turn becomes a task and where a project scope is bound; posture changes that are not one task's state; the `schedule_task` admission and duplicate gates and their refusals; admission facts for a requested subagent (census, caps, constraint); resolution of a terminal event into durable truth and delivery; and what a running worker reports about itself
  │   ├── task_dispatch.py     ← Pure admitted-event→worker-payload construction, including the identical top-level/metadata depth and configured-route projections consumed by workers
  │   ├── log_addressing.py    ← Explicit audience for task-scoped live log events: `address_task_event` (lineage from the RUNNING row, project binding wins, explicit chat_id preserved — 0 is `HIDDEN_CHAT_ID`, the hidden partition of the Skill Review panel and headless runs without a registered project; A2A frames are suppressed at the `push_log` choke, not by dishonest addressing), `make_server_log_sink`, `address_handler_push`
  │   ├── steering.py          ← Owner steering-message delivery to running tasks: mailbox routing to the drive the worker drains, plus a typed refusal while a cancel intent is pending (steering is fenced during a stop — what makes the owner-stop single-turn rail safe)
  │   ├── telemetry_events.py  ← Durable handlers for RARE typed telemetry-only worker events (merged into `EVENT_HANDLERS` like `_CEH`): a type-agnostic passthrough appends each row verbatim to events.jsonl, beside the `task_message_injected` sibling shaping the A2A row; membership is bounded by contract — every dispatch is a durable append, so high-rate narration never joins this registry
  │   ├── git_ops.py           ← Git operations (clone, checkout, rescue, rollback, push, credential helper) and the shared bounded local-Git process runner
  │   ├── git_ops_remotes.py, git_ops_rescue.py, git_ops_reset.py, git_ops_updates.py ← Git-operation leaves re-exported through `git_ops.py`: personal persistence remote (`origin`) configuration and push; the rescue/snapshot machinery destructive tree movement takes first; checkout/reset admission, dependency sync and safe restart; and managed-update status, official tags and update preparation
  │   ├── update_source.py     ← Official update-source selection + network policy through the bounded Git runner
  │   ├── update_recovery.py   ← Exact owner Restore/promotion: pinned prior HEAD, rescue-before-reset, one captured SHA for local/remote promotion
  │   ├── update_merge.py      ← Managed-update engine: exact-target 3-way plan, direct clean fast-forward, stash-first reviewed assisted merge (both lanes stash dirty work before the authoritative replan; VERSION + carrier tokens projected before the M0 pin), transaction (`m0_tree`, `tests_evidence`, `stash_sha`/`local_work_carrier`, `failed_update_ref`), verified rollback/smoke, phase-dispatched boot recovery (the marker-cleanup phase only retries the tx-marker unlink when the repository already holds its final state); disclosed residual: M0 is a pin-once forensic baseline in the resolver-writable tx marker — review discloses it and does not re-verify
  │   ├── update_candidate.py  ← Candidate/carrier primitives (re-exported by update_merge.py): private-index tree serialization, rerere-neutral merges (prevent silent rr-cache resolution replay), failed-update preservation branches and marker-guarded stash restore; publishes forensic tests_evidence from the runner's process-held proof, never durable-file reuse authority (§6 Git and commit review)
  │   ├── update_carriers.py, update_merge_plan.py ← Carrier-aware managed-update conflict resolution (the version/contract carriers a merge may not resolve by hand), and merge planning with live materialization
  │   └── update_merge_policy.py ← Presentation-only doc/code/hot conflict labels; every conflict uses the same reviewed assisted path
  │
  └── ouroboros/               ← Agent core (runs inside worker processes)
      ├── config.py            ← SSOT: paths, settings defaults, load/save, PID lock (§7)
      ├── settings_defaults.py, settings_scales.py, model_slots.py, review_model_routes.py, runtime_limits.py ← The settings vocabularies `config.py` re-exports, one owner each: the key set with its shipped defaults plus `RETIRED_SETTING_KEYS`/`RETIRED_COMMA_LIST_SETTING_KEYS` (`settings_defaults.py`); the closed scales and setting-effect vocabulary (`IMMEDIATE_SETTINGS` / `RESTART_REQUIRED_SETTINGS`); model-slot resolution and the frozen `ResolvedModelTarget`; the reviewer model lists the API-pinned review surfaces run; and the numeric runtime knobs with their clamps. `ouroboros.config` stays the one import surface — a new key and default belong to the leaf, never to the facade
      ├── version.py           ← Version string from the VERSION file with importlib.metadata fallback
      ├── secret_masking.py    ← Exact Settings/MCP wire-placeholder emitters/recognizers + top-level secret repair before env overlay and persistence
      ├── settings_integrity.py ← Task-local in-memory ordinary-settings read view and strict settings-snapshot integrity pin; `OUROBOROS_SETTINGS_SHA256` enables the trust root
      ├── credential_shapes.py ← Credential leaf shapes plus physical owner credential locations; ordinary root reads never import shape policy
      ├── update_channels.py   ← Closed Stable/QA/Development channel mapping and update-network defaults
      ├── update_letter.py     ← The update letter: first-parent range material (EVERY commit subject of the range with its full sha + the README history rows their diffs added, each naming its commit; only bodies and the oldest row texts are bounded, and disclosed; the paragraph's shape is asked of the model, never policed by the host), one accounted LIGHT-slot call with the ordinary task context, `state/update_letter.json`, one projection shared by the Updates payload and the Runtime context `official_update` fact
      ├── colab_bootstrap.py   ← Google Colab source-mode bootstrap: official update source, stable local `ouroboros` branch, Drive-backed settings/data, personal origin, no-UI server command, native Telegram setup
      ├── cli.py               ← Source/headless CLI over gateway tasks, logs, settings, skills, marketplace, local-model, and MCP wrappers
      ├── packaged_cli.py      ← Packaged desktop CLI bridge: resolves bundle roots, bootstraps the launcher-managed repo, delegates to cli.py
      ├── packaged_cli_install.py ← Packaged CLI installer planning/execution for user-local command shims
      ├── agent.py             ← Task orchestrator; the dispatch-note pair lives in `subagent_dispatch_notes.py` (same-name re-exports). The outer catch around `run_llm_loop` delegates terminal projection to `_task_exception_terminal`: `_LoopExitContext.attach_exception_evidence` retains the loop's own accumulated objects on the original exception, while missing captures are explicitly unknown. A failed cold-source read never supplies fabricated zero counters or unverified checkpoint bytes; the trace summary, loop usage, task metrics, ephemeral chat/history counters and post-task summary preserve that absence. The loop's tally rides `loop_outcome.usage` (ABI-3's honest loop plane); the top-level `total_rounds`/`prompt_tokens`/`completion_tokens` stay the LEDGER's answer from `reconstruct_task_cost`. An internal `task_exception` is `failure.kind = "runtime"`, never a fabricated provider failure
      ├── agent_startup_checks.py ← Worker-boot verification: dirty repo, version sync, budget, memory files, health checks
      ├── agent_task_pipeline.py ← Task execution pipeline orchestration; freezes one shared non-final subtree-cost snapshot for summary/reflection before the terminal checkpoint records final spend; hands the summary and reflection prompts the commit/advisory review lens PLUS the task's own acceptance-panel projection, and an absence statement names the lens it describes; calls the swarm-efficiency rollup owned by task_finalization.py at pipeline end
      ├── agent_dispatch.py, post_task_synthesis.py ← The agent's delegated-child dispatch seam, and the post-task synthesis workers the pipeline runs after a terminal result
      ├── task_finalization.py ← Terminal delivery + sealed final ground truth: live final-answer delivery before blocking post-task (final event selected by the finalizing task's id; buffered copy retained under one `delivery_id`), the sealed final package (submitted final text, the durable result's artifact manifest, and task-related completion observations) fed to summary/reflection as a prompt input, never a validator; owns the per-task `swarm_efficiency` rollup (subagent_count / fanout_count / fanout_interval_sec_total / `lanes_requested` — a rollup built from pre-dispatch fanout events cannot truthfully report effective lanes, which are per-child dispatch facts; `planned` stays null, never inferred as 0 from absent events; host-attested Swarm intent is the typed metadata `force_plan_source == "swarm"`, never prompt inspection, and a Swarm task that fanned out nothing records a minimal `no_fanout_observed` block instead of silence); a fanned-out root additionally carries a `depth` block (`requested_depth`/`permitted_depth`/`attempted_depth`/`achieved_depth` plus a typed status, `host_visible_only`) built from the root contract and its subtree's depth provenance, so a root that never carried its own request recovers it from the children it scheduled, and the terminal task_summary row reports it only when a request exists
      ├── mutation_attribution.py ← Root-task baseline capture in the existing task result; clean-at-baseline Git candidate projection; terminal projection includes the committed interval delta
      ├── process_interpreters.py ← Interpreter resolvers for the user process launch surfaces: one-time pre-guard unversioned-Python resolver + post-gates Node ladder (PATH-first health probe, bundled fallback, attested child-env PATH prepend; the probe EXECUTES a candidate, so it runs only after the dispatch gates approve the call)
      ├── post_task_checkpoint.py ← Durable root post-task phase/final-cost checkpoint shared by task finalization and Project naming recovery
      ├── presence_profile.py  ← Strict reviewed `presence:` behavior-profile parser (instructions, context topics, runtime defaults, portable capability requests)
      ├── presence_runtime.py  ← Symbolic `main`/`light` defaults; owner-local overrides clamped to the global round limit
      ├── presence_capabilities.py ← Host-owned installation selections for portable Presence capability requests
      ├── presence_authority.py ← Immutable positive capability ceiling compiled from one reviewed profile and its selected exact targets
      ├── presence_bindings.py ← Owner-created revocable transport-room → behavior-skill bindings with exact origin/destination identity
      ├── presence_admission.py ← Fresh review/enablement/profile/state admission + immutable per-turn Presence snapshot
      ├── presence_context.py  ← Presence instructions, exact event facts, declared knowledge-topic projection
      ├── presence_runner.py   ← Fresh-agent Presence turns: cross-process installation cap, per-conversation serialization, idempotency, typed result, dialogue provenance
      ├── dialogue_provenance.py ← Shared exact transport-provenance rendering for history, memory, and consolidation
      ├── extension_companion.py ← Host-supervised companion processes for transport skills (§12)
      ├── extension_reconcile_queue.py ← Durable worker→server extension reconcile markers + server pickup loop
      ├── event_bus.py         ← Typed in-process event bus for skill subscriptions
      ├── evolution_checkpoints.py ← Append-only campaign/eval checkpoint ledger for evolution progress
      ├── evolution_fingerprint.py ← Canonical fingerprint for evolution-campaign objectives; SSOT for repeat gating
      ├── improvement_backlog.py ← Durable advisory improvement backlog: recurrence-counted dedup (bump count/last_seen, never drop), priority+recurrence+recency ranking, close-on-commit `close_backlog_items`, size-triggered `groom_backlog`; parser-safe locked writer; entries carry priority/kind
      ├── loop.py              ← High-level LLM tool loop and its one-shot finalization nudges, ordered nanny → red-verification → masked-verification → no-op-attempt; continuous `FINAL ANSWER:` latching captures the latest typed candidate every round (tool-count-stamped, no prose mining) so review/nudge/forced-finalization paths never erase a structured answer; all marker prompting is gated on `task_contract.answer_protocol="final_answer_line"` via the `answer_protocol_active` SSOT (the gate is sufficient — an empty `expected_output` cannot suppress it) while the latch/extractor stay unconditional; `outcomes.extract_final_answer` (re-imported here) structurally rejects the outcome-tier ledger identifiers as answers — internal enum vocabulary is never a deliverable, and `solved` stays extractable as an ordinary English word. The nanny nudge (`loop_nudges._nanny_finalization_message`, re-exported here) fires once for a harness child finalizing with ZERO durable start attempts (blocked and uncustodied attempts count, so an exact-route startup fault is not accused of skipping delegation); it reads custody evidence from the canonical (budget) root via `delegate_custody.custody_root` and branches PENDING ≠ FAILED — a started-but-unsettled run gets a wait reminder, never a failure accusation, which would invite a duplicate concurrent run; an actually-injected nudge is stamped by the WORKER as a durable custody row, and a COMPLETED harness child with zero runs carries the typed `nanny_finalized_after_nudge_without_delegation` disclosure (stamped by `subagents._disclose_native_only_substrate` at the completion seam; visibility, never a gate); a configured session child finalizing without a succeeded leaf run carries the typed `CONFIGURED_ACTOR_INCOMPLETE`/`CONFIGURED_ACTOR_UNKNOWN` fact (`subagent_bootstrap.actor_first_unresolved_fact`; host children ride along as auxiliary `direct_child_statuses`, never a substitute for the leaf), and a successful run's later silence stays proportional to the measured burn (the `NANNY_METERED_OVERRUN` reminder in `loop_nudges`, metered by `nanny_pacing.nanny_metered_since_delegate_activity`/`nanny_burn_phrase`)
      ├── loop_acceptance.py, loop_acceptance_review.py ← Acceptance machinery (moved whole out of `loop.py`, which keeps the checkpoint, run-record and message rails; every name with an external caller is re-exported from `loop` so external callers, patch sites and the acceptance-writer inventory keep one import surface — a moved name nobody outside references is not re-exported). The fence and its obligations: eligibility, begin/end/supersede, subtree snapshots, the final-answer latch, the closed typed reason set `ACCEPTANCE_DECISION_REASONS`, the sole decision merge point `_set_acceptance_decision`, and obligation collection/reopen/disposition. The run: the host evidence packet (`_build_host_acceptance_evidence` over `review_evidence.build_task_acceptance_evidence`, plus the forced-rail child debt and the unhashed dialogue history), the one substantive panel (`_execute_task_acceptance_panel`: reviewer rows, wave-budget admission, the free zero-physical refusal, the exact-hash wallet stamp, the timing event), the bounded `acceptance_dialogue_history`, and the paid identity + free-replay `_refuse_identical_acceptance`; the `dialogue_status` reducer is `review_verdict.aggregate_dialogue_status` (the vote SSOT, re-exported through `review_substrate`; these are its consumers). `loop_acceptance_review.acceptance_retrieving_work_order` renders the delivery-conditional work order of the retrieving rows (session: FULL packet + absolute pointers + access disclosure; native: packet without its freely degradable tail + real data root) onto `ReviewRequest.slot_session_tasks` — the FULL packet stays the `evidence_refs` authority
      ├── loop_llm_call.py     ← Single-round LLM call + usage accounting
      ├── loop_transport.py    ← Transport-outage wait episodes and provider-failure terminal text; bounded backoff with free redial
      ├── loop_delivery.py ← Delivery candidates and delivery control: the candidate dataclass with inherited host-control-episode provenance, the hold-control literals, the control prompt and its cycle, whole-body classification and the duplicate-key/trailing-object parsers, child-result dispositions, the delivery evidence state, acceptance bindings, candidate publish/replace/degrade, the subagent handoff and the no-tool final; `loop` re-exports the underscore names
      ├── loop_budget.py, loop_forced_finalization.py, loop_messages.py, loop_model_call.py, loop_nudges.py, loop_round_limits.py ← The rest of the round driver, one rail per leaf, all re-exported from `loop`: the budget rails (the shared warm/cold post-tool decision, cost ceilings and tree accounting, soft landing, budget-exceeded handler, resource cleanup, service finalization); forced finalization of a task that ran out of road (orphan notes, child claims and the absorption gate, swarm-action enforcement, owner-directive drain, the one forced model call); owner-message text plumbing (extraction, append-or-merge of user turns, stale-image eviction, owner-directive bookkeeping, round-progress text); the per-round model call (context-fit identification, measurement and memory, dispatch, main-context reclaim, overflow-retry predicates, the cross-model fallback chain); the mid-task steering notes (self-check, time and cost milestones, nanny economics, round checkpoints, plan-forcing and finalization nudges); and the round-limit and terminal-drain handling (owner-stop drain and its window, incoming-message drain, round compaction and its usage accounting)
      ├── task_pacing.py       ← Task-pacing SSOT: deadline/cost milestones, finalization reserve, BudgetSnapshot, acceptance launch/improvement rails — `review_launch_allowed` / `improvement_pass_allowed`, the floor evaluated ONCE per panel at loop admission, no review-duration prediction (owner R52/R55; semantics: §6 Task lifecycle; cap SSOT: `review_cycles.py`, §6 Review stack)
      ├── vision_routing.py    ← Send-time image routing SSOT: inline vision vs generic captions vs placeholders on a per-send message copy (`OUROBOROS_IMAGE_INPUT_MODE`, `OUROBOROS_MODEL_VISION`)
      ├── fallback_cooldown.py ← Per-process 429-aware cooldown for the `OUROBOROS_MODEL_FALLBACKS` chain: a transiently-failed model is parked for a short window so fallback walks and repeated rounds skip it; advisory, default-on, fail-soft, passive heal; per-process only — honestly not a swarm-wide governor
      ├── model_concurrency.py ← Per-(model,use_local) `BoundedSemaphore` capping concurrent provider calls (`OUROBOROS_MODEL_MAX_CONCURRENCY`, default 3) so a task's own loop + subagent threads + status pings cannot self-DoS one model's rate limit; excess threads WAIT deadline-bounded; wraps only the provider call in `loop_llm_call.call_llm_with_retry`; per-process only
      ├── project_naming.py    ← SSOT for LLM-first project naming: bounded light-model title with deterministic fallback, shared by the proactive card namer (supervisor/worker_chat_lane.py), turn-into-project conversion (gateway/projects.py), and `ensure_project_scope`; the provider call goes through the model_concurrency slot
      ├── loop_tool_execution.py ← Tool dispatch and tool-result handling
      ├── deadline_utils.py    ← Shared deadline parsing/remaining-time helpers + the transport-vs-logical wait seam for loop milestones and process-tool/review timeouts
      ├── observability.py     ← Private forensic execution ledger: redaction, gzip CAS blobs, call manifests, trace refs
      ├── model_send_seal.py ← The runtime invariant `model-visible ⟺ logged` for `model_send`: a reconstruction mismatch is a typed durable fact, and the call is NOT blocked — dispatch authority stays with the pre-existing in-memory identity re-check
      ├── cancel_intents.py    ← Durable cancel-intent projection: compact locked `state/cancel_intents.json` of ACTIVE intents (request id, claim owner/pid + claim GENERATION fencing every mutation, `scope` recording single-vs-cascade so a watchdog replay re-runs the right shape) + forensic `cancel_intent` ledger rows; the ONE ingress `request_cancel` for the agent tool, HTTP single/cascade, and boot migration of legacy latch files — intent never rides the canonical task status; reads are strict and fail closed per §10 (typed `CancelIntentProjectionCorrupt`; enforcement degradation is owner-visible, never a silent "no intent"); a quarantined malformed row discloses once per row content — a ~20 s watchdog must not append the same disclosure forever, and a restart re-announcing once is honest; owns `claim_is_abandoned` and the `allow_settled_target` live-ownership exception (§10, cancellation custody)
      ├── owner_hurry.py       ← Owner "hurry": a typed TASK-LOCAL acceleration latch, never a chat message; the durable `owner_hurry` projection is written by `update_json_locked` touching only its own keys — never `write_task_result`, whose status-regression guard could drop concurrent terminal fields — keyed by the real attempt identity `task["_attempt"]`; while latched, the next acceptance panel is skipped with zero reviewer calls (`acceptance_skip_applied`), remaining improvement passes overlay to 0 through `effective_budget_profile` (the immutable task_contract is never rewritten), and force-plan becomes task-locally advisory; the effect DIES WITH THE ATTEMPT (`retry_reset` on every same-id requeue producer), a never-applied request is marked `not_applied_before_terminal`, and the non-chat `owner_hurry` events are hidden from chat by `log_events.js`
      ├── owner_quiz.py        ← Owner-quiz lifecycle projection: worker-side `record_asked`, request-id-idempotent first-answer-wins `record_answered` (option index validated against the STORED labels), structural-only `reconcile_terminal` (open → expired_terminal at task done; no host TTL) which also closes the PAIRED `owner_wait` under the same task-result authority — repairable on a second pass after a partial pair write, including an answered quiz whose worker resume was not yet granted, while preserving the recorded answer — `quiz_states` replay; same locked-writer idiom as owner_hurry, touching only the `owner_quiz` and paired `owner_wait` keys
      ├── owner_wait.py        ← Native owner-answer continuation: completed-tool source checkpoint, original-process sleep, and same-task recovery only through an acknowledged planned-restart handoff; active capacity belongs to supervisor/worker_owner_wait.py
      ├── routing_wait.py      ← Root-parameterized SSOT of the durable routing-receipt waits (`wait_for_promotion_admission`, `wait_for_routing_annotation`); tools/control.py keeps thin wrappers, so the gateway picker dispatcher confirms clicks through the SAME receipts the LLM routing tools poll
      ├── outcomes.py          ← Typed task-outcome and acceptance-decision authority keeping the lifecycle/execution/objective/review/artifact/verification/child-absorption axes separate; policy denials, cosmetic exits, and ignored outcomes never masquerade as genuine tool failures; receipt reconciliation lives in `_outcome_receipts.py`, trace classification in `_outcome_tool_errors.py`; a verification ledger above the inline threshold rides as a stub whose `summary` is re-projected from the refreshed artifact file at finalization, and the stub is never a source for entries or outcome axes (the ledger embeds the task contract, which its entry count excludes, so a stub is the normal shape for a swarm root)
      ├── outcome_receipt_store.py ← Durable verification-receipt path/append/read authority + exact-row union of forked-child and canonical replicas; owns the zero-run WRITE enum (`incomplete`/`unknown` only — a zero-run "complete" is unverifiable self-report); outcomes.py re-exports the compatibility names
      ├── depth_evidence.py    ← Pure requested/permitted/attempted/achieved depth projection for root acceptance; missing admitted permission stays evidence-unknown rather than reconstructed from mutable live config
      ├── _outcome_receipts.py ← Receipt parsing and the ONE canonical receipt identity (`receipt_canonical_identity` → `ReceiptIdentity`; invariant in §10): three independent components — `criterion_id`; structurally canonical `check` text PAIRED with its `check_rendering` stamp (quoted shell punctuation is data, not syntax, and receipts from different renderings are never the same verification — the stored string alone cannot say which renderer wrote it); and the raw-sorted `canonical_path_set` (whitespace untouched — a leading space is a legal filename byte); `ReceiptIdentity.key` selects ONE typed (kind, value) and sameness is that key's equality, never a match across kinds — the parts are disclosures, never the comparison; the per-kind normalization answer lives in the closed `IDENTITY_KINDS`/`KIND_NORMALIZES_COMMAND_TEXT` table, so a fourth kind must state its own answer in its own row rather than inherit a default; the outstanding sets `unreconciled_failed`/`unreconciled_masked` scan every candidate against ALL later reconcilers and collapse repeated failures of one check onto the freshest receipt; the shared disclosed projections (`receipt_identity_projection`, `disclosed_list_projection`) make every bound explicit — exact omitted counts plus a hash over the injective serialization, string bounding via the SSOT `utils.truncate_review_artifact`, never a hand-rolled slice; `verification_receipt_ledger_row` splats that projection, so a new receipt key is dropped unless added there
      ├── _outcome_tool_errors.py ← Leaf SSOT for tool-trace status vocabularies and execution-axis classification; outcomes.py re-exports
      ├── code_intelligence.py ← Internal code inventory: derived-only file facts, hashes, polyglot symbol/import/call extraction via tree-sitter with Python on stdlib `ast`, a visible `structural_unavailable` fallback when a grammar is missing, and an incremental JSON cache (no raw source)
      ├── code_intelligence_architecture.py ← Architecture facts over the pinned domain/contract/persistence carriers: `owner_of`, the domain quotient, and the facade inventory (`docs/v7next/FACADE_INVENTORY.md`)
      ├── code_search_rg.py    ← Optional ripgrep-backed search for search_code; every match is post-filtered through the protected/secret gates
      ├── pricing.py           ← Exact-route best-effort provider-catalog lookup with nullable estimates; no static model tariffs (they go stale) and not the monetary ledger
      ├── usage_accounting.py  ← Append-only physical-model-attempt monetary authority: reserved→dispatched→settled|unresolved (or reserved→released), short cross-process check+append+fsync lock, conservative global/root admission, validated replay/torn-tail quarantine, compatibility projections, resumable legacy import; application candidates carry exact raw/context identities + a pre-dispatch manifest on the same attempt id
      ├── _usage_response.py   ← Pure provider-response usage normalization for physical accounting; a zero-usage body error settles at a confirmed $0 to release the reservation, so a provider storm cannot manufacture phantom budget exhaustion. It is the one NORMALIZER of a provider's usage block for accounting (its importers are `usage_accounting.py` and `loop_llm_call.py`). Not the only READER of that block: every provider adapter reads the raw `usage` dict for its own response envelope, so a "consolidation" here would centralize a read that was never centralized
      ├── _usage_rows.py       ← Pure row arithmetic (summaries, limit/integrity decoration, physical-call counts, breakdown buckets, the exact Skill Review wave/slot projection); no I/O or locks; re-exported by usage_accounting.py
      ├── _usage_rows_memo.py  ← Validated-rows memo + fingerprint-keyed render cache + in-lock warm read cache; every cache resumes via the substrate's `LedgerResumeState` fingerprint and falls back to the authoritative locked read on any doubt
      ├── _usage_cache_splits.py ← process-local `(task, provider, route identity, review surface)` last-observed prompt-cache split; non-durable and re-exported from `usage_accounting.py`; a lost entry only re-prices a full cache write
      ├── skill_review_usage.py ← Read-only cached projection of final physical-attempt rows for one exact `(review_skill, review_wave_id)`; no second ledger or persisted totals
      ├── usage_ledger.py      ← Durable append-only ledger substrate: cross-process locking, atomic append+fsync, row/transition validation, torn-tail quarantine; one-way seam — accounting imports it, never the reverse
      ├── usage_compaction.py, usage_legacy_import.py ← Seq-preserving compaction of that monetary ledger, sitting BESIDE the substrate rather than inside it, and the one-time legacy usage-telemetry import
      ├── cost_projection.py   ← The ONE projection of task cost for every producer surface: `accounted_upper_bound_usd` is the honest name for the settled+reserved+unresolved upper bound the ledger reports; the retired `cost_usd[_with_children]` spellings are read-only tolerance for stored legacy records (a diverged stored pair resolves deprecated-wins) and are never emitted — the write/read seams strip them; null projects as None (never $0.00); finality is never fabricated; the `COST_OPENNESS_FIELDS` accounting markers ride beside every amount; producers pass their source through `cost_projection`/`with_cost_aliases` instead of hand-assembling the pair
      ├── delegate_custody.py  ← Durable custody for delegated (Claudexor) runs: the SSOT is the `delegate_run_*` rows in the canonical event log plus ONE compact incident projection `<drive_root>/logs/containment_faults.jsonl` — the event log grows without bound, and a tail-bounded scan can bury an unresolved fault; OWNED/FOREIGN/UNKNOWN ownership replay survives worker restart; the per-intention invocation id rides the wire as `Idempotency-Key` (the deterministic per-logical-start hash is only the pending-invocation LOOKUP identity; reuse only via the explicit `retry_of` token); run settlement is decoupled from registration cleanup — `settled` follows the idempotent ledger row while the owned-registration obligation survives on `project_owned` with its own sharer tie-break and sweep, and the terminal audit discloses `deferred_project_retirements` additively; typed cancel vocabulary (confirmed | requested | failed | containment_fault_run_may_still_be_live) with durable faults riding the health invariants; one `daemon_says_absent` predicate decides everywhere that a 404 is the daemon ANSWERING the resource is gone, never a failure to find out; an ABSENT custody log is a positively-established clean state while an EXISTING-but-unreadable one audits as typed `delegated_run_state_unknown:custody_log_unreadable`, never cleanly reconciled; patch-apply intent rows (`delegate_run_patch_apply_started`/`_resolved` + the `patch_apply_pending` replay flag) make a crashed disposition typed-AMBIGUOUS instead of falsely rejected; `run_not_owned` refusals disclose `owner_task_id`/`run_settled`/`run_terminal_state`, and `run_ownership_unknown` names `get_task_result` as the ownership-free cross-task read
      ├── delegate_custody_reconcile.py, delegate_state_sweep.py ← Reconciliation of delegated runs — the settle-or-cancel sweeps and their recovery — and the terminal-plus-age sweep of the recovery/supervision state they leave behind
      ├── delegate_custody_usage.py ← Pure usage and terminal-state projections over delegated-run custody rows; the cross-process custody usage lock; the reviewer usage observers (one `llm_usage` row per ledger attempt, rows for physically dispatched failed sends, one row per delegated run across executors)
      ├── delegate_hold.py     ← Unknown-provider hold: parks the task in supervised_wait, waits for the leaf wake, never resends
      ├── delegate_source_coverage.py ← Oversized-work-order source custody: canonical interval union/completeness, strict durable receipt bounds, replay-safe start binding, durable delivery confirmation for safe receipt retry, terminal cannot-verify projection, apply refusal — incomplete source cannot authorize a terminal PASS/apply; reuses `get_task_result` + the existing interaction seam; no alternate store
      ├── delegate_evidence.py ← Read-side execution-evidence projection over the custody rows (`task_execution_evidence`: started/settled/succeeded/failed counts, terminal-state axis, `evidence_read_failed`, disclosed subscription spend, `nanny_nudge_recorded`, and `delegate_start_attempted` counting blocked and uncustodied attempts too, so a refused-but-obedient nanny is never disclosed as nudge-ignoring); owns the stamp writers `record_nanny_nudge_stamp`/`record_start_blocked`; projects `applied_access_profiles` — the access the engine actually served, read off SETTLED rows only (empty = no receipt disclosed it, never "no access") — and `acceptance_patch_dispositions`, the bounded section over `delegate_run_patch_verdict` rows (cap 20 with the exact omitted count, `unreviewed_delegated_apply` headline); absence of the section means NO disposition was recorded, never "reviewed clean", and an unreadable custody log is the typed `evidence_read_failed` marker, never an empty-therefore-clean section
      ├── synthesis_cost_text.py ← Synthesis-prompt renderers for the pre-synthesis cost/outcome snapshot over the SSOT `cost_display`; re-exported by agent_task_pipeline.py
      ├── llm.py               ← Multi-provider LLM routing (OpenRouter/OpenAI/compatible/Cloud.ru/MiniMax/DeepSeek/GigaChat/Anthropic); canonical conversations stay function-shaped while the physical-send seam delegates exact-route request adaptation to the request-wire leaves below
      ├── llm_routing.py, llm_attempt.py, llm_messages.py, llm_capability_policy.py, llm_fallback.py, llm_pricing.py, llm_openai_compatible.py, llm_anthropic.py, llm_gigachat.py, llm_local.py, llm_claudexor.py ← The client's leaves behind that facade, which re-exports the historical names: provider target resolution, client construction, route affinity, remote dispatch and subscription metadata; physical-attempt candidates and send-time prompt-cache policy; transcript shaping for the wire and the reasoning-artifact contract; route capability metadata and the learned parameter/effort policy; the recovery ladder a failed or poisoned send is retried on; the live provider price catalogs (OpenRouter, Cloud.ru) and settled-cost projection; and the wire lanes — OpenAI-compatible, native Anthropic, GigaChat, local llama.cpp and caller-owned Claudexor model operations, with their route-specific context contracts
      ├── llm_stream.py        ← Complete Chat Completions and native Messages SSE assembly inside one physical attempt; private wire/partial evidence and terminal framing, with existing response normalization and monetary custody
      ├── net_transport.py     ← Shared httpx transport construction for remote LLM clients; TCP-keepalive socket options
      ├── model_wait.py        ← Quota/auth waits bound to the existing task or phase owner: same-call continuation, role overrides, quota-aware clocks and typed stop/deadline propagation; its completed-call choices and quota union also survive native owner-wait continuation
      ├── transport_custody.py ← Typed transport facts for the physical-attempt custody seam
      ├── openrouter_attribution.py ← Canonical OpenRouter application attribution, centralized so forks do not compete under the same external application identity
      ├── openai_chat_custom.py ← Pure direct-OpenAI Chat function→custom codec: deterministic compact schemas, exact full-schema/catalog binding, tool-choice projection, prior-call replay, canonical response normalization, parser-issued validation sidecars; no Responses transcript or second stored history
      ├── openai_chat_dispatch.py ← Direct-OpenAI Chat policy leaf: custom+requested reasoning first, exact-dialect fallback with the same reasoning, then task-local explicit `none` only when the physical-attempt rail still permits it; owns private sidecar consumption + the bounded schema-error continuation
      ├── request_wire_contract.py ← Provider-neutral exact-route request profile: closed `set_value`/`drop_field`/registered `replace_dialect` actions, 14-day success-only evidence store `data/state/request_wire_compatibility.json`; task-local explicit `none` can never become durable dispatch authority — a task-local availability fallback must not teach the route
      ├── request_wire_resolution.py ← Deterministic request-profile composition with source-predicated effort bounds/transitions; contradictions fail open to the requested effort with `conflict=True`
      ├── request_wire_receipts.py ← Factory-bound wire candidates + semantic-success receipts (exact serializer digests; tool-choice semantics cannot change)
      ├── request_wire_attempt.py ← Physical-attempt validation — a settled accounting attempt for the exact candidate; public `usage.request_wire` disclosure
      ├── request_wire_custom_validation.py ← Custom-call validation; a validation failure may prove wire acceptance but cannot authorize tool execution (`allows_execution` gate)
      ├── request_wire_recovery.py ← One sync/async wire-recovery machine: durable evidence applied immediately before send, reactive evidence committed only after terminal semantic success, ordered terminal disclosures
      ├── anthropic_native_custody.py ← Whole-block replay custody for Anthropic native reasoning (same provider/endpoint/API/model); opaque type/order/size/digest projections
      ├── reasoning_artifacts.py ← Sealed-vs-portable reasoning-artifact classification, shape-first and fail-closed; only sealed artifacts pin an endpoint — portable reasoning stays failover-eligible; the `SIGNED_PORTABLE` roster is a decaying external provider fact (inventory: docs/DEVELOPMENT.md), extended only by a fresh cross-provider replay probe
      ├── llm_observability.py ← Persists public call projections; strips private sidecars from durable records while returning them in-process
      ├── llm_probe.py         ← Oversized-context evidence probe + Provider Test transport with physical accounting; no retry, fallback, or learning
      ├── mcp_client.py        ← MCP client: parses MCP_SERVERS, validates transports, masks tokens, prefixes tools `mcp_<server>__<tool>`, guarded SDK import; MCP descriptions/results stay untrusted data (§6)
      ├── safety.py            ← Safety supervisor call with a bounded newest-first context budget (the omission marker is reserved INSIDE the budget); a 429 on the safety check is an infrastructure fact about the supervisor, not a verdict about the tool call — one deadline-capped retry, then the typed `⚠️ SAFETY_UNAVAILABLE` non-verdict telling the agent to retry the same call, not reword it; process-local storm latch answers in-window checks without provider calls; durable `safety_check_rate_limited` audit event; structured insufficient-quota keeps its PERMANENT classification and still blocks as a verdict; the serialized SUBJECT has its own 250k-char budget (`_SAFETY_SUBJECT_CHAR_BUDGET`, rendered `ensure_ascii=False`) and an over-budget subject is refused fail-closed with the typed `⚠️ SAFETY_SUBJECT_TOO_LARGE_BLOCKED` denial plus a durable `safety_subject_too_large` event — never truncated, because anything past a cut would run unreviewed; the fail-open cases are owned by prompts/SYSTEM.md
      ├── consciousness.py     ← Background thinking loop with live progress emission (§6)
      ├── consolidator.py      ← Dialogue consolidation with a generation-aware cursor over the ordered archive chain; an unfindable generation appends a loud durable `[MEMORY GAP]` block, never a silent offset reset
      ├── memory.py            ← Scratchpad, identity, chat history
      ├── memory_journal_compaction.py ← Digest-only compaction of old memory-journal snapshots: the digest replaces the snapshots it summarizes, never a silent drop
      ├── project_facts.py     ← project_id resolution (explicit `--project-id` or workspace-path hash); per-project knowledge dir `projects/<id>/knowledge` isolated from `memory/knowledge`; journal/workpad helpers
      ├── task_tree_ledger.py  ← Append-only `data/task_trees/<root>/blackboard.jsonl`: EPHEMERAL swarm coordination in typed validated rows (prose is never authoritative), distinct from the durable project journal it is mirrored into by `tools/project_journal.py::mirror_tree_coordination_to_journal`; size-capped writes; exposed via `tree_note`/`tree_read` with the tail injected each turn; pruned on root terminal
      ├── projects_registry.py ← Durable `data/state/projects.json`: 80-char names, `active|deleting|tombstoned`; deletion preserves bindings/history/folder/memory; a tombstone blocks resurrection; reconcile registers existing stores and NEVER prunes
      ├── project_dialogue.py  ← Read-only chat lens + append-only `logs/chat_annotations.jsonl`; the sidecar never owns routing except the `needs_manual_target` decision card (token+options validate the click; first-wins closing rows); `build_owner_message_ref` mints refs at ingress
      ├── project_lease.py     ← One-writer-per-project lease in `assign_tasks`; same-project subagent swarms exempt; `""` is no lane
      ├── context.py           ← Main agent context assembly; places ordered subagent-catalog JSON under `## Available subagents` in the semi-stable block
      ├── main_context_authority.py ← Deep-copies the context authority; replaces only oversized raw result strings with source-resolvable narrative or a typed gap
      ├── client_surface.py    ← Closed-key bounded client-surface normalizer; surface identity excludes viewport/narrow_layout; mailbox surface-change note
      ├── context_fit.py       ← Deterministic Max/Low context projections from one immutable core with labelled measurement + typed reclaim deficit; no routing/retry/global-mode authority; owns the message-side transcript cache seal (exactly ONE message-side breakpoint: the task-contract boundary until a rolling tool-result seal qualifies, migrated in the same call so the four-breakpoint cap can never drop it)
      ├── context_budget.py    ← Context budget vocabulary + typed reclaim SSOT (owner-Low 200K economy target); owns `estimate_message_chars` (images counted at `IMAGE_BLOCK_CHAR_EQUIVALENT`, replayed `reasoning_content` counted on the DeepSeek echo lane), the bounded basis of the local compaction proxy; remote fit and the density witness measure on `context_fit.estimate_context_prompt_tokens`
      ├── context_mode_compat.py ← One-window compatibility shim for the retired persistent context auto-Low state
      ├── capability_evidence.py ← Sourced capability evidence in `data/state/capability_evidence.json` (confirmed/asserted/unprobeable/failed): authorizing readers require fresh evidence, and unknown fails the ≥1M gates closed; owns `observe_token_density` — the density witness calibrates on the bounded-proxy basis the fit estimator measures, while budget reservation keeps RAW, because over-counting money is the safe direction and the two consumers split on purpose; owns `cold_start_density_probe` too — the one bounded exact-model send that sources a witness when a review pack is refused or degraded under the cold floor; exact-route dispatch authority lives in `request_wire_compatibility.json`
      ├── context_layout.py    ← Doc-layout SSOT: tier-0 always full; ARCHITECTURE full in Max and a lossless fence-aware navigation map in Low; DEVELOPMENT full-or-pointer by task binding; reduction is by relocation with a visible pointer, never silent truncation
      ├── context_compaction.py ← Atomic-unit compaction: exact checkpoint, gap-free map/fold, provenance capsules, transactional apply on the caller's basis; an unfinished Anthropic native unit is ineligible; opaque custody never enters summarizer text
      ├── context_health.py    ← Health invariants for the reading task (`build_health_invariants`); delegated-run obligations stay globally visible — a preserved-and-invisible result is how work rots on disk — while the instruction is ownership-aware, so a non-owner is never handed a call that structurally refuses: it states the rule statically (owner-only while the owner task is live, a live top-level holder of the same target once it is terminal) and reads no per-orphan task result. The caller threads its own `active_root`, so when that root already satisfies the recorded target under the same predicate the apply gate asks (`delegate_shared.orphan_apply_target_ok`), the static rule is followed by the concrete `integrate_delegated_patch` call — one comparison over a value the caller already holds, still no per-orphan task result. `build_health_invariants` runs ONCE per task attempt from `build_llm_messages` (and once per Background Consciousness wake), so the block is a task-start snapshot and does not refresh mid-task
      ├── context_runtime_facts.py ← The runtime section's FACT builders: what the host can honestly say it knows about this turn
      ├── headless.py          ← Child-drive isolation, workspace patch artifacts, memory export helpers; a sensitive-looking untracked credential is excluded per-file and disclosed as `sensitive_blocked`
      ├── headless_status.py ← Artifact and task lifecycle vocabulary shared by the headless owners
      ├── workspace_patch_rules.py ← Pure patch-exclusion rules (env/cache sets, junk regex, lockfiles, credential-shaped names); the I/O checks + `untracked_capture_veto_reason` stay in headless
      ├── workspace_patch_capture.py ← Workspace patch capture: the patch artifact, its manifest, and its git plumbing
      ├── coop_checkpoint.py   ← Quiescent checkpoint commits of cooperative trees (detected by supervisor/events.py, run off the drain thread): a tree qualifies only through a MUTATIVE child's `write_root` — owner-attached folders are never auto-committed; credential-shaped files excluded + disclosed; quiescence revalidated pre-mutation; a root whose owner is mid merge/rebase/cherry-pick/revert is SKIPPED with the operation named in its receipt (`skipped`), because staging an interrupted operation consumes its MERGE_HEAD and commits a half-resolved tree
      ├── delegate_output.py   ← Atomic staged full outputs under `delegated_runs/<run>.json` (sha256+length); `acknowledge_staged_output_read` hooks read_file's task_drive path; once-per-run `delegate_run_output_consumed` row — a disclosure, never a gate
      ├── delegate_containment.py ← Engine-derived isolation facts: a home-isolation breach is exactly two facts (`harness_home_isolated: false`, or applied home == the operator's own); a home nested under `$HOME` is disclosed-unconfined (`home_nested_under_operator_home`), never relabelled as isolation; absence is reported unproven
      ├── delegate_progress.py ← Transport read bound + transient Git-object retry (`poll_bound`); renewal/wake policy lives in delegate_supervision; publishes event-local executor observations from the already-polled owned timeline, not a current-executor authority
      ├── nanny_pacing.py      ← Metered-silence pacing: only `BASELINE_RESET_TOOLS` (`delegate_start`/`schedule_subagent`) reset the burn; supervision verbs advance the round baseline while dollars accumulate — coordination never buys metered silence
      ├── delegate_interactions.py ← Child-interaction custody: reported-interaction memo, bounded display scalars with whole answer keys, typed `waiting_on_user` with an immutable spill file, strict `_delegate_answer` validation with typed rejections (`subscription_window_exhausted` carries `reset_at`; transport death/5xx is `delivery_unknown`)
      ├── delegate_shared.py   ← Shared delegate leaf (`_fail`/`_emit`/`_owned_run` — OWNED/FOREIGN/UNKNOWN from durable rows — plus `orphan_disposition_status`, the disposition-only OWNED upgrade for a terminal owner's orphan held by a top-level principal; `_owned_run` governs wait/cancel/answer and is deliberately NOT widened); one-way seam, facade re-exports
      ├── route_spec.py        ← Neutral route primitive: kind/target/pin normalization + effort validation; semantic owners keep their own spelling
      ├── configured_subagents.py ← Canonical `OUROBOROS_SUBAGENTS` parser/serializer: strict validation, stable ids, fingerprinting; owner free text is never host-parsed
      ├── subagent_runtime.py  ← Immutable task-start subagent snapshots, exact `subagent_id` selection, typed alternatives, bounded deterministic legacy-input seam
      ├── subagent_route_health.py ← Route health: the ONE manifest reader behind every delegated dispatch
      ├── subagent_work_order.py ← Complete chosen work-order compiler and normalized host authority without arbitrary admission cuts; exact-source readers and legacy source-response validation share the unchanged renderer
      ├── subagent_bootstrap.py ← Host pre-start of the exact snapshotted leaf BEFORE the first metered round, through the same wrapper as `delegate_start(prompt="")`; branch order recovery → fences → blocked → pre-start, and a fence-wake outranks every terminal because a fence may hide a live run; the host never waits (`configured_session_started` receipt); only a definite typed refusal ends unrun at $0 — everything ambiguous wakes the model, because a false "spent nothing" terminal over a possibly-live run is the one direction classification must never fail toward
      ├── delegate_supervision.py ← Event-only sleeping-nanny loop: quiet windows renew without a model call; terminal/interaction/fault/addressed/control (or one reasoned checkpoint) triggers a durable wake with fresh coordination context (parent intent, time, tree spend, host-visible descendants and root review capacity — every fact observed READ-ONLY, so a metadata-poor task reports `time.state = "not_set"` rather than latching an anchor from a poll; polling writes nothing of its own and inherits only the canonical usage-ledger reader's bounded maintenance — the torn-tail quarantine after a SINGLE crash mid-append, which every reader performs identically (a crash inside that repair itself, a torn quarantine sink, is a known residual: issue #586), the empty `state/` directory the reader's lock lives in on a never-initialized root, and removal of a stale `usage_attempts.lock` older than the reader's 90 s stale window (`usage_ledger._locked` → `platform_layer.acquire_exclusive_file_lock`, whose stale-age branch unlinks the lock file and retries) — each pinned by a regression; an absent ledger answers known-zero settled spend through that same canonical reader); replay returns the stored snapshot
      ├── delegate_start_instructions.py ← Stable host start instructions + a complete separately-hashed coordination appendix; host pre-start sends no appendix
      ├── delegate_recovery.py ← Narrow exact-leaf recovery for proven crash + planned self-restart; validates bindings; vetoes every no-resume cause
      ├── delegate_registration_policy.py ← `persistent_registration` + the STARTED-row field tables
      ├── delegate_pending.py  ← Durable pending-invocation replay preserving the original idempotency key + canonical start body
      ├── delegate_terminal.py ← Terminal reconciliation + custody-audit persistence, split by surface: counters stay a frozen historical snapshot while `actual_substrate` and the envelope mirror are rewritten from live custody, so the executor chip reads live truth while history stays a snapshot; audit-only in both directions — unreadable custody proves nothing; `refresh_recently_settled_terminals` rides a durable byte-offset cursor (`state/delegate_terminal_refresh_cursor.json`, 5 MB per tick, deferred map for still-running parents)
      ├── subagent_dispatch_notes.py ← Dispatch-time executor notes for delegated children (configured-nanny charter note; non-configured branch keeps "decide your delegation plan first"); agent.py keeps re-exports
      ├── subagent_messages.py ← Bounded durable child-message identity shared by the final frame, recovery, compact persistence, and replay; `executor_observation_meta` separately validates/copies task-bound progress actor facts without changing final-lineage fields
      ├── subagents.py         ← Subagent envelopes + bounded legacy compatibility; `configured_subagent` snapshots dispatch through subagent_runtime
      ├── subagent_worktrees.py ← Worktree lifecycle + durable registry `state/subagent_worktrees.json` with ops lock and startup orphan reconciliation; `provision_genesis_project` (never registry/GC); execution snapshots pinned by `refs/ouroboros/delegated/`; standalone payload snapshots (CAS hash, writer-race abort); removal only explicit or custody-cross-checked startup GC — fail-closed: an unreadable custody log replays as "no open runs", so the prune skips entirely rather than destroy a child's only copy of its work (prune events are emitted at the server.py call site)
      ├── artifacts.py         ← Attachment staging into `artifact_store/attachments/` (secret-source skip, bounded, read_file manifest); artifact records exclude attachments + `chat_media/`; scratch fingerprints (`.scratch_manifest.json`, both roots) gate patch exclusion only while content matches; the undeclared-output guard stat-verifies post-exec; `delegated_capture_read_target` rebinds `delegated_runs/` reads to the canonical drive
      ├── retention.py         ← Unified GC retention SSOT: clamp/age-cutoff + legacy-key seed picker
      ├── workspace_preflight.py ← Read-only external-workspace git/manifest/toolchain snapshot used by gateway task creation
      ├── project_sources.py   ← Folder attach validation (realpath, not the home root, no repo/data overlap); opt-in `init_git`, NEVER auto-init; atomic server-side clone with `GIT_TERMINAL_PROMPT=0` and typed `auth_required`; provenance + `clone_url` recorded; attaching IS the trust grant (`trusted_at` automatic)
      ├── promotion_source.py  ← Promoted-task source admission off the event-drain loop, only after an executor/id reservation
      ├── workspace_admission.py ← Shared admission for `/api/tasks` + promotion: disjoint git root, Project binding, `workspace="none"`, bounded preflight, loud refusal of a broken binding; empty-Project promotion provisions idempotent genesis, and a failure is typed `workspace_provisioning_failed` on the promotion path (supervisor/workers.py) — never a fallback onto the system repo
      ├── local_model.py       ← Local LLM lifecycle (llama-cpp-python)
      ├── local_model_autostart.py ← Local model startup helper
      ├── deep_self_review.py  ← Deep self-review of the whole system against BIBLE.md on the configured `deep_review` reviewer row (`reviewer_slot_config.deep_review_slot`; absent = the packed api row synthesized from `OUROBOROS_MODEL_DEEP_SELF_REVIEW`). THREE deliveries by the row's `retrieves` predicate. A direct api row is the PACKED review: deep-review atlas + the full memory whitelist on a ≥1M route → one `chat_observed` call, byte-identical on the wire to the pre-row review (golden digest test); a bounded in-prompt OMITTED section is reserved inside the fixed budget; `atlas_assembly_failed` (over hard budget, or a REQUIRED artifact omitted) ships no pack at all — the compact manifest is the atlas default, so there is no compact retry rung — and a final-shrink rebuild (tighter hard budget by the measured overage) replaces the historical fatal 'Review pack too large' error, the gate remaining as the fail-closed last assertion; a required set that does not fit under a COLD density cap (no fresh exact-model witness — the refusal precedes every send, so the model could never record one) gets ONE bounded probe send on the exact model (the shared rung `capability_evidence.cold_start_density_probe`, the same one the commit gate runs: a slice of the real pack — `review_helpers.density_probe_sample` — `DENSITY_PROBE_MAX_TOKENS` output tokens, through `chat_observed`) that records the witness, then ONE rebuild under the recalibrated cap; a pack that still does not fit is the typed `deep_self_review_pack_unfit` refusal whose text asks the owner to switch the `deep_review` row to a retrieving delivery or a larger-window model — the host never falls back to another delivery on its own; centrality ranking (reverse in-degree) is deep-review-only. A configured-subagent api row is a NATIVE inspection episode and an `agent_session` row a delegated session, both through `review_execution._review_route_executor` exactly like the advisory: a hand-built `ReviewRequest(surface="deep_self_review")` whose route-owned task carries the role prompt, up to seven memory files inline byte-exact — EVERY whitelisted entry with its disposition (`inlined` / `missing` / `empty` / `oversized` / `read_error`) in the task's memory section, in the usage fact `deep_review_memory` and in the header (`memory=n/7`) — BIBLE.md as a MANDATORY full read and ARCHITECTURE/DEVELOPMENT/CHECKLISTS as `generate_doc_nav_map` navigation maps; `policy["output_contract"]` = the report contract plus the deep review's coverage-header sentence (the shape is `report` either way) and `policy["native_data_root"]` = the real runtime root (readable by the reviewer's tools; the memory whitelist itself arrives inline); a `ReviewSlot` with an explicit logical window (the task's absolute ceiling narrowed by the owner deadline); `record_reviewer_slot_executions("deep_self_review", …)` for «Выполняется как» — recorded on all three retrieving outcomes (responded, empty response, executor exception) from a usage that already carries `deep_review_memory`, which the returned usage keeps too (a typed failure included, together with the executor's failure custody); the durable execution projection itself persists route/model/status/capability_delta and typed failure facts only, never a deep-review-only field; prompt/response custody through `persist_call`. Every delivered report is prefixed by the host provenance header (`<!-- deep-review provenance: delivery, model, memory[, memory_missing/memory_empty/memory_oversized/memory_read_error], coverage, incomplete, attestation[, rounds, tool_calls, receipts, end_reason, transcript, landing — native only] -->` + one human line; every comment value AND every external value on the human line sanitized and bounded (`_header_value`); `incomplete` is derived per delivery — the native episode's typed `native_incomplete`, the packed call's provider stop marker — the OpenAI-compatible `response_finish_reason == "length"` or the direct-Anthropic message `stop_reason == "max_tokens"` — as `output_reserve`, a session's completeness `unobserved`); after a native episode BIBLE.md coverage (`read` / `partial(fraction)` / `missing` / `unobserved`) is derived by `_native_read_coverage` from the executed repository-root `read_file` receipts on the reader's `opened_path` and `opened_root` (contract and edge rules: its docstring and the «Deep self-review» section) and anything but `read` is disclosed in the header and as a typed `capability_delta`, never refused; a session's coverage is `unobserved`. Availability is route-aware (`deep_review_route`): the packed row keeps the ≥1M floor (a window CONFIRMED below 1M by the shared `reviewer_window` resolver is a typed refusal — the pack is never shrunk to a smaller window; an unknown window keeps the full-window assumption, disclosed as `window=assumed_1000000`) and the `OPENAI_BASE_URL` trust rule, a native row needs its model's credentials, a session row the substrate's `route_health`. Every failure returns TYPED usage (`execution_status=infra_failed` + reason code) so `agent.py` never overwrites `memory/deep_review.md` with an error
      ├── review.py            ← Repository size-ratchet inventory, code collection, and complexity metrics; official CI enforces the shrink-only module/function/byte ceilings while every local surface only warns (§6)
      ├── size_ratchet_manifest.py ← Generated data-only size-debt manifest (regenerated by scripts/regenerate_size_ratchet.py)
      ├── review_execution_projection.py ← Pure read-side reviewer-execution projection: bounded rows, 2000-char post-redaction string bound, unknown shapes ship as disclosed JSON; kinds `api` | `harness` | `native` — the native tool-round episode is an API execution with a different DELIVERY, projected as its own kind so the owner can tell a retrieving review from a packet review on the same model
      ├── preflight_runner.py  ← Hermetic pre-commit test runner: ONE hardened raw-bytes `git diff --binary … HEAD` capture, because a staged+unstaged diff pair cannot faithfully materialize unmerged entries; capture/apply failure is a typed `PREFLIGHT_CANDIDATE_ASSEMBLY` block, never a test verdict; node lane first, then the two-pass parallel/serial pytest split under one budget (`LANE_EXCLUSION_EXPR` is the marker-lane SSOT); a dead xdist worker or a missing required plugin is a distinct named block, never a retry or silent serial fallback
      ├── preflight_node.py    ← Content-keyed Node test lane: bundled signed node first then PATH, floor 20.11; typed `PREFLIGHT_NODE_MISSING`/`PREFLIGHT_NODE_TOO_OLD` hard blocks and `NODE_TESTS_FAILED` on a red suite, never a silent skip; both CI jobs mirror `cd web && node --test tests/*.test.js`
      ├── review_substrate.py  ← Review slot coordinator (the row-identity mint and paid stamp live in `review_dispatch.py`, the slot builders in `reviewer_slot_config.py`, the verdict reducers in `review_verdict.py` — all re-exported here): duplicate ids run as independent slots; per-actor records keep transport/parse/verdict/coverage/quorum/hashes distinct with a compact projection outward; acceptance enforces adaptive quorum, one substantive interaction per actor (≤2 physical sends on a packet row; one bounded episode on a native row; one delegated session on a session row — a retrieving row's verdict is equally authoritative, owner R1), provenance, and public-info anti-cheat; task acceptance and plan review follow each configured row's delivery kind (`api_chat` packet in-process, `agent_session` retrieving reviewer, a configured-subagent api row as the native episode) — task acceptance through the same `reviewer_slot_config.triad_delivery_slots` builder as plan and skill review (owner R2, 2026-09-01; the earlier api-row-only pin and its default-panel projection are gone); a session's narrative is canonicalized by the surface's output SHAPE (`triad_review.review_output_shape`: `array` — bare `[]` or a findings array, so a clean verdict survives `empty_array_is_verified_clean` unchanged; `object` — the whole acceptance verdict; `report` — passed through verbatim, no schema asked); the delivery seam below it is owned by review_execution.py (§6 Review stack)
      ├── review_custody.py    ← Process-local physical review custody: independent deadlines, late-result settlement, stable retry identity, duplicate-dispatch suppression; not durable scheduling
      ├── review_owner_custody.py ← A paid attempt records its owner `(server session, pid)` before fan-out; only confirmed pid death or a later-generation observation settles tokenless rows; adds no scheduler
      ├── review_execution.py  ← The ONE review-delivery seam below the substrate: closed route vocabulary with THE delivery-class predicate beside it (`delivery_retrieves(route, subagent_id)`: a hosted session or a configured-subagent api row retrieves the subject itself and never receives the packet — `ReviewSlot.retrieves`, `ConfiguredReviewerSlot.retrieves`, admission, packet fit and the surfaces' request builders all call this one definition), immutable `ReviewAssignment` bound once via `_review_route_executor`, the single physical seam `_execute_slot_attempt`, typed `ReviewAttemptResult`; no cross-transport fallback — a route that cannot deliver raises on its own slot (`ReviewRouteUnavailable`), and the durable prompt record and both physical sends share one byte-identical rendering; `ApiChatReviewExecutor` renders lazily and memoizes (digest-pinned); `AgentSessionReviewExecutor` runs one delegated read-only session — `outputSchema` is sent only when the route's own live manifest declares it and trusted only on `outputConformance == "passed"`, else strict parsing then light-model extraction disclosed as `capability_delta`; `review_output_contract(request)` is the ONE governance text every delivery honours (the api pack renders it into its byte-stable segment; a surface hands the same text to its retrieving rows as `policy["output_contract"]`), and `ROUTE_OWNED_POLICY_KEYS` (`output_contract`, `native_data_root`) are consumed by the retrieving executors and never rendered into the api pack's Policy JSON; per-row delivery via `OUROBOROS_REVIEW_ROUTES`/`OUROBOROS_SCOPE_REVIEW_ROUTES`, session target `OUROBOROS_REVIEW_SESSION_ROUTE` falling back to `OUROBOROS_SUBAGENT_HARNESS`; task acceptance follows its configured rows like every surface (owner R2 — the former `api_chat` pin is gone). Disclosed residual (pre-existing release behaviour, b9bcc2da; issue #588; out of this change's scope): a compatibility transport that raises with positive physical capture invokes the paid stamp after the send; if the tree's last paid cycle is consumed concurrently at that late stamp, the wallet refusal replaces the captured exception and the substrate may resend
      ├── review_native_episode.py ← NativeToolRoundReviewExecutor: one read-only inspection episode for configured-subagent api rows, including advisory; window-derived transcript bound, owner deadline and paid ledger, no round cap. Host-observed read receipts and typed terminal custody keep retrieval distinct from assembled coverage (§6 Review delivery).
      ├── review_verdict_extraction.py ← Session/native verdict canonicalization: strict parse first, then light-model extraction to the review's own contract; branches on the surface's output SHAPE (`triad_review.review_output_shape`): `array` keeps the historical findings ladder; `object` (task acceptance) keeps the WHOLE verdict object on the schema, strict and extraction branches (an acceptance object is never reduced to its findings list, and `[]` is not a strict object verdict); `report` (deep self-review) passes the product through verbatim, never extracted
      ├── review_session_custody.py ← Exact delegated-review recovery validation + pre-POST durable invocation checkpoint; no scheduler or state store
      ├── review_slot_cancel.py ← Slot-cancel honesty: a cancel outcome reports only what it PROVED ("host-cancelled" only on a confirmed own-state receipt; confirmed failed/interrupted is attributed to the run's own terminal); a succeeded run whose result read fails is typed `ReviewSessionSucceededResultUnavailable` after one bounded retry, never "may still be live"
      ├── review_actor_aggregation.py ← Contract aggregation for completed review actor rows; demotes non-contract-valid responses
      ├── review_session_usage.py ← UsageScope-to-custody attribution for delegated review sessions
      ├── review_thread_continuity.py ← Thin Claudexor thread operations for delegated plan reviewers
      ├── commit_admission.py  ← Deterministic commit-admission SSOT: release checks + auto-sync, staged-Python syntax compile, `run_tests_preflight_with_proof` forwarding actual runner evidence to ordinary/managed consumers; the advisory and commit gates both delegate here
      ├── reviewer_slot_config.py ← Structured reviewer-slot SSOT: stable ids, route targets, per-slot effort, disclosure-only execution records, save/runtime validation (`reviewer_slot_save_check`, whose only disclosure is the one-time R12 notice `acceptance_delivery_disclosure` when a save first makes the triad retrieve); a row is EITHER an inline route OR a `subagent_id` (materialized at load; unresolvable = typed refusal); `is_session` is transport while `retrieves` is delivery class; advisory shares the row vocabulary (+`enabled`, `disabled_reason`), and so does the optional `deep_review` singleton (fixed id `deep_review_slot_1`, no `enabled`; `deep_review_slot()` returns the saved row or the packed api row synthesized from the legacy `OUROBOROS_MODEL_DEEP_SELF_REVIEW` key, whose own effort outranks `OUROBOROS_EFFORT_DEEP_SELF_REVIEW` only when set); malformed config refuses EVERY surface — commit/scope/advisory/plan/skill review, deep self-review and task acceptance (owner R3; the former legacy/default API panel residual is gone); `triad_delivery_slots` is THE triad-row builder: plan review, skill/commit review (as aligned vectors through `commit_triad_delivery`) and task acceptance all read the rows through it, so no surface reads a projection of the panel instead of the panel; the legacy comma keys remain a runtime projection of api model ids for legacy consumers only
      ├── review_state.py      ← Durable advisory pre-review state (`state/advisory_review.json`)
      ├── review_state_model.py, review_state_records.py, review_state_custody.py, review_records.py, review_verdict.py, review_projection.py, review_evidence_sections.py ← The review ledger and its vocabularies, re-exported through `review_state.py`: the in-memory ledger and every transition it permits; the record types and the pure rules that shape them; durable custody of in-flight invocations plus attempt-history hygiene; the typed panel records and hardness vocabulary shared by every review surface; the pure reducers from panel actor rows to a verdict, a tier and a capsule; panel identity and the compact redacted projection of a run; and the bounded, provenance-tagged sections of the task-acceptance evidence packet
      ├── review_cycles.py     ← Shared paid-cycle cap SSOT (`OUROBOROS_REVIEW_MAX_CYCLES`, positive int or `unlimited`); the four per-gate meanings live in the module docstring (§10)
      ├── review_dispatch.py   ← Review row-identity mint + the write-ahead PAID stamp: typed pre-start refusals stay $0 while a late worker or crash cannot race away the durable paid fact; acceptance binds a strict exact-hash tree-wallet claim to that same stamp on every delivery its rows run — the API ledger transition of a packet row, each paid send of a native episode, `START_REQUESTED` of a session (owner R11: the paid identity is material, not route; one idempotent claim per panel); a wallet/cancellation veto at the claim releases the reservation (the launch floor is the loop gate's, evaluated once at admission; the R23 clamps bound a running panel), and, on the canonical paths, no reviewer transport proceeds on any row (the compatibility positive-capture path disclosed under `review_execution.py` — issue #588 — is the known exception)
      ├── reviewer_window.py   ← ONE typed `ReviewerWindow` per route (window/status/stale/observed_at + computed `blocking_authority_allowed`); metadata-only probe, per-route-locked, rate-limited by the evidence TTL — never a process-lifetime memo that would outlive the record; fail-closed sub-floor with no evidence; reserves scale to sub-1M windows
      ├── triad_review.py      ← Shared review primitives: JSON-array extraction (repo + skill), per-actor records, quorum/degraded accounting; owns `REVIEW_JSON_ARRAY_CONTRACT`/`REVIEW_JSON_MATRIX_CONTRACT`; a clean verdict is the WHOLE response `[]` (± one fence, ± `NO_FINDINGS`) — a refusal cannot be distinguished from a benign preamble by structure, so prose or a bare sentinel is a parse failure; the contract text lives beside `empty_array_is_verified_clean` so the two cannot drift; also owns `REVIEW_OUTPUT_SHAPES` / `review_output_shape(surface)` — the ONE form fact (`array` | `object` | `report`) the retrieving-route canonicalizer, the session output schema (`review_execution.review_session_output_schema`) and the strict parser branch on; shape is form only — acceptance rules, tier classification and quorum authority stay with their owners
      ├── onboarding_wizard.py ← Shared desktop/web onboarding bootstrap + validation (§2)
      ├── subscription_install_presets.py ← Pure sibling install compilers from one normalized draft + one discovery snapshot; output is linear, unpinned, exact-discovery-backed, all-or-nothing
      ├── settings_setup_contract.py ← SSOT for the setup contract, derived bootstrap state, payload validation, and the moved `TOTAL_BUDGET` resolver authority `resolve_total_budget_usd`
      ├── owner_mailbox.py     ← Per-task user message mailbox (compat module name); full revocation-aware drain and wait-local proven-empty peek
      ├── launcher_bootstrap.py ← Bundle-to-repo bootstrap + managed sync helpers (used by launcher.py)
      ├── launcher_onboarding.py ← First-run onboarding as the desktop launcher presents it (serves the gateway /onboarding page; §2)
      ├── launcher_server_reaper.py ← POSIX same-install server discovery, pre-signal descendant capture, root-first termination, live identity revalidation, bounded survivor reporting; PID-lock-owning launcher only
      ├── launcher_windows_runtime.py ← Windows-only pythonnet/pywebview runtime preparation
      ├── provider_models.py   ← Model-ID helpers; the `ACTIVE_MODEL_SETTING_KEYS` vs `LEGACY_MODEL_SETTING_KEYS` split keeps Heavy out of startup/Provider Test/new consumers while migration/history still read it
      ├── runtime_mode_policy.py ← Protected-path policy (safety-critical files, frozen contracts, release/managed invariants) shared by the registry, git tools, and gateway guards
      ├── schedule_contract.py ← Schedule id, 5-field cron, IANA timezone validation SSOT
      ├── reflection.py        ← Execution reflection and pattern capture
      ├── post_task_evolution.py ← The worker writes a durable promotion signal; the supervisor idle tick applies it via the existing gated enqueuer (one-shot autostop); never enqueues from the worker, never fires from evolution/subagent tasks
      ├── repo_remotes.py      ← Role-based remotes: `managed` is the read/update-only official source; `origin` is the personal target auto-configured from the GitHub token
      ├── review_evidence.py   ← Same-execution commit-review source selection, redacted browser/vision calls and actual same-round image-attachment observations, with canonical source handles and exact pending reuse (§6, Git and commit review); bounded provenance-tagged task-acceptance evidence from effective task/plan claims, verification support, artifacts, tool trajectory, obligations, and retrieval facts; ingress claims win over the current closed plan wave, projected without mutating the live task contract; for harness-dispatched tasks the packet carries a host-attested `substrate_execution` section and the sibling `delegated_patch_dispositions` from delegate_evidence — VISIBILITY ONLY, zero typed rules tie substrate to the verdict: acceptance judges quality, never the execution route, and because `integrate_delegated_patch` has no review facts on its path the packet ATTESTS the apply rather than inventing a review
      ├── review_evidence_refs.py ← Leaf SSOT of the evidence-ref vocabulary + exact-membership resolver; unsupported claims cannot certify (`CLAIM_ID_UNSUPPORTED`)
      ├── review_status_projection.py ← Leaf commit-review status projection over `review_state` records (`build_review_projection` / `build_review_status_payload` and the typed run-failure reason renderer); re-exported by `review_evidence` so the historical import sites keep resolving
      ├── semantic_dedup.py    ← LLM-first semantic dedup, fail-open None; consumed by improvement_backlog + review_state
      ├── betterleaks_runtime.py ← Pinned Betterleaks runtime resolver (six platform artifacts, packaged-resource-first)
      ├── skill_loader.py      ← Skill discovery over `data/skills/{native,clawhub,ouroboroshub,external}` + `OUROBOROS_SKILLS_REPO_PATH`; `.self_authored.json` marker; per-skill state under `data/state/skills/<name>/`
      ├── skill_readiness.py   ← Execution readiness and phase-specific next actions from review, hash, enablement, grants, dependencies and peer conflicts; acceptance also exposes actual loaded revision/state without inferring a passed functional test
      ├── skill_dependencies.py ← Shared dependency-spec resolution and installed-readiness probe for skills
      ├── skill_repair_admission.py ← Selected-skill development admission: immutable `base_content_hash` and observed `expected_content_hash`; verify before each payload operation and advance after opaque process work with attribution unproven, without a long shell lock or rollback
      ├── skill_owner_attestation.py ← Owner attestation lane: the owner may skip the expensive LLM skill review for skills they authored themselves
      ├── skill_publish_snapshot.py ← Immutable captured-byte authority for publication
      ├── skill_publish_scanner.py ← Betterleaks scan adapter; literal `high` findings block
      ├── skill_publish_result.py ← Typed publish attempt/receipt + finalization veto
      ├── skill_publish_github.py ← GitHub publication transport after the local gates
      ├── skill_publish_eligibility.py ← Passive publish visibility + `task_start_allowed`
      ├── skill_review_status.py ← Verdict aggregation → `executable_review` (anchors the §13 readiness statuses); current Advisory author acceptance may remain valid while the original critic hash/status is stale, while Blocking requires fresh critic authority
      ├── skill_review_passes.py ← One multi-model pass or chunked quorum; reserves the complete operation roster before dispatch, rejoins exact logical waves from lifecycle history, and binds process-local custody to wave/chunk rather than a restart index
      ├── skill_review.py      ← Skill review orchestration: preflight + advisory critic via `run_advisory_critic`; tri-model gate against the Skill Review Checklist (docs/CHECKLISTS.md) + docs/CREATING_SKILLS.md
      ├── skill_review_prompt.py, skill_review_packs.py, skill_review_output.py, skill_review_rebuttals.py ← The skill reviewer's leaves: its prompt contract, governance context and waves; the reviewable payload — what a reviewer may see and how much of it; the parsed findings, aggregate verdict and rendering; and the review-history evidence a reviewer reads before re-judging
      ├── skill_review_history.py ← Write-ahead review-history marker (`physical_attempt_v1`) and idempotent `state/skill_review_root_tasks.jsonl` projection; append failures emit typed `skill_review_history_append_failed`
      ├── skill_review_cycles.py ← Paid skill-review cycle counting, exact prior-wave selection, $0 replay and typed exhaustion; shared cap SSOT is review_cycles.py
      ├── extension_loader.py  ← Extension loading: in-process pure-Python via `PluginAPIImpl`, child-process proxies for isolated-dep/native extensions
      ├── extension_process_runner.py ← Extension child processes: scrubbed env, per-skill deps, timeouts, graceful host errors
      ├── extension_route_stream.py ← Portable stdio response frames and ASGI relay; child standard responses retain headers/HEAD/Range/background work, parent backpressure and exact-bundle cancellation use the existing runner process owner and supervised_futures
      ├── extension_ui_validation.py ← The one host-owned declarative-schema-v1 widget validator
      ├── extension_isolated_deps.py ← Legacy/forced in-process bridge for isolated-dependency extensions
      ├── extension_health.py  ← Durable process-qualified per-skill health at `data/state/skills/<name>/health.json`; server observation is authoritative, worker observation is a handoff-qualified view
      ├── extension_plugin_api.py, extension_registry_state.py, extension_liveness.py, extension_child_catalog.py, extension_import_staging.py, extension_surface_names.py ← The extension runtime's leaves: the `PluginAPI` object handed to one extension's `register(api)`; the process-wide registries of the surfaces live extensions own; the liveness authority for one extension (what it should be and what it is); host-side validation of the surface descriptors a child catalog run returns; the staged import trees for in-process extensions and their reclamation; and the provider-safe naming and syntax rules for extension surfaces
      ├── skill_token.py       ← Opaque Host Service token minting/validation (§12)
      ├── marketplace/         ← ClawHub + OuroborosHub: `clawhub.py`, `ouroboroshub.py` (hub update via adopt transaction with VERIFIED `rolled_back`/`rollback_errors`, `.pre-adopt` retention disclosure), `fetcher.py`, `adapter.py`, `install.py`, `install_specs.py` (normalized third-party dependency install metadata for bounded per-skill prefixes), `isolated_deps.py`, `provenance.py` (+ publication receipt `state/skills/<n>/ouroboroshub.json`)
      ├── skill_lifecycle_queue.py ← Single FIFO skill-mutation lane + event snapshot (§13)
      ├── skill_lifecycle_actions.py ← Shared grant/toggle effects and owner-action admission for UI, launcher, CLI and task adapters; existing chat/quiz/mailbox source resolution, exact target/revision, no new permission store
      ├── skill_uninstall_state.py ← Marketplace-uninstall tombstones and explicitly authorized local payload/state deletion; separate operations with separate retention contracts
      ├── skill_review_runner.py ← Writes `review_job.json` + `skill_review_*` events; separate review, dependency and process-qualified extension outcomes; unchanged verdict replay can resume dependencies without another panel or overriding owner disable
      ├── server_auth.py       ← Non-localhost network gate via `OUROBOROS_NETWORK_PASSWORD` (warns when unset; see §8 packaging note)
      ├── server_control.py    ← `restart_current_process` + `execute_panic_stop`
      ├── server_entrypoint.py ← CLI parsing + port binding helpers
      ├── server_runtime.py    ← Startup/onboarding wiring + WS liveness
      ├── server_web.py        ← `NoCacheStaticFiles`, web-dir resolver and fixed-source `read_author_kit_assets(repo_dir)` for optional author-owned routes; no endpoint or cache
      ├── server_process.py, server_liveness.py, server_maintenance.py, server_restart.py, server_owner_routing.py, server_routing_context.py ← Server leaves the composition root calls: the facts one server process shares with every leaf; wedge detection for the supervisor generation; the upkeep a generation owes the drive; the restart-adjacent helpers used at shutdown, the owned-work stop of the owner's manual Restart and the planned restart's engine-pin daemon stop; where one owner message goes; and the bounded facts one owner turn is allowed to address
      ├── task_continuation.py ← Durable review continuation state
      ├── task_results.py      ← Durable task results `task_results/<id>.json`; locked `task_acceptance_review_accounting` — the claim is minted at first physical reviewer dispatch, and a claim without a recoverable terminal host run is UNKNOWN, never permission to re-dispatch (double-spend fence); its read-only root review-capacity projection for configured-session wakes is WALLET and cancellation only (`root_task_id`, `cap_cycles`, `claimed_cycles`, `remaining_cycles`, `binding_seen`, `dedupe`, `state`, `reason` — no time axis: the launch rule `task_pacing.review_launch_allowed` is evaluated once per panel at loop admission (owner R55; the paid claim inside the dispatch stamp checks cancellation and the wallet only), and a descendant reads its own window from the coordination `time` fact)
      ├── task_result_schema.py ← Task-result schema admission: the `_schema_version` stamp, the classifier, and the quarantine an unstamped, future, malformed or retired-key row lands in
      ├── task_status.py       ← Effective-status SSOT, lineage, bounded waits; worker-side `task_has_live_queue_ownership` (§10, cancellation custody); the DESTRUCTIVE orphan predicate keeps the same fail-open-toward-liveness polarity — an in-process direct actor (`supervisor.active_activity` registry, deliberately absent from PENDING/RUNNING) and a missing/invalid/stale queue snapshot can never prove a task dead
      ├── git_shell_policy.py  ← Structural git argv classifiers for the shell guards
      ├── protected_artifacts.py ← Execute-only black-box policy for protected artifacts
      ├── shell_parse.py       ← One shell normalization and POSIX wrapper vocabulary (`POSIX_SHELL_HEADS`: sh/bash/zsh/dash/ash) shared by guard and execution (`recover_stringified_argv`, `normalize_check_argv`, `shell_tokens`, `shell_segments`, `canonical_command_text`): quoted shell punctuation is data, not syntax — over-splitting on a quoted `&&` is the fail-safe direction; `split_redirections` is the ONE redirect grammar read by `writer_target_rows`, whose per-segment `(argv, targets, inline_code, unprovable)` facts include bounded shell-`-c` recursion, stdin-program heredocs only when no inline/script operand exists, Python AST targets/independent UNKNOWN, sequential `cd`/`pushd`/`env -C` cwd changes, and non-concrete find/xargs placeholders; uncertainty widens only its own row/body mentions, while the separate mention lane keeps unmangled Windows drive/UNC spellings
      ├── argv_budget.py       ← Argv admission counts encoded bytes of argv PLUS environment (ARG_MAX charges both; per-arg `MAX_ARG_STRLEN`, Windows unit limit); asked by skill_exec before exec
      ├── workspace_executor.py ← Workspace process backends: `local` and network-none `docker_exec`
      ├── deliverables_paths.py ← Lexical + case-folded deliverables path views
      ├── tool_capabilities.py ← SSOT for the core/parallel-safe/untruncated/stateful-browser tool sets
      ├── tool_access.py       ← ToolProfile × ResourceRoot × Operation matrix, affordance map, closed-enum `required_capabilities` check
      ├── tool_access_types.py, tool_access_roots.py, tool_access_paths.py, tool_access_user_files.py ← The access matrix behind that facade: the closed access vocabulary and the profile × root × operation policy matrix; who is acting and where each resource root physically lives; the physical path primitives; and the `user_files` confinement with its secret-name policy and path resolution
      ├── tool_policy.py       ← Round-one tool visibility (the sets live in tool_capabilities)
      ├── browser_policy.py    ← The browser tool's target and control-request policy: task-granted concrete origins, metadata/private/reserved refusals, and the three-valued Ouroboros control-service identity (`runtime_service_kind`: proven kind / unknown / none) that both the URL decision and the owner-operation request shapes consume; `tools/browser.py` keeps the Playwright lifecycle (§6 Web access mechanisms)
      ├── skill_payload_binding.py ← Skill payload targeting: `.seed-origin` distinguishes native vs external; read/list/search only for read profiles; bounded manifestless skill_publish recovery
      ├── utils.py             ← SSOT for atomic JSON, timestamps, hashes, sanitization, subprocess helpers, `truncate_review_artifact`
      ├── world_profiler.py    ← Generates WORLD.md
      ├── contracts/           ← Frozen ABI package (§11)
      │   ├── tool_context.py  ← ToolContextProtocol
      │   ├── tool_abi.py      ← ToolEntryProtocol + GetToolsProtocol
      │   ├── chat_id_policy.py ← SSOT for human-visible vs synthetic chat ids (§12)
      │   ├── task_contract.py ← Frozen task-contract normalization and effective acceptance-claim binding (semantics: §11.1)
      │   ├── task_constraint.py ← `VALID_WRITE_SURFACES` + surface/write_root validation, fail-closed
      │   ├── skill_payload_policy.py ← Payload path resolution/confinement/sidecar detection
      │   ├── skill_manifest.py ← Unified skill manifest parser (`VALID_SKILL_TYPES`: instruction|script|extension)
      │   ├── schema_versions.py ← Opt-in `_schema_version` stamping helpers (§11.2)
      │   └── plugin_api.py    ← PluginAPI, ExtensionRegistrationError, FORBIDDEN_SKILL_SETTINGS, VALID_EXTENSION_PERMISSIONS, VALID_EXTENSION_ROUTE_METHODS
      ├── gateways/            ← Thin outbound transport adapters; no business logic
      │   └── claudexor.py     ← Loopback daemon client: discovery (`discover_daemon_at` reads `<config_dir>/daemon/control-api.json`), handshake, runs, cached quota GET + explicit quota POST; the token stays inside; account-surface translations are read-only; prefers the OWNED daemon via `claudexor_daemon.owned_daemon_provisioned`
      ├── claudexor_runtime.py ← Reviewed engine pin (version/SHA/URL/SHA-256/size/protocol/Node/entrypoints, nullable CLI): seed-or-download, verify + staged extract + probe + atomic promote under `data/state/cx`; no mutable `current` pointer or background updater — the reviewed pin IS the next-spawn selection, and a `null` CLI honestly identifies a pre-CLI closure; read-only resolvers; `OUROBOROS_CLAUDEXOR_BIN` stays an explicit operator override
      ├── claudexor_daemon.py  ← Installation-owned Claudexor lifecycle over `data/claudexor`: lazy first use from worker/server, authenticated attach, purpose-bound startup custody and independent startup/admission waits; `stop_outcome` owns explicit same-home CLI shutdown and measured/Popen fallback (§9). Atomic ownership marker publication, metadata-only runtime status, provisioned warmup through the same ensure, and `install_missing_harness_cli`
      ├── gateway/             ← Gateway Boundary v1: browser-facing route ownership + frontend contract SSOT (see below)
      │   ├── contracts.py     ← Active WS/HTTP envelope contract owner
      │   ├── decision_contracts.py ← Typed request/response contracts for decision families, re-exported by contracts.py; each ingress owns runtime validation
      │   ├── endpoint_index.py ← `HTTP_ENDPOINTS` index (re-exported by contracts.py); routers own the Route objects
      │   ├── schema.py, task_list_scan.py ← The executable gateway contract — JSON Schema derived from the TypedDicts, validating ingress — and the stat-invalidated compact result facts shared by list ordering, SSE discovery and Main routing
      │   ├── router.py        ← Starlette route collector for /api/* and /ws
      │   ├── ws.py            ← WS manager, extension WS dispatch, broadcast
      │   ├── state.py         ← /api/health + /api/state
      │   ├── tasks.py         ← Headless task create/list/get/cancel/events; cancel accepts `stop_policy` (empty = immediate; `finalize_then_cancel` → 202 + open intent → supervisor/owner_stop.py; unknown → 400)
      │   ├── task_events.py   ← Task-event SSE endpoint: legacy GET ranks plus read-only POST v2 physical-chain cursors
      │   ├── task_hurry.py    ← POST hurry ingress: exact one-field `{request_id}` body — extra fields refused, because hurry carries no text by design and a smuggled field must not become a side channel; queue-owned admission atomically initializes an absent pooled lifecycle through task_results.write_task_result(create_only=True), preserving existing rows and excluding direct turns; idempotent projection (semantics: owner_hurry.py)
      │   ├── task_decision.py ← ONE `POST /api/decisions` ingress with family-parsed ids (`quiz:` here, `routing:` → routing_decision.py, `interaction:` reserved); writes `KIND_QUIZ_ANSWER`, broadcasts `quiz_state` (lifecycle: owner_quiz.py; ABI: §11.1)
      │   ├── task_model_wait.py ← Shared model-wait decision effects over the existing mailbox, with live-owner/revision checks, optional role persistence and history-row projection
      │   ├── routing_decision.py ← Validates a click against the durable `needs_manual_target` row, recovers the original text, dispatches the existing `steer_task`/`promote_chat_to_task`, confirms through routing_wait receipts; replay-stable derived identities
      │   ├── logs.py          ← Read-only runtime log tail
      │   ├── onboarding.py    ← `POST /api/onboarding/complete`: install-time latch, shared validation, live engine read, preset compile, single settings write under lock through the shared bounded writer seam; a typed 503 persists nothing (§2) — except 503 `settings_save_timeout` (`saved: null`), the seam's unknown outcome
      │   ├── onboarding_host.py ← GET /onboarding: side-effect-free wizard page served as ES modules
      │   ├── owner_settings.py ← Settings-lock-as-precondition + `CommitBoundary`
      │   ├── settings.py      ← /api/settings + /api/owner/*; `GET /api/reviewer-slots` with row limits (triad 10 / scope 4 / advisory 1 / deep_review 1) and typed `config_error`, never a 500; the response carries the deep self-review singleton — the saved `deep_review` row, or the packed api row synthesized from `OUROBOROS_MODEL_DEEP_SELF_REVIEW` labeled `synthesized_from` so the editor can say it is not saved yet — attached beside a `config_error` too as the legacy-derived repair placeholder, never an effective row (none is effective until the malformed setting is repaired; `deep_review_slot()` raises)
      │   ├── presence_settings.py ← Owner-facing runtime overrides for reviewed Presence behavior skills
      │   ├── control.py       ← /api/reset, /api/command, /api/git/*, /api/update/*, /api/evolution-data HTTP handlers
      │   ├── schedules.py     ← Cron schedule HTTP surface
      │   ├── files.py         ← File Browser + chat upload
      │   ├── ui_preferences.py ← `state/ui_preferences.json`: widget order, per-card widget start-mode overrides (`widget_start_mode`, values from `extension_ui_validation.WIDGET_START_MODES`), nested subagent expansion
      │   ├── models.py        ← Model catalog + provider probes + local-model lifecycle
      │   ├── extensions.py    ← extensions/skills HTTP surface (GET /api/extensions, GET /api/extensions/<skill>/manifest, GET /api/extensions/<skill>/module/<entry:path> — reviewed module sources served from the live loader registration, ALL /api/extensions/<skill>/<rest:path>, POST /api/skills/<skill>/toggle, POST /api/skills/<skill>/delete, POST /api/skills/<skill>/review, POST /api/skills/<skill>/grants)
      │   ├── extension_receipts.py ← Process-qualified extension index/toggle/reconcile receipt projection
      │   ├── widgets.py       ← GET /api/widgets: the Widgets card list projected from the in-memory extension snapshot (live UI tabs + the owning skill's live payload `content_hash` as `revision`) with no discovery, reconcile, hashing or writes on the read path; homes the `WidgetTab`/`WidgetsResponse`/`ExtensionLiveSnapshot` TypedDicts that `gateway/contracts.py` re-exports, importing no transport at module level
      │   ├── skill_publish.py ← Read-only publish preflight with scan cache; one five-state response; no task or GitHub effect
      │   ├── marketplace.py   ← ClawHub + OuroborosHub HTTP surface
      │   ├── mcp.py           ← MCP HTTP surface backed by the shared MCPManager
      │   ├── claudexor_accounts.py ← Agent accounts HTTP surface (Settings → Agents → Accounts): six thin proxies — counted by HANDLER, two serving more than one route or action — over the owned daemon's account truth: GET /api/claudexor/status[?include=models] (side-effect-free daemon/runtime state + harness catalog + credential profiles + quota windows, each facet stamped with its read state; the additive `unified_accounts` feature fact reads the engine's own /v2/operations catalog — `get:account-pools` present = the unified account model, an unreadable catalog fails closed to the legacy rendering); POST /api/claudexor/wake (owner-initiated daemon start); POST /api/claudexor/login (one Connect intent: install/repair the managed runtime, start or attach the owned daemon, create or re-adopt its setup job; a structural `not_supported` missing-binary create response invokes `install_missing_harness_cli` — claudexor_daemon.py owns the whole operation — and creates the same login job exactly once more); GET/DELETE /api/claudexor/login/{job_id} (canonical snapshot/cancel); POST /api/claudexor/login/{job_id}/input (the owner's answer to a waiting engine prompt); POST /api/claudexor/login/{job_id}/reconcile (explicit proof-of-empty after `termination_unconfirmed` — a terminal job is not automatically release proof: custody holds until reconcile records `status=empty`); DELETE/PATCH /api/claudexor/credential-profiles/{harness}/{profile_id} (one route, two row actions, the engine's strict `{enabled}` body passed through). Zero auth logic and zero vendor recipes live here; the browser never sees the daemon token; the `{job, cursor, sequence, deviceCode?}` envelope passes through verbatim; `harness_login_cards.jobDetail()` renders the escaped untruncated message only beside a settled non-success verdict
      │   ├── claudexor_quota.py ← Explicit owner quota-refresh transport: POST /api/claudexor/quota/refresh discovers the already-owned daemon, performs the mandatory handshake (ordinary 60 s control-plane read bound), and delegates exactly once to the engine's quota POST (90 s foreground bound); the envelope returns verbatim; no lifecycle start, cached status composition, quota policy, retry, or daemon token crosses this boundary; GET /api/claudexor/status stays passive
      │   ├── host_service.py  ← Loopback-only Host Service API (§12)
      │   ├── history.py       ← Chat history + cost breakdown factories
      │   ├── cost_breakdown.py ← Ledger-derived dashboard buckets and root-task detail breakdown over the same physical-attempt authority; `history.py` and `tasks.py` keep their historical import seams
      │   ├── projects.py      ← GET/POST /api/projects, /from-task, /update, /delete
      │   └── _helpers.py      ← Shared request-root/coercion/JSON error envelope and `run_sync_to_completion`, the settled worker wait for request-owned blocking work
      ├── tools/               ← Auto-discovered tool plugins (registry.py owns discovery; frozen module list for packaged builds)
      │   ├── registry.py      ← Tool registry SSOT: loads tool modules, exposes schemas, executes safely; owns the shell-guard/process-tool membership sets
      │   ├── core.py          ← File/data tools (read_file, write_file, list_files) + code search and digest helpers
      │   ├── core_file_tools.py, core_secret_paths.py, core_artifacts.py ← Core-tool leaves: the read/list file tools with the shared resource-access helpers, and the verbs a task uses to put something in front of a human; `core_secret_paths.py` owns the restricted-subagent physical read-denial policy (owner secrets/control state, child and canonical data roots, repository credential locations, listing redaction); list/search/query prepare shared locations once per call while target resolution and file identity remain live
      │   ├── shell.py         ← Process tools `run_command`/`run_script` (in-process `_active_subprocesses` tracking; §9)
      │   ├── shell_guards.py  ← Shared process-tool inspection paths and shell-guard helpers (re-exports write_shape)
      │   ├── registry_core.py, registry_guards.py, registry_guard_process.py, tool_context.py ← The registry's leaves: the execution authority (load modules, expose schemas, dispatch safely); the host-owned pre-dispatch guards (capability/resource, ephemeral, managed-update and skill-payload constraints); and the process/shell guard helpers — self-change tripwires, the `_is_pure_read_inspection` HEAD allowlist with option denial, and light-mode repo snapshots; and `ouroboros/tools/tool_context.py`, the concrete `ToolContext` + `BrowserState` imported by `registry.py` and `registry_core.py` (the protocol it satisfies is `contracts/tool_context.py`)
      │   ├── git.py           ← Git/write tools with the advisory, triad, and scope review commit gates (§8)
      │   ├── git_plumbing.py, git_repo_edit.py, git_vcs_ops.py, git_review_cycle.py, git_evolution.py ← The git tool's leaves, re-exported by the facade: the low-level plumbing shared by the git owners; the uncommitted repo write and exact-match edit surface; generic VCS inspection and rollback; staging plus the advisory/triad/scope review and reviewed-material binding the commit gate consumes; and evolution-campaign authority at the reviewed-commit and publication boundaries
      │   ├── search.py        ← Web search tool (OpenAI Responses API, LLM-first overridable defaults)
      │   ├── browser.py       ← Playwright browser tools with per-ToolContext lifecycle and thread affinity (§6)
      │   ├── vision.py        ← Vision LLM tools for browser screenshots and uploaded images
      │   ├── vision_process.py ← Tracked vision-child IPC: validated model/physical-attempt receipts, result recovery and parent-owned cancellation
      │   ├── knowledge.py     ← Persistent topic-based knowledge files with an auto-maintained index
      │   ├── memory_tools.py  ← Memory registry tools for tracking data sources, gaps, and trust
      │   ├── health.py        ← Codebase health tool: complexity metrics and self-assessment
      │   ├── compact_context.py ← LLM-requested tool-history compaction trigger; stores the pending request for the next round
      │   ├── control.py       ← Control tools: restart, timeout settings, scheduling, review, chat history, model switching; publishes the strict `subagent_id`+objective `schedule_subagent` contract, and `wait_task` emits the burst/absorb advisory and compact wait projections (§6)
      │   ├── control_delegation.py ← Delegation-budget and in-task project-scoping affordances (`ensure_project_scope` handler)
      │   ├── control_events.py, control_routing.py, control_runtime.py, control_scheduling.py, control_subagent_spec.py, control_task_results.py ← The control tools' leaves: emitting one control event and waiting for its durable handler outcome; routing real work out of a conversation lane into a supervised task; runtime self-control (restart, promotion, evolution, memory, model); scheduling one live subagent — what the parent asked for and nothing else; the published `schedule_subagent` parameter surface and its validation; and absorbing a child — reading one result, or waiting on a batch of them
      │   ├── tool_discovery.py ← Tool-discovery meta-tools: confirm registration, no delayed capabilities
      │   ├── tool_result.py, tool_catalog.py, tool_resolution.py ← Dispatch-side vocabulary: the typed internal tool result with its byte-compatible legacy text adapter (the producer stamps outcome facts; a note append degrades rather than rewrites); the intrinsic tool descriptors shared by tool modules and registry dispatch; and argument normalization with physical target binding
      │   ├── review_response.py ← Pure response-envelope projection for multi-model review rows
      │   ├── shell_process.py, shell_effects.py, shell_outputs.py ← The command-running substrate: the process execution shared by every command tool; what a command did to its working tree and which of it was throwaway; and declared process outputs — resolution, fingerprints, artifact registration, and the per-path export-eligibility rules that reuse the workspace-patch credential-shape SSOT
      │   ├── plan_review_artifacts.py ← Exact plan-review waves and full operative specs in existing source handles, bounded successor index, and reviewer-continuation inputs; reconstructs the API transcript
      │   ├── evolution_stats.py ← Generates evolution.json metrics from sampled git history
      │   ├── owner_delivery.py ← Owner event delivery: live queue XOR `pending_events` fallback with sticky deferral, preserving narrative order after the first live failure; lineage stamped; background-consciousness frames unstamped and deferred
      │   ├── deliverables_shell.py ← cp/mv/ln into deliverables with symlink checks
      │   ├── shell_audit.py   ← Post-exec custody audit for process tools
      │   ├── process_facts.py ← Typed process-fact seam consumed by loop_tool_execution for the same call; the regex harvest stays a read fallback
      │   ├── write_shape.py   ← `interpreter_write_shape`/`non_interpreter_write_shape` — the write-shape classification SSOT
      │   ├── extension_dispatch.py ← Extension tool dispatch (contracts preserved; discovery stays in registry.py)
      │   ├── release_sync.py  ← `sync_release_metadata` (version carriers) used by commit-admission preflight; `_preflight_check` uses `check_history_limit`; agents may call it directly; the carrier-span SSOT (`VERSION_CARRIER_SPANS`, `substitute_carrier_spans` — the ONE span primitive the managed-update resolver and the commit-triad pack cut share — and the `carrier_only_change` predicate: a carrier changed only inside its declared version spans)
      │   ├── review_synthesis.py ← Shared synthesis helpers; the parser/aggregator lives in plan_spec.py
      │   ├── ci.py            ← CI trigger/monitoring
      │   ├── claude_advisory_review.py ← `preflight_review` with the callable `advisory_review` compat alias; admission policy lives here; api_chat rows ride review_native_episode, agent_session rows ride the AgentSessionReviewExecutor; a native episode that ends on its own transcript bound reaches the caller as the typed non-blocking `ADVISORY_SKIPPED: native_transcript_bound_exceeded` (keyed on the episode's structured `native_transcript_cap_exceeded` code, never on message text; not the provider window vocabulary) carrying the bound, the refused chars and the paid rounds, and every episode exception keeps `failure_custody()` as the advisory meta's `usage` (never an empty `{}`); on the native route the documents the prompt's MANDATORY FULL READ pointers name are measured from the files at prompt-build time (`preflight_review_prompt._mandatory_read_corpus_chars`, wire chars) and declared to the episode as its mandatory reading, so the bound is lifted to hold them when the advisory model's window allows and otherwise the prompt's MANDATORY READ budget section and the usage both carry the typed `native_mandatory_read_exceeds_bound` — never a silent full-read contradiction
      │   ├── preflight_review_prompt.py, preflight_review_run.py ← The preflight (advisory) pre-review: prompt assembly with its git captures, and the run with its result parsing; its changed-context pack applies the shared span-only release-carrier cut over the pair the advisory reviews — HEAD→working tree, the text the pack reads — with the same `PACK EXCLUSION NOTE` (the native episode keeps `read_file` for a withheld carrier)
      │   ├── recent_tasks.py  ← Read-only context recovery
      │   ├── commit_gate.py   ← Commit gate: `_record_commit_attempt` (LLM claim synthesis), `classify_review_block`/`attempt_block_class`, `check_identical_verdict_refusal`, `count_paid_review_cycles`/`check_review_cycles_ceiling`, `commit_review_contract_fingerprint`
      │   ├── git_rollback.py  ← Wraps `git_ops.rollback_to_version`
      │   ├── git_pr.py        ← Five PR tools (non-core)
      │   ├── github.py        ← Issue + PR tools (frozen tool module): shared process binding selects the active Project; explicit repo flows through every subcall; missing implicit Project targets fail visibly. Discovery reads token sources or native CLI configuration without an authentication probe; explicit Hub/API transport retains its own target. `_gh_run` is the structured transport read (exit code, bounded redacted stderr head, observed HTTP status, failure class) that `_gh_cmd` projects onto the string ABI; the transport publishes only its own target refusals (`GH_TARGET_INVALID`/`GH_TARGET_REQUIRED`) into the tool-result sidecar; a process outcome publishes nothing there, because the publication transaction owns its own final result.
      │   ├── parallel_review.py ← Triad + scope review orchestration: assembly of both packets, the money admission call (`review_admission.py`), the scope-first hold, and the executor transitions that carry the admitting usage scope (`contextvars.copy_context`) into every seat
      │   ├── plan_review_references.py ← Reference projection that also writes its own provenance rows (`append_jsonl` into `logs/progress.jsonl`, `emit_log_event`), never a second plan authority
      │   ├── plan_review.py   ← `plan_task` engine: evidence, packet, fan-out over the review substrate, `plan_review_state` v2, the shared `OUROBOROS_REVIEW_MAX_CYCLES` cap, free identical replays; no scouts, Atlas, or plan_class
      │   ├── plan_review_runtime.py ← Plan-review deadline rail, `ReviewSlot` rows, `plan_slot_fit` + `preflight_oversize`, health snapshot + `plan_wave_replay_decision`, `plan_review_advisory_open` emitter
      │   ├── plan_spec.py     ← Pure plan-spec parsing/aggregation (`resolve_constitutional`); no I/O
      │   ├── plan_evidence.py ← Bounded plan-evidence manifest; the runtime data plane is denied
      │   ├── plan_packet.py   ← Reviewer packet; the W3 governance pack inlines BIBLE + ARCHITECTURE in full for self-modification plans, nav maps otherwise
      │   ├── plan_render.py   ← Wave view + `PLAN_REVIEW_CONTROL_JSON` footer; no independent behaviour
      │   ├── review.py        ← Acceptance review + multi-review adapters
      │   ├── review_multi_model.py, review_file_pack.py, review_prompt_text.py ← Commit-triad leaves: the sync/async multi-model fan-out entry with its per-row dispatch through the review substrate and their shared limits; reviewable file classification and the packs read from the working tree (the touched pack takes the caller's `exclude_paths` — the advisory seam's shape — and marks a withheld path once; `span_only_release_carriers` is the ONE carrier-cut predicate the advisory, triad and scope packs share — release carriers changed only inside their declared version spans on a VERSION-bearing change, over the pair each pack reviews — and `pack_exclusion_note` its one disclosure (`CARRIER_CUT_REASON`); `triad_pack_exclusions` adds the commit triad's second class, governance docs byte-identical to the inlined prefix copy, with the `PACK EXCLUSION NOTE` the call site appends after the OMISSION NOTE; a managed subject keeps every full text); and the fixed reviewer prompt vocabulary plus the sections built from prior rounds
      │   ├── review_context_atlas.py ← Repository atlas for scope_review + deep_self_review (plan review does not consume it)
      │   ├── query_code.py    ← Read-only code intelligence; `root=user_files` with path guards, denied to subagents
      │   ├── edit_ops.py      ← `apply_patch` + `edit_batch` with shared `_syntax_check`/`_unified_diff` backing write_file
      │   ├── media.py         ← `ocr_pdf`, `youtube_transcript`, `extract_video_frames` (dependency-optional, typed capability envelopes; frames under `artifact_store/video_frames`)
      │   ├── verify.py        ← Independent check execution through the SAME pre-exec guard machinery (shell-guarded, deliberately NOT a process-command tool); receipts append to `<drive_root>/task_results/artifacts/<task_id>/verification_receipts.jsonl`; `expected_match` kinds: substring (default) · exact · exact_line · json_equals · bytes_equal; an owner-settings change is detected and reported as a typed note, never auto-reverted — an auto-revert would undo concurrent differences without proving causation, and POST-execution checks cannot gate a receipt already written, so verify rides the PRE-execution guards; verification reads only public task info (anti-cheat); the exit-masking sensor feeds the advisory nudge without changing status, and `run_command`/`run_script` read the SAME sensor to append ONE advisory note plus `exit_masking_reasons` to a masked GREEN result envelope — status, `is_failure` and the reported returncode unchanged, no receipt written, outside the verification ledger and receipt reconciliation, using the shared typed shell lexer so subshell grouping does not hide masking and quoted operators remain literal arguments; `delegation_zero_run` writes only `incomplete`/`unknown` and only after the custody scan proves no open run — a self-reported "complete" with zero runs is unverifiable authority
      │   ├── review_helpers.py ← Shared review helpers (governance-doc loading, checklist section slicing, prompt-size SSOT, the density-calibrated input cap and the probe sample `density_probe_sample` — a slice of the real atlas content — every packed review surface measures on)
      │   ├── review_binary_context.py ← Staged/parent Git object metadata for review packs
      │   ├── review_subject.py ← `ManagedReviewSubject`; `capture_review_diff` stays byte-identical for non-managed callers
      │   ├── review_admission.py ← Pre-dispatch admission for both packets ($0 on a deterministic block): `fit_triad_prompt`, typed `not_dispatched` seats, typed oversize outcome; the commit gate's money admission (`commit_gate_paid_seats` prices every paid seat, `admit_commit_gate_wave` admits the wave whole through `review_wave_budget_gate`); the commit gate's cold-start density rung (`density_probe_before_size_refusal`, the packed deep self-review's rung shared) before either packet is refused or degraded for size on a cold store
      │   ├── review_revalidation.py ← Review-contract fingerprint revalidation
      │   ├── scope_review.py  ← Enforcement/budget-aware whole-repo scope reviewer (§6 Review stack)
      │   ├── scope_review_session.py ← Session scope delivery from the SAME `build_scope_review_prompt` (canonical docs as nav maps); coverage manifest is forensics, never a gate; session admission per BIBLE P3 (≥200K sourced window)
      │   ├── scope_window.py  ← `scope_window` resolution + ReviewerWindow constants
      │   ├── scope_review_contract.py ← Pure scope-item parser (`normalize_scope_items`); also consumed by scripts/validate_scope_receipt.py
      │   ├── scope_review_pack.py, scope_review_budget.py ← Scope-review assembly and its money: the touched context (a span-only release carrier on a VERSION-staged commit keeps no snapshot by design — `_carrier_span_only_paths`, the shared cut over the same HEAD→index pair the pack reviews: named in the dedup note, declared diff-only to the atlas with that reason so the durable coverage row stays truthful, traced as the ladder's first entry; never for a managed subject or an artifact the atlas owes in full), atlas and guaranteed-fit ladder, and the token limits, reserves and oversize classification that bound them
      │   ├── services.py      ← Service mini-manager with process-group cleanup
      │   ├── skill_exec.py    ← list_skills/skill_review/toggle_skill/skill_owner_action/skill_exec; runtime allowlist python/python3/bash/node/deno/ruby/go; gated by enablement + fresh review + hash
      │   ├── skill_publish.py ← Thin publish transaction over the four leaves; success is PR-receipt-gated
      │   ├── skill_preflight.py ← Read-only skill preflight; a module widget's declared `render.entry` is checked for existence and containment inside the skill directory, and parsed with classic-script grammar because the widget frame runs it as an inline script
      │   ├── project_journal.py ← journal_write/read, workpad_read/write, journal_tail_digest (over-limit rejected); owns `mirror_tree_coordination_to_journal`, the durable-journal mirror of tree coordination
      │   ├── presence.py      ← configure_presence, initiate_presence, typed completion/cancel
      │   ├── task_tree.py     ← tree_note/tree_read (storage SSOT: task_tree_ledger.py)
      │   ├── followup.py      ← One deferred follow-up into `state/scheduled_tasks.json`; exactly one trigger (`once` ISO or 5-field cron+tz); cap 2 pending, over-limit refused; preserves the SOURCE address (top-level `project_id` so `resolve_project_id` sees it, plus the originating `chat_id`) instead of defaulting to the global owner chat — an unscoped task's follow-up still takes the existing `owner_chat_id` default. The existing scheduler consumes a one-shot on the typed tombstoned-Project refusal, retaining its failed task and last error; deleting Projects and transient refusals remain retryable
      │   ├── join_ledger.py   ← Child-result absorption: validates lineage + exact hashes; dispositions integrated/irrelevant/deferred; `CHILD_RESULT_STALE`; keeps peek_task/discard_child_result
      │   ├── delegate.py      ← Delegation facade verbs `delegate_start` (with `retry_of`), `delegate_wait`, `delegate_cancel`, `delegate_answer`; the host pre-start rides the same `delegate_start(prompt="")` wrapper and the shared `subagent_runtime.exact_start`; supervision/recovery/custody/transport live in the named leaf modules (§6)
      │   ├── delegate_integration.py ← Delegated-patch integration: `_mutation_authority`, `_provision_snapshot` (registered before the start intent), retry-binding validation, `_capture_terminal_patch`; the skill-payload cluster (`_payload_mutation_authority`, `_rebind_payload_reference`, `_write_payload_patch_artifacts` — git diff --binary; reserved paths refuse the WHOLE apply as `blocked_reserved_paths` with the candidate preserved) and `integrate_payload_patch` (CAS, index-free git apply in NO-REPOSITORY mode — `GIT_CEILING_DIRECTORIES` pinned at the resolved PARENT of the payload, because git still searches the ceiling entry itself, so an ancestor Git worktree above the runtime data root cannot make git treat the payload as a subdirectory prefix and silently skip every hunk at rc=0 while `--numstat` prints nothing; the same env bounds the `--numstat -z` touched-path reader; a post-apply live hash equal to the BASELINE with a non-empty touched set is refused typed as `INTEGRATE_APPLY_NO_OP` with the apply intent resolved and nothing disposed; the busy-check refuses a second delegation only while a run whose OWNER TASK IS STILL LIVE, or whose terminality cannot be proven, has open custody on the same payload, QUEUES the extension reconcile request via `request_extension_reconcile`)
      │   ├── delegate_payload_patch.py, delegate_terminal_evidence.py, subagent_integration_delegated.py ← Delegation leaves: the skill-payload patch pipeline (capture artifacts, guards, the live apply); the terminal story of ONE delegated run as the parent reads it; and `integrate_delegated_patch`, the delegated-run disposition seam
      │   ├── subagent_integration.py ← integrate_subagent_patch (sha256, 3-way --index, protected-path gated, genesis refused), external-workspace audited verdict, `coop_already_in_tree` no-op, compare_subagent_patches
      │   └── patch_verdict.py ← The ONE verdict writer for both patch pipelines (re-exported as `_write_verdict`): verdict subjects are minted and classified by the writer (`run_<rid>`), never prefix-matched by readers; each decision lands twice — artifact + typed `delegate_run_patch_verdict` custody row — so the acceptance packet reads one replayable store; a failed artifact write is disclosed on the row
      ├── delegate_start_claims.py ← One short pre-transport transaction serializing the zero-run/custody recheck + `START_REQUESTED` append; the nested payload claim is taken only when selected; transport and waiting stay outside claim locks
      ├── process_containment.py ← Env-token container membership (`OURO_PROC_CONTAINER_*`; /proc environ on Linux, `ps -E` on macOS, kill-on-close Job Object on Windows) with live-state read at reap — containment is unconditional because a surviving descendant can become invisible to ordinary parent-child traversal once the controller exits, and Windows spawns suspended-then-adopt so a child cannot execute before Job membership takes effect; an attributed alive-or-undeterminable member is an honest hard-block answer, never a kill guarantee; unreadable strangers are warnings, never members by uid/start-time alone (an unobserved detached descendant that hides its token remains a disclosed detection gap); `process_group_has_live_members` excludes zombie-only service groups while retaining live children and unknowns; policy layered over platform_layer
      ├── process_custody.py   ← spawn_supervised + durable process_ledger.jsonl; reap_orphaned_processes checks strict identity and retained_purposes across generations; start_parent_lifeline watches the spawner; quiesce_custodied_services checks all group members. live_daemon_root_pids/live_kept_service_pids select teardown exclusions; process_stop_snapshot binds asynchronous stop fallback to the observed rows; stop_ledgered_processes requires measured identity and confirmed exit. Complements platform primitives and existing panic tracking (§1 Runtime topology; §9 Shutdown).
      ├── platform_layer.py    ← Cross-platform process helpers, the descendant-enumeration seam, the Windows Job Object ABI
      ├── verified_download.py ← Shared exact-size/digest verification and atomic cached byte delivery; consumers retain their own transport timeout and error vocabulary
      └── node_runtime.py      ← Execution-probed Node runtime health (`node_runtime_health`, memoized by path/mtime/size — a missing binary is never cached, and a timeout verdict is re-probed only by a larger budget), `select_skill_node_runtime`, `skill_node_emergency_path_dir`

Devtools boundary

devtools/ (including devtools/benchmarks/cybergym/) lives outside the runtime and package discovery: no runtime imports, normal review, artifacts in an external output root; a sentinel-marked isolated root suppresses rotation warnings. devtools/e2e_live/ is the live E2E stand: K staggered isolated real servers running the owner-shaped scenarios SM1 (a design-system-consistent brand-accent change landed as a reviewed release, the product's own review policy shaping the work), SW1 and SK1 with acceptance over durable artifacts and a browser probe, admitted through the same seed gate and manifest seams as the benchmark launchers (DEVELOPMENT "Live E2E stand").

Gateway Boundary v1

ouroboros/gateway/ is the single inbound browser/CLI boundary: contracts.py owns the envelopes (with the endpoint index in endpoint_index.py), router.py collects the routes, and files.py/host_service.py stay separate trust boundaries. The ouroboros/contracts/api_v1.py compatibility re-export is gone with the envelope aliases, so gateway/contracts.py is the one envelope owner; the contract is also EXECUTABLE — gateway/schema.py validates ingress against JSON Schema derived from those TypedDicts. Domain handlers translate transport into calls on existing runtime owners and must not acquire a second copy of queue, review, settings, or lifecycle policy. The facade exists for dependency direction: the UI evolves without importing the agent body, and the runtime evolves without ad-hoc browser contracts.

gateway/owner_settings.py is the ONE owner-scoped settings WRITE seam (the generic POST, the single-decision endpoints, and onboarding; membership = calling _owner_update_settings, directly with a transform or through _owner_write_settings with a whole document). The settings lock is a precondition — a timed-out acquisition refuses before any precondition or write — and CommitBoundary marks the commit instant so a later-step failure is reported as that step, with saved a field on BOTH sides of the boundary: an envelope that merely omits the field would be ambiguous. The invariant prose lives in the module docstring.

Frontend calls go through web/modules/api_client.js with the JSDoc mirror web/modules/api_types.js; gateway parity tests pin the mirror. Extension HTTP lives under /api/extensions/<skill>/… with namespaced WS dispatch.

CLI / Headless Boundary

ouroboros.cli is a client of the same gateway/queue — no second task engine. Its parser is the command-surface SSOT (server, run, tasks, chat, logs, evolve, schedule, settings, skills, marketplace, local-model, MCP); streaming commands reserve stdout for the final answer/patch/result/JSONL and send progress to stderr.

POST /api/tasks creates an ordinary managed root; GET /api/tasks is a non-materializing list; GET /api/tasks/<id> returns the effective durable result; /events is the archive-aware SSE stream (§3 Chat); /artifacts/<name> serves simple filenames confined to data/task_results/artifacts/<task_id>/ — a stored arbitrary path is not a download capability. The CLI refuses any delegation_role other than root, the gateway rejects caller lineage/subagent labels, and only schedule_subagent creates children. Reserved service metadata is written after caller metadata. Admission reserves the task id plus a worker-pool slot under one queue lock; a failure rolls back only the token-owned row with a loud typed refusal. The parsed request's complete admission and attachment staging run off the HTTP event loop through gateway._helpers.run_sync_to_completion; a cancelled HTTP waiter retains that worker until durable admission or rollback settles, without cancelling the admitted task. Detail reads and SSE terminal materialization use the same settled wait, and v2 closes its row iterator only after the outstanding read finishes. Attachments are copied into the effective task drive before enqueue; artifact-store references are not host-path authority.

Workspace tasks default memory_mode=forked; shared is rejected for an external workspace and materialized on a forked child drive for project scope — the stored memory_mode reports what was requested while drive_root reports where the task executes, so isolation does not depend on relabelling the request.

--detach returns only after durable admission; --no-stream polls to completion; waiters treat a result terminal only after the artifact state leaves pending/finalizing, and an explicitly-partial cost gets a bounded 60-second finality wait before partial flags stay visible. ouroboros run exits 0 only for a completed lifecycle, a clean execution axis, no failed/degraded objective, and a finished artifact bundle — strict exit semantics keep shell automation from interpreting "the model answered" as "the requested workspace deliverable exists"; --patch/--patch-out are stricter still (failed/missing patch, no-change, empty payload, or unfinished finalization is an error).

CLI schedules and skill-manifest schedules enqueue ordinary supervisor tasks — no parallel scheduler. resync_skill_schedules() mirrors manifests (executable skills with supervised-task permission only) into the same table after lifecycle changes and on ticks; a blank timezone means the DST-aware system zone, with a fixed-offset fallback only when the zone is unrecoverable; the active schedule digest rides task/consciousness context.

Packaged CLI artifacts are a tiny wrapper + installer, not a second PyInstaller runtime: packaged_cli locates repo.bundle, its manifest, and python-standalone, bootstraps the launcher-managed repo, and invokes the same ouroboros.cli under the embedded interpreter with canonical env. Packaged server is refused — it would bypass launcher-owned bootstrap, process identity, restart, and cleanup. run --start is loopback-only, starts the desktop app when no ready gateway answers, follows data/state/server_port, and waits for /api/health + supervisor_ready. Nested AppImage extract-and-run gets a private TMPDIR because the type-2 runtime keys extraction by TMPDIR + digest; a marker-gated AppRun custodian removes the verified extracted child after the launcher exits.

Release builds also carry Node and ripgrep. Skill-side Node resolves bundled-first via node_runtime.select_skill_node_runtime() (rolling back to a healthy PATH node when the bundled candidate fails the health probe), while the four generic process launch surfaces run the opposite policy (process_interpreters.resolve_process_node): a PATH candidate that passes the execution health probe stays byte-identical in argv and child env, and the bundled runtime substitutes only when that candidate is missing or probe-dead, attesting the child-env PATH prepend — never inside a non-local executor backend. The Node downloader verifies the official archive against published SHASUMS and the macOS signing pass re-signs it under the hardened runtime; ripgrep is archive-hash verified, and search_code still pre-enumerates policy-approved files before invoking it, so bundling a faster binary does not widen search authority. Every bundled consumer searches bundled_resource_bases(): OUROBOROS_BUNDLE_DIR → frozen root → interpreter-ancestor roots → source checkout — server/CLI children run from the managed repo with neither _MEIPASS nor an in-bundle module path, and ancestor recovery covers older launchers starting newer checkouts.

The embedded interpreter must never write into the signed application (codesign seal): entry processes suppress bytecode before project imports, and embedded_python_env() redirects bytecode to data/state/pycache and user installs to data/state/python-userbase (pip_install_target_args() adds --user only for python-standalone). Disclosed residual: the userbase outranks bundle site-packages and nothing prunes or versions it — recovery is manual (remove the directory, relaunch).

Workspace binding changes the contextual repo, never the system repo for BIBLE/prompts/review governance. /api/tasks and project-room promotion share workspace_admission.validate_workspace_root() (exists, exact worktree root, disjoint from the system repo and data drive, resolved/bidirectional/case-folded); an empty Project binding is idempotently provisioned as a standalone git repo unless workspace="none". Binding changes the default file/process/VCS target plus memory/lease/preflight/finalization; it does not remove top-level tools or downgrade the Architecture context in Max mode (root=system_repo stays; root=skill_payload takes bucket+skill_name). The workspace executor is a process-routing boundary: executor_ref is host-owned, mappings must cover the workspace without overlapping system repo/data, network=none only when the backend implements it, and executor processes enter durable custody records.

Workspace preflight snapshots git state (bounded porcelain rows), manifests/scripts, and tool availability into the full workspace_preflight.json artifact with a bounded summary in metadata; tools_on_path/tools_missing_from_path are named that way because shutil.which measures PATH presence, not executability, and the structured keys stay frozen because they ride durable, replaying task metadata; a collection failure is a disclosed error summary, never a fictitious full artifact.

Completion compares against the captured preflight base (task-local commits stay in the delta, not git diff HEAD); the patch is bound to task_constraint.base_sha; a moved HEAD fails closed only for self_worktree (a shared tree relies on reverse-patch verification); an unborn repo diffs against the canonical empty tree. Patch capture streams the tracked binary diff plus admitted untracked files, excluding scratch/cache/junk/incidental-lockfile entries with per-file reasons; otherwise eligible oversized/binary untracked outputs ride complete manifest+zip file artifacts, and tracked files whose old or current size exceeds 50 MiB stay in the same file-reference manifest (deletions explicitly name the baseline and absent current file) instead of a giant Git patch; generated output (dist/, build/) is governed by the project's own .gitignore, honoured through --exclude-standard, not by a host name rule — git-ignored files are outside the capture universe and are not listed as exclusions; a sensitive-looking untracked credential is excluded per-file and disclosed as sensitive_blocked. workspace_patch.json is written for EVERY workspace finalization (including no-change and failed) and is the truth source for CLI strict-patch — it distinguishes omitted vs no-op vs failed; workspace.patch exists only for ready_with_changes.

Forked/empty task state lives under data/state/headless_tasks/<task_id>/data: a forked drive copies identity.md, WORLD.md, registry.md (a project fork carries memory/knowledge/patterns.md and omits global knowledge; an empty drive starts blank); dialogue, scratchpad, mailbox, and history never cross. The child drive is execution state: the result copies back to the canonical root, declared artifacts rebase to data/task_results/artifacts/<task_id>/ (missing source = copy failure; collisions get a deterministic suffix), and verification-receipt replicas union with exact-row de-dup. Once the canonical result is terminal, late copy-back and effective reads pass through the same pure field-custody projection — the parent-owned terminal marker and cost/round/token fields cannot be overwritten. memory_export.json is an explicit artifact, never merged automatically.

Complete input sets are preserved in the existing artifact store. task_contract.attachment_manifest is a preview of at most 25 rows; an additive attachment_manifest_ref names the full immutable JSON under source_handles/context_checkpoints with count, size and SHA. Pure contract normalization preserves this shape; inheritance, owner-mailbox reads, physical retries and copy-back resolve the complete file closure through artifacts, verifying captured bytes before reuse. Legacy inline lists remain readable. Ordinary artifact and input copy failures share child_ref_promotion pending custody, retry and cleanup protection for both headless and direct-task drives. A preview never substitutes for an unreadable full reference. Native-image bounds and transport-specific Telegram limits remain separate from complete work-order delivery.

Automatic genesis output listing is discovery, separate from artifact custody. Unreadable directories and changing files remain explicit rows; complete and gap_count disclose coverage and a bounded artifact error note survives projection. An available listing with gaps does not turn an otherwise successful capture into FAILED. Actual copy/ZIP capture still requires stable source bytes, descriptor and path identity and any expected digest; growth beyond initial regular-file size fails immediately instead of waiting indefinitely for EOF.

The two router continuation tools require an explicit predecessor_task_id ("" = fresh; omission or null refused before lookup, enqueue, or spend); predecessor source metadata survives snapshot/restore, and Main receives a defensive provider-only copy (inline threshold, persisted narrative, bounded legacy row, or an explicit gap — never a raw head/tail substitute), while exact reads and work orders stay full.

System self-modification, external workspace, and genesis remain distinct task classes; a genesis child gets an inspectable deliverable_manifest.json listing with streamed sizes and hashes for readable, stable regular files and explicit gap rows otherwise; symlinks stay un-followed. Listing gaps do not fail the task, while actual artifact copies still require verified stable bytes. Global/system installs stay runtime-policy reviewed, and sudo is always non-interactive (sudo -n).

Immutable artifact identity is owned by the existing artifact-record merge: collection, effective results and manifest refresh retain the captured size/SHA and disclose changed bytes. Rebased immutable files preserve that identity and name; an already verified canonical copy survives a later child change. A failed copy remains a per-file failure and a pending ref in the existing child-copy projection, keeping both GC roots until the file is promoted, while other files materialize; the existing artifact-bundle owner reports failed/missing and public, routing and terminal-event projections use that aggregate. The stored capture lifecycle remains distinct so pending-ref retry can rebuild the bundle after successful copying; no artifact observation rewrites task lifecycle, price or objective. Mutable outputs retain their normal refresh/version behavior. The task-artifact endpoint is a synchronous Starlette route so materialization, private-source reads and full-file verification run in its existing worker pool, including HEAD requests.

Startup file recovery follows prior-process custody and precedes the actual drive-prune pass; unknown ownership or unresolved/protected sources defer task-source pruning for that pass. GC removes a headless child drive only when the canonical parent is terminal, artifact finalization is terminal, retention has elapsed, and the recorded child path matches the expected directory — everything needed after child-drive deletion must cross the canonical handoff before a task is presented as settled; canonical results, artifacts, genesis repos, and memory exports survive.

Runtime topology

Two continuity roles: launcher.py owns the PID lock, bundle bootstrap, the server process, presentation, the restart signal, and cleanup (the packaged launcher runs outside the managed repo); server.py is the self-editable inner runtime. Native packages ship an opt-in systemd user unit as an alternate ingress, not a third role — deliberately without a restart policy, because the launcher owns managed restart, the crash fuse, and panic-to-complete-stop (KillMode=control-group).

Spawn custody: POSIX children start in a new session/process group; Windows creates the server suspended, assigns a kill-on-close Job, then resumes — failure to establish Job custody refuses to run. The launcher Job permits explicit breakaway; only the shared daemon requests it, while ordinary generation children remain covered by Job close. Lazy worker and direct-server starts use the same platform helper. An old packaged or external Job that cannot confirm breakaway retains the previous working spawn with a disclosed survive-close limitation; managed source updates cannot replace the immutable launcher. The launcher records data/state/server_process.json (PID, pgid, server/repo paths, requested and actual ports, argv, creation time) and re-proves identity before cleanup. Forced tree/group cleanup excludes the shared daemon subtree; existing listener sweeps retain their separate scope.

Same-install reaper (launcher_server_reaper.py): holding the PID lock licenses the reap, which runs at main() preflight and at the top of every launcher generation. A PID is proven only on three live facts — the exact <REPO_DIR>/server.py argv token, OUROBOROS_DATA_DIR, and OUROBOROS_MANAGED_BY_LAUNCHER=1 — revalidated immediately before the signal, with descendants captured before the root signal and the whole pass bounded to three rounds. Kills require the byte-exact /proc environment: ps -E output never authorizes a kill, because argv is indistinguishable from an env assignment there, so non-/proc hosts stay report-only. Server selection does not require a custody row — missing server records are the defect being repaired. Caller-supplied retained daemon roots filter descendants before final root revalidation. POSIX-only; Windows generation children die with the launcher Job while the shared daemon breaks away; never on panic or window-close. The startup stray check is report-only and annotates same_install/foreign.

Durable process custody (ouroboros/process_custody.py): spawn_supervised() records every long-lived child in data/state/process_ledger.jsonl — {pid, pgid, fingerprint{start_time, cmd_sha256}, purpose, scope task|session|daemon, owner_task, session_id}. The custody reaper runs at server startup and on the 10-minute supervisor tick and kills only entries whose generation or task owner is gone, by STRICT fingerprint — never by command-line class, so dev and packaged instances can coexist. The POSIX start-time fingerprint is a downgrade-safe /proc-first + ps -o lstart= pair (a rollback meeting an unknown token would prune every row WITHOUT a kill, orphaning processes; the mint order and platform details are specified under "Platform substrate" below). A recorded bare tick never authorizes a kill. Current-generation session processes take a cheap non-zombie liveness check that only ever KEEPS; daemon entries and explicitly retained legacy installation-process records are kept, with skill companions the exception — reaped on owner-uninstall or foreign generation, log-only by default (process_would_reap), fail-safe keep-all on an unknown live-skill set. start_parent_lifeline() gives our python entrypoints a watchdog that group-suicides when the spawning parent dies: inside a multiprocessing child it waits on the spawner's parent sentinel (EOF on supervisor death under fork, spawn and forkserver alike — under forkserver the ppid is the forkserver, which outlives a dead supervisor while any worker holds its alive pipe, so a ppid watch would never fire), a plain subprocess falls back to its ppid. _active_subprocesses, existing port sweeps, and generation Job Objects complement durable custody. Every worker tree-kill (supervisor/worker_pool_lifecycle.kill_worker_tree: pool shutdown/restart/update, unready replacement, cancel and timeout) and custody reaping of a stale session ancestor preserve retained daemon subtrees. Windows uses native creation FILETIME plus canonically quoted psutil argv for new fingerprints and one PID/PPID snapshot for selective tree termination; it never applies taskkill /T to an ancestor of a retained branch. Legacy empty-birth Windows rows keep their earlier liveness-only retention but cannot authorize forced signals (§9).

Launcher lifecycle: the lifecycle thread removes stale port state, starts the server, follows the actual port file, and waits for health. Exit 42 requests a managed restart (refresh bundle metadata/remotes, sync dependencies); a failed dependency install retries once, then startup proceeds under the crash fuse — an offline install with already-satisfied dependencies may still be healthy. Five ordinary crashes within 120 seconds stop automatic restart; a panic exit performs full cleanup and terminates the outer process rather than the retry loop.

Presentation: the Linux browser fallback checks DISPLAY/WAYLAND_DISPLAY before touching pywebview; GTK needs a live Gdk.Display.get_default(), while Qt is trusted on env alone — probing it constructs a QGuiApplication that can itself abort, so the probe would cause the crash it exists to avoid. On probe failure the same launcher supervises the same server, prints the authoritative URL, and opens the system browser best-effort — the browser is the owner's application, deliberately outside process custody and teardown. A repeated launch that loses the PID lock soft-polls the port file (~10 s) and opens the last-read loopback URL best-effort. Browser-mode SIGINT/SIGTERM handlers only set the shutdown event; sys.exit paths leave PID-lock release to the registered atexit owner, because a second release could unlink a newer launcher's lock.

Extension children, delegated runtimes, services, the local model, and companions all sit beneath these roles: every long-lived process enters the custody ledger or a process group. Disclosed residual: shutdown admission is not atomic with publishing a spawned child (a signal can land between Popen and the record) — tracked, not a reason for a second launcher.

Platform substrate

platform_layer.py owns OS observations and lock/process primitives; their callers own policy. kernel_file_locks_enforced probes once per real directory under a module lock, using a scratch file rather than a failed live acquisition. Unsupported-lock errors and ENOLCK select a recorded name-only tier; an unprobeable directory stays enforced for that call and is retried later. Contention retries, other kernel errors fail closed, and callers may refuse specific name-tier causes. The name tier uses O_EXCL plus identity recheck/unlink without kernel exclusion; monetary compaction refuses it while appends retain their existing policy.

A won file lock must still have a readable descriptor/path inode identity; a creator evicted before acquiring its hold recontends, and unreadable identity is not proof. Owner-aware stale recovery never evicts a live writer on age alone. POSIX judges and unlinks the stale inode under its flock; Windows unlocks/closes the probe before identity-checked deletion and relies additionally on open-handle deletion refusal (CPython omits FILE_SHARE_DELETE), not POSIX's single-reclaimer guarantee. Refresh returns ownership, not a courtesy heartbeat: losing it requires abandoning protected work. Release unlinks only the held identity, before close on POSIX and after unlock/close on Windows. Windows contenders can transiently hold the path open, so unlink retries within its bound; swallowing that refusal would strand a stamp bearing a live owner's PID. LockFileEx covers one fixed byte beyond the short owner stamp, keeping the stamp readable despite mandatory locking and supporting empty files. Its OVERLAPPED is rebuilt, not cached per recyclable fd. Violation means contention; invalid-function/not-supported map to unsupported-lock errnos; other Win32 errors preserve their classification.

Birth identity is separate from PID presence. Linux prefers ticks plus boot id, then the ps wall-clock token, then separator-qualified bare ticks as a disclosed cross-boot collision limit; recovered observation capability can change the token form. The custody ledger keeps the legacy start_time spelling and the optional boot-qualified sibling so an older reader does not silently discard every owned row after rollback. Windows uses exact creation FILETIME; full command identity remains independent. Presence uses non-signalling process handles on Windows and retains access-denied/unknown presence rather than authorizing cleanup. Win32 calls declare full-width HANDLE arguments/results and snapshot last-error at the call; omitted declarations truncate 64-bit handles. Job termination/close checks false BOOL returns as failures, not just exceptions, so survivors remain disclosed.

Bundled resources use the §1 CLI/headless lookup order rather than assuming the managed server runs inside the frozen app. Node health/selection stays in node_runtime; the platform's lazy PEP 562 re-exports avoid its eager import cycle while preserving existing caller names. Platform TCP keepalive is specified with the shared transport in §6; kernel dead-peer detection does not shorten a cognitive operation's deadline.

Data layout (~/Ouroboros/)

~/Ouroboros/ is the default application root; APP_ROOT, DATA_DIR, and SETTINGS_PATH are independently env-overridable (ouroboros/config.py, §7).

~/Ouroboros/
├── repo/                          ← the self-modifying git repository; launcher-managed git clone keeps server.py in sync (never copied per launch — §2)
│   ├── ouroboros/                 ← core package (module map above)
│   ├── supervisor/                ← supervisor package
│   ├── web/                       ← Web UI; ES-module pages under web/modules/
│   ├── docs/                      ← ARCHITECTURE.md (this map), DEVELOPMENT.md (engineering handbook), CHECKLISTS.md (review checklists SSOT), CHECKLISTS_ARCHIVE.md, CREATING_SKILLS.md, DESIGN.md, DEPLOYMENT.md
│   └── prompts/                   ← SYSTEM.md, SAFETY.md, CONSCIOUSNESS.md
├── data/
│   ├── settings.json              ← user settings (API keys, models, budget)
│   ├── task_results/              ← durable task results (task_results/<id>.json, every write stamped `_schema_version: 1`; an unstamped, future, malformed or retired-key row is QUARANTINED with log-only visibility and keeps its id occupied rather than being re-minted); artifacts/<task_id>/ holds .artifact_manifest.json (private metadata) + artifact files; .scratch_manifest.json declares ephemeral scratch {abs_path: sha256} excluded from patch capture only while content matches
│   ├── artifact_versions/<task_id>/ ← artifact recovery history, last 5 versions per name
│   ├── task_drives/<task_id>/     ← task-scoped scratch, including live per-call manifests; startup prunes terminal tasks after the headless retention window
│   ├── task_trees/<root>/blackboard.jsonl ← append-only swarm blackboard + beacons; tree-scoped and ephemeral (task_tree_ledger.py), pruned on root terminal
│   ├── state/
│   │   ├── state.json             ← runtime state + compatibility cost projection; never the monetary authority
│   │   ├── queue_snapshot.json    ← durable PENDING/RUNNING recovery projection + worker counts + explicit worker_pool_disabled_reason
│   │   ├── usage_attempts.jsonl   ← append-only monetary authority: per-attempt id + state transitions; a settled attempt with cost=None and a numeric reservation bound counts at the bound
│   │   ├── skill_review_root_tasks.jsonl ← append-only compact index produced by `skill_review_runner._append_terminal_history` through `skill_review_history.append_history_once` and consumed with a bounded tail by `skill_readiness._skill_names_from_review_history`; derived projection of per-skill `review_history.jsonl`, retained append-only, with a 20 MB warning at `context_budget.SKILL_REVIEW_ROOT_TASKS_WARN_BYTES`
│   │   ├── usage_attempts.quarantine.jsonl ← loud quarantine of a proven-corrupt final row; the validated prefix stays readable
│   │   ├── usage_import_watermark.json ← resumable idempotent legacy-import watermark
│   │   ├── request_wire_compatibility.json ← cross-process locked, schema-versioned 14-day exact-route wire evidence (request_wire_contract.py)
│   │   ├── capability_evidence.json ← sourced model-capability evidence (capability_evidence.py)
│   │   ├── process_ledger.jsonl   ← durable process-custody ledger (process_custody.py; Runtime topology)
│   │   ├── server_port            ← active HTTP port for launcher/browser handoff
│   │   ├── server_port.bindings.json ← informational endpoint snapshot owned by `server_process.py`: the main, Host Service and local-model owners publish their actually bound host/port with the owning pid and process fingerprint while they hold it (`record_service_binding`/`clear_service_binding`; compare-and-remove, so a late close cannot erase a replacement); a browser identity fact, never a grant or a custody ledger
│   │   ├── server_process.json    ← launcher-owned server identity record for relaunch cleanup
│   │   ├── advisory_review.json   ← durable advisory/review ledger (runs, attempts, obligations, commit-readiness debts)
│   │   ├── deep_self_review_context.json ← last Atlas manifest + model metadata
│   │   ├── code_intel/<repo_key>/inventory.json ← code-inventory facts; no raw source cache
│   │   ├── evolution_metrics_cache.json ← per-tag metrics cache regenerated by /api/evolution-data
│   │   ├── evolution_campaign.json ← campaign objective/progress/history/budget
│   │   ├── evolution_checkpoints.jsonl ← append-only per-cycle checkpoints
│   │   ├── post_task_evolution_request.json ← worker-written one-shot promotion signal; consumed + deleted by the supervisor idle tick; dropped while evolution_owner_stopped
│   │   ├── post_task_evolution_counter.json ← per-drive every_n counter
│   │   ├── scheduled_tasks.json   ← cron (5-field + tz) and one-shot {type:"once", run_at} schedules; consumed one-shot receipts age out past the unified GC retention
│   │   ├── claudexor_rotation_provisioning.json ← receipt of the last rotation-reconcile settings POST
│   │   ├── update_letter.json     ← the last update letter (key = base/target/channel/ref, state, text, `last_good`); kept after apply and projected as pending/applied/superseded/other against the live HEAD (update_letter.py)
│   │   ├── projects.json          ← Project registry: immutable id/chat identity, working folder, lifecycle/routing fence, revision; tombstones are durable and never age-pruned
│   │   ├── project_task_bindings.json ← schema v1 root↔Project bindings with REQUIRED typed origin; one-way enrichment; tombstoning never removes a binding
│   │   ├── ui_preferences.json    ← owner-local layout preferences + monotonic project_seen_revision ACKs
│   │   ├── cancel_intents.json    ← compact locked projection of ACTIVE cancel intents; the forensic trail is typed cancel_intent rows in logs/supervisor.jsonl, never read back (cancel_intents.py)
│   │   ├── terminal_deliveries.json ← bounded delivered-dedupe + PENDING terminal-answer outbox (terminal_delivery.py)
│   │   ├── extension_companions.json ← runtime snapshot of live companion processes
│   │   ├── extension_reconcile/   ← worker-written markers consumed by the server lifespan pickup task
│   │   ├── review_continuations/  ← durable blocked-review continuations (+ corrupt/ quarantine; archived/ holds settled un-resumed rows ≥7 days, never deleted)
│   │   ├── workspace_executor_processes/ ← durable local/docker executor cleanup records
│   │   ├── consciousness_observations.jsonl ← append-only inbox; rows retained until a settled successful cycle appends an ACK; malformed rows stay visible as source gaps
│   │   ├── headless_tasks/<task_id>/data ← forked/empty child execution drives whose live per-call manifests are promoted at terminal; until then their refs are not resolvable by the canonical reader (#805; CLI / Headless Boundary above)
│   │   ├── pycache/               ← embedded-interpreter bytecode (packaged builds; CLI / Headless Boundary above)
│   │   ├── python-userbase/       ← embedded-interpreter user installs (packaged builds)
│   │   ├── betterleaks/           ← versioned scanner runtime + archive cache, created only by the explicit source-checkout installer
│   │   ├── cx/                    ← managed Claudexor store: immutable <version>-<sha12>/ trees each with managed-runtime.json + node/, cache/ of verified archives, install.lock
│   │   └── skills/<name>/         ← per-skill state plane: review.json (content_hash, findings, reviewer_models, raw actor records, advisory_result — findings stay authoritative), owner_attestation.json (owner-issued marker; removal invalidates, content edit stales via content_hash, the agent can never forge it), review_history.jsonl (append-only terminal history; raw reviewer text never exposed to chat), accepted_rebuttals.json (injected into later review prompts), deps.json (isolated-dependency install fingerprint), auto_repair.json (marketplace auto-repair dedup by payload hash), ouroboroshub.json (publication receipt; malformed reads as published=null + typed diagnostic), health.json (server-authoritative health plus worker qualifier; flags live→broken regressions across restarts), auth_token.json (content-hash-bound Host Service token), enabled.json ({"enabled": bool, "updated_at": iso_ts, "actor": host_actor}), extension_calls/ (transient per-call child-process payloads), __extension_imports/<pid>-<uuid>/skill/ (staged import trees — §13)
│   ├── claudexor/                 ← Ouroboros-owned Claudexor home (CLAUDEXOR_CONFIG_DIR): daemon descriptor/token, credential profiles, runs, ouroboros-owned.json, daemon.log; never the operator's ~/.claudexor
│   ├── memory/
│   │   ├── identity.md            ← durable identity
│   │   ├── scratchpad.md          ← auto-generated from scratchpad_blocks.json (rendered newest-first; FIFO eviction of the oldest blocks until BOTH the 10-block count cap and the SCRATCHPAD_MAX_CONTENT_CHARS content cap hold)
│   │   ├── dialogue_blocks.json   ← consolidated dialogue memory blocks (dialogue_summary.md remains a read-only legacy fallback when present)
│   │   ├── dialogue_meta.json     ← consolidation cursor/metadata for the dialogue blocks
│   │   ├── WORLD.md               ← host profile generated on first run
│   │   ├── knowledge/             ← topic files + auto-maintained index; patterns.md (Pattern Register), improvement-backlog.md (backlog SSOT), *_journal.jsonl + *history.jsonl provenance
│   │   ├── deep_review.md         ← written by the deep-self-review task
│   │   ├── registry.md            ← memory awareness map
│   │   └── owner_mailbox/         ← per-task user message files
│   ├── projects/<id>/knowledge/   ← per-project facts + provenance sidecars; logs/task_reflections.jsonl holds full reflections with a bounded pointer row in the canonical log
│   ├── observability/             ← canonical private forensic ledger: blobs/<sha256>.json.gz compressed CAS payloads (0600) + terminal-promoted calls/<task_id>/<call_id>.json manifests
│   ├── services/<task_id>/<service>.log ← service runner logs; public tool output is bounded redacted tails + private blob refs
│   ├── logs/
│   │   ├── chat.jsonl             ← canonical chat: one logical message stored once, projected into Main/Project lenses
│   │   ├── chat_annotations.jsonl ← compact routing status by client_message_id; presentation-first, with the token-bound `needs_manual_target` decision card as the one routing-authority exception
│   │   ├── progress.jsonl         ← runtime ledger: progress/thinking stream
│   │   ├── events.jsonl           ← runtime ledger: lifecycle, llm_round/llm_usage, errors
│   │   ├── tools.jsonl            ← runtime ledger: tool calls
│   │   ├── supervisor.jsonl       ← runtime ledger: workers/supervisor
│   │   ├── task_reflections.jsonl ← canonical reflection log
│   │   └── containment_faults.jsonl ← append-only compact projection of containment incidents (delegate_custody.py)
│   ├── archive/                   ← rotated logs, rescue snapshots, archived managed repos
│   └── uploads/                   ← chat file attachments (paperclip)
├── Deliverables/                  ← bare user_files filenames land here (OUROBOROS_DELIVERABLES_ROOT; sibling of projects/, outside repo/ and data/, never GC-pruned)
└── ouroboros.pid                  ← launcher PID lock; platform lock auto-released on crash

Every entry of this tree is probed against the tree and the runtime sources by the generated docs/v7next/DATA_LAYOUT_INVENTORY.md, so a durable file renamed in code while its row here survives is red, not silent.


2. Startup / Onboarding Flow

Packaged startup is an ordered ownership transaction. The launcher prepares the platform UI runtime (or runs the Linux browser-mode probe), acquires the single-instance lock, verifies Git, and validates and bootstraps the embedded managed-repo seed — server preconditions, so they precede it. It then removes only identity-proven stale server state and ports, starts the lifecycle thread, and waits on /api/health at the authoritative port from data/state/server_port; only then is first-run onboarding presented against that live server, before the pywebview shell or browser presentation opens. The server starts the gateway first and the supervisor/worker pool only when provider configuration is structurally sufficient.

Onboarding runs after the gateway because connecting an agent subscription is a live /api/* conversation, not a form field — and a gateway without a supervisor is exactly what the readiness predicate produces, so no second server, mode, or onboarding state machine exists. ouroboros/launcher_onboarding.py owns the presentation (readiness decision, setup window, window-lifecycle bridge); when completion reports a boot-pinned value changed, the launcher recycles the managed server rather than counting the exit as a crash. Neither launcher nor server boot normalization may CREATE settings.json (the install-time latches below are gated on its absence).

has_startup_ready_provider() is a structural gate, not a network, credential, entitlement, model, or local-process probe: any non-empty recognized remote configuration (including DeepSeek) or active task-capable local routing flag passes (key list: server_runtime.has_startup_ready_provider; USE_LOCAL_HEAVY is legacy migration input only, LOCAL_MODEL_SOURCE alone insufficient). When the gate is false the server marks startup complete without workers so the web UI serves the blocking onboarding overlay; a later successful settings save hot-starts the supervisor.

Every host renders one served page: GET /onboarding returns onboarding_template.html with the settings_setup_contract bootstrap injected, linking wizard CSS and web/modules/onboarding_wizard.js as static assets so steps import the same modules as the rest of the UI — an inlined srcdoc string cannot. The desktop setup window opens that URL, the blocking overlay frames it, a browser owner opens it directly; GET /api/onboarding is the readiness probe (204 once the gate passes, otherwise the page). Steps: Accounts, Models, Review, Budget, Summary; context mode remains outside this wizard. Accounts includes optional agent connections and mounts the shared login cards in full mode because compact omits the paste-code entry a Claude login needs when its localhost callback cannot complete; its account facts come from the shared Claudexor status store and become the completion payload's subscriptionsConnected declaration — a request to look at the daemon, never an authority.

Completion is one HTTP conversation on every host — POST /api/onboarding/complete (gateway/onboarding.py; its docstring carries the numbered step order) — so providers and runtime mode cannot be left half-saved. Its single write, re-proved under the settings lock, persists settings, next-boot runtime mode, the fresh-install OUROBOROS_SAFETY_MODE=light default no other endpoint may author, the one-shot preset marker, and the durable completion fact; only then does the supervisor start. Install time means three proofs together (gateway/onboarding.py::preset_eligible): onboarding never completed here, no preset generation applied, and no settings.json yet — three, because "no startup-ready provider" alone is a state an old install reaches whenever its key stops working. GET /api/onboarding normalizes what it displays but never persists: a read that created settings.json would silently disqualify the install-time latches. A completion failing AFTER bytes reach disk reports that it saved plus the failing stage; a completion whose body outlives the shared writer bound (settings.py::_run_settings_writer, the seam every settings writer runs through) answers 503 settings_save_timeout with saved: null and the wizard offers "Check status" (a re-read of the readiness probe) instead of a blind re-save; a 2xx is not a completion — only the exact success envelope is, because the saved runtime mode and restart receipt live in that body. On a boot-pinned change the framed overlay shows its restart card and a plain browser tab shows the saved-but-restart-required screen (the launcher recycle: above). The overlay sandbox grants popup permission because the agent sign-in link is the step's primary action.

ouroboros/subscription_install_presets.py is a pure compiler for sibling task actors, reviewer slots and model assignments, not a reviewer/subagent matrix; its docstring carries the emission rules, the ratified Claude/Codex/Cursor reviewer subset, and the Agy-neutrality contract. Normalized API/local settings alone suffice for API-only and local-only installs; a declared subscription makes gateway/onboarding.py read a live account snapshot and the declared raw-model source catalog before the settings transaction — the authority for supported harnesses, exact model ids (Agy's automatic row: gemini-3.8-flash-high), and enabled accounts. Durable seat facts (kind/enabled/verified), not the hour's quota reading, decide the once-only preset: a spent Claude window at save time must not yield a Codex-only preset permanently. Task-actor rows stay unpinned and grow linearly, never as a powerset; on the fresh-install path reviewer slots are subagent_id references into the roster the preset ships (a seat matching no task actor mints a review-<harness> roster row), while an owner-configured roster stays validate-only. Missing required exact discovery is a typed pre-write refusal; no partial preset persists.

POST /api/onboarding/subagents/preview runs the same compiler against the open draft without persisting; the editable result becomes OUROBOROS_SUBAGENTS, and an owner-edited draft is validated and preserved, not regenerated. Completion persists actors, reviewer disposition, preset receipt, and completion facts through the one-write document-lock/fingerprint boundary; network discovery happens before the lock. A daemon failure keeps the wizard open with an explicit finish-without-agent-defaults path; a read or failed refresh never rewrites saved intent.

Validation is structural: at least one exposed remote configuration, selected managed model source, or local model source; local-only setup routes at least one active lane locally; Main required, while Light, Vision, Consciousness, and Fallback keep inheritance/empty semantics and Heavy is readable only for bounded migration into an explicit API actor; enforcement and runtime mode are closed enums, budgets finite and positive, the MiniMax region closed, a Hugging Face local source needs a filename. Credential length is checked only on fields changed in the payload: rejecting an unchanged short legacy value would discard the whole form, including its own repair.

Provider readiness and provider defaulting are separate. With no OpenRouter, legacy OpenAI base, or OpenAI-compatible endpoint, exactly one registered direct provider receives provider-prefixed defaults and migration of untouched shipped/legacy slot values; OpenAI, Anthropic, Cloud.ru, GigaChat, MiniMax, and DeepSeek each use their own registered defaults. Multiple direct providers stay owner-editable, and OpenRouter keeps router-style routing. An arbitrary OpenAI-compatible endpoint gets no guessed model ids — compatible servers have no universal safe name; the owner selects explicit openai-compatible::... routes. A local-source install with no remote provider clears only untouched shipped remote Light/Fallback values that would be unreachable; owner-authored values and explicitly local slots are preserved. This is migration of defaults, not a model allowlist, and never proof the local server is running.

scripts/build_repo_bundle.py creates the packaged seed only from a clean named checkout, writes a git bundle of that commit/tags, and records schema, version, source SHA, release tag, bundle hash, and managed branch/remote metadata (the release-tag check itself: §8). The launcher validates the manifest fields and the bundle SHA-256 but no per-file member set: clone-time Git verification proves the manifest source object exists and checked-out HEAD equals it.

ensure_managed_repo() owns packaged checkout bootstrap: a first install clones the bundle into a temporary checkout, verifies and configures the pinned source SHA and managed branches/remotes, then moves it into repo/ (an existing legacy non-git directory is archived first). Once a managed git checkout exists, a changed application manifest does not archive or replace its working tree: bootstrap atomically refreshes managed metadata and the official managed remote in place, preserving the local branch tip and owner edits. Ordinary restart performs no network fetch — network movement to an approved official SHA belongs to the pinned managed-update path in supervisor/git_ops.checkout_and_reset, not to bootstrap — and origin remains optional personal persistence, not the official update authority.

Bootstrap also creates the initial world profile when absent and seeds launcher-owned native skills without resurrecting an intentionally deleted seed. launcher_bootstrap._per_skill_version_resync replaces existing marker-owned payloads when their manifest versions differ in either direction. Equal versions retain installed bytes and lifecycle state; the shared content hash reports payload drift or an unavailable comparison without blocking startup. Bundled payload changes increment the skill manifest version independently of the application version. Dependency installation runs only when checkout/bootstrap metadata changed and its boolean result reaches the launcher; after exit code 42 the launcher refreshes bundle metadata and synchronizes dependencies before starting the edited body. A failed install gets one visible retry after five seconds; a second failure stays in logs while startup continues under the five-crashes-in-120-seconds fuse — silently losing a pip failure makes a later ImportError inexplicable, while refusing every offline restart would break a checkout whose requirements are already present.

Managed supervisor bootstrap is the sole owner of destructive dirty-tree recovery. Before any reset/clean, supervisor.git_ops writes a merge-aware rescue directory: porcelain status, changes.diff captured as raw bytes with a hardened argv/environment (supervisor/git_ops.py) because it is the only carrier of a resolution stash cannot capture, a stash-created rescue ref when possible, copied untracked files with completeness metadata, unpushed-commit evidence, rescue_meta.json. An incomplete snapshot blocks rescue_and_reset — not permission to discard what could not be captured. Normal managed bootstrap then cleans back to the local branch's own HEAD, not to managed/<branch>.

Managed update is the second user of that machinery, with the opposite failure policy. Every destructive rollback shares one choke point in rollback_managed_update, and both it and boot-resume re-materialization take a FRESH rescue before the first destructive command, because the pre-update snapshot predates the merge and holds none of the resolver's work. This hook is fail-open: a rescue that cannot be taken never blocks the rollback (logged and disclosed), one durable supervisor.jsonl line is written at capture time, before destruction, so the record survives a crash mid-teardown, and a git status that cannot answer counts as dirty. The rescue is deliberately NOT linked to an active evolution transaction — that would flip the campaign's cycle to abandoned for an unrelated reason. The update transaction persists a pointer to what was rescued: a replayed rollback does not duplicate a snapshot, a retry re-rescues the tree it actually finds, and the resolver's objective names the latest rescue directory — re-materialization re-creates a dirty tree WITHOUT replaying the rescued edits and must never be read as their return. Rescue Git processes are bounded by OUROBOROS_RESCUE_GIT_TIMEOUT_SEC with process-tree termination on timeout.

An active Evolution transaction or managed-update merge uses rescue_and_block: evidence links to the transaction and the tree is left intact, pausing Evolution rather than erasing partially resolved work; with no such owner, startup uses rescue_and_reset. Source/local-development startup skips the managed checkout/reset path (dependency sync plus import test only). Worker startup checks are diagnostic and warning-only: launcher-management environment variables propagate into worker, review, and test subprocesses, so per-constructor auto-rescue would let an incidental child steal or clean another actor's in-progress edits.

server.py establishes OUROBOROS_AGENT_PYTHON from its actual interpreter immediately after binding the repo import root, before workers or review subprocesses start. Hermetic commit/review preflight uses that handle (then sys.executable, then python3) so tests run in the environment containing Ouroboros dependencies; plugin verification is part of that preflight, not a launcher claim that every interpreter was live-probed.

User process tools have a separate surface-aware resolver (ouroboros/process_interpreters.py, which carries the priority ladder): for exact unversioned python/python3 on the four public process launch surfaces, registry pre-dispatch resolves once before deterministic guards so guard and handler see byte-identical argv. Absolute or versioned interpreters, shell bodies, and non-Python commands stay literal at pre-dispatch (the post-gates Node ladder: §1 process_interpreters.py row); resolution emits secret-free provenance, never silently installs dependencies, and fails closed only when a system-owned interpreter cannot be proven.

3. Web UI Pages & Buttons

The Web UI is a build-free vanilla-JavaScript SPA (web/index.html, shared CSS, web/modules/*). No TypeScript or bundler step, so the running interface stays inspectable and editable by Ouroboros without regenerating opaque artifacts. web/app.js owns top-level page, Project-panel, and mobile-navigation state; feature modules own their presentation. Every long-lived UI acquisition carries a disposer bound to the lifecycle that created it (docs/DEVELOPMENT.md "UI resources carry a disposer").

The same SPA serves the desktop shell, ordinary browsers, Docker/web deployments, the Telegram mini app, phones and the Linux browser fallback: the fallback changes presentation only — no second API, onboarding contract, runtime identity, or state owner — and the browser stays the owner's application outside process custody (§1).

The desktop shell exposes a small MainApi JS bridge (window.pywebview.api, launcher.py): three native confirmation methods (runtime mode, reviewed-skill auto-grant, skill key grant), download_file_to_downloads, open_file_with_default_app, open_external_url (absolute http(s)/mailto), and save_bytes_to_downloads for live base64 payloads. The loopback-file methods share one guard: loopback host, exact server port, and a path allowlist of /api/files/download, /api/extensions/..., and /api/tasks/.... Because the embedded WebView has no new-window or download delegate, ui_helpers.js installs a shell-only link interceptor when the bridge is present, in BOTH top-level documents (SPA and framed onboarding wizard), routing each URL class to the matching bridge method. Bridge methods are feature-detected per call because the packaged launcher updates only on reinstall while the served frontend updates with the managed repo; missing methods degrade to copy-link-plus-toast or the file-helper fallback chain.

The authenticated Telegram proxy's existing X-Ouroboros-Telegram-MiniApp: 1 presentation marker makes server_web.make_index_page include the asynchronous Telegram SDK and host hint in the main document. Ordinary documents remain unchanged. An unavailable SDK reports a click failure and never replays it after loading; the marker grants no authentication authority.

One shared WebSocket serves the whole application; Projects open no independent sockets, and REST remains the recovery and durable-read path (protocol and connection order: the WebSocket protocol subsection of §4).

Navigation and shared UI contracts

Primary navigation exposes Chat (Main), a collapsible Projects group, Files, Skills, Widgets, Dashboard, and Settings; About is a Settings sub-tab. syncNavigationState() is the single presentation state machine for active page, active Project, Projects expansion, mobile drawer, and panel backdrop — independent toggles must not leave multiple rows active or a hidden surface looking selected. Sidebar and Project panel widths are owner-local UI preferences, not runtime settings. The application has exactly one client-side route: a #<page> fragment for the injected page ids, honoured once on load and validated against the existing #page-<name> section, so an unknown fragment is ignored rather than painting a blank surface. It is never written back on navigation — the packaged desktop shell and the Linux browser fallback have no address bar to read it from, and the Telegram mini app always loads at / — so a browser gains a shareable /#widgets link while no other surface changes. Projects are panels, not routes, and the sidebar is the document's one <nav> landmark.

Each active or deleting Project has a sidebar row with pointer- and keyboard-operable open/rename/delete; the backend owns the 80-character name limit and lifecycle truth. Unread Projects sort ahead of read, then by durable activity; a deleting Project becomes non-openable and stays visibly transitional until the server publishes authoritative registry state. On narrow screens navigation is an explicit drawer and the Project chat a full-width overlay; no gesture-only navigation layer competes with message scroll, text selection, or the software keyboard.

Shared frontend primitives keep pages from acquiring competing contracts — frontend work must not reimplement supervisor, review, marketplace, extension, and provider semantics per page:

  • web/ui.css ← shared values and explicit field/button/status/popup classes; index.html and onboarding_template.html load it first, before their page sheets. style.css keeps shell composition and tab chrome; page-specific layout does not duplicate the palette. DESIGN §8 names the migrated families/regions, not an all-page migration.
  • page_header.js / page_icons.js ← page headers, tab-strip markup and bindTabStrip; navigation/header icons. The binder synchronizes selected class, ARIA and roving keyboard focus, revealing selection only within the strip; each page keeps panel/loading ownership, and silent select() supports programmatic navigation without a second callback.
  • ui_interactions.js ← bindDialogFocus, bindMenu and geometry-only bindPopoverPosition. Callers mount/portal and remove their own markup; measured --ui-popup-* values drive the common fixed, viewport-bounded .ui-popup. Menus own action focus/dismissal; editable listboxes reuse geometry alone. Every binding returns teardown, with no global overlay registry.
  • scroll_fade.js ← bindScrollFade projects actual hidden content onto the scroll body's edge attributes, keeping a fully visible first/last edge readable; the binding owns its observers, scroll listener and pending frame.
  • api_client.js / api_types.js ← browser API calls with typed error propagation; mirrors of the browser-facing contract shapes.
  • ui_primitives.js ← self-contained renderSafeField, collectSafeFieldValues, escapeHtmlAttr, normalizeTone and setInlineStatus; no shell/API or document-global initialization. Existing ui_helpers.js / utils.js imports re-export the same functions, so author pages do not need a second field renderer.
  • ui_helpers.js ← host-bridge, shared row/badge helpers and keyboard menu-lock behavior; the design-system action button for host-stamped system chat rows (createSystemMessageAction); installs its Alt guard on both top-level documents.
  • skill_card_renderer.js ← installed-skill cards.
  • hub_sync.js ← the one catalog×listing hub-card verdict (install/installed/update/adopt/wait_pr/none plus badges) for the OuroborosHub tab and My-skills badges; joins the catalog with /api/extensions by canonical name.
  • client_surface.js ← the send-time sending-surface snapshot (raw observables, no device taxonomy) that chat.js spreads into each chat frame.
  • chat_markdown.js ← the chat rich-markdown renderer: marked+DOMPurify, a chat-local URL policy (external http(s)/mailto plus the canonical /api/files/download form), KaTeX (single-$ unsupported), lazy mermaid, bounded ```chart rendering, and an enhance/destroy lifecycle whose disposers chat.js invokes on bubble removal. Vendored renderer versions: web/vendor/VENDOR-MANIFEST.md.
  • log_events.js ← event classification, shared technical outcome reducers, and one factual task-presentation projection for Chat and Logs: task truth maps only to Working / Done / Done with warnings / Failed / Cancelled; it owns no actions, incidents, or notifications, and compact headlines never expose raw reason codes.
  • toast.js, masonry.js, widget_frame.js, widget_job.js, CSS tokens ← notifications, layout, framed-widget bootstrap/lifecycle, bounded widget request/job policy.
  • task_control_menu.js ← the shared task stop/hurry dropdown, verbatim on Chat live cards and the Activity tab: "Wrap up" / "Hurry up" / "Stop now" (frozen owner wording); a host-attested budget-paused member swaps the working pair for "Resume" (POST /api/tasks/{id}/resume, refusals surfaced verbatim). Eligibility gates differ per surface; actions, endpoint bindings (stop_policy mapping, stable per-task request_id retry), locking, and typed refusals do not. Choosing an action executes it immediately; dismissing continues the run; a pending cancel leaves only the hard escalation; "Hurry up" acknowledges via local toast, never a chat message.

confirm_dialog.js::openConfirmDialog is the one browser-dialog authority: confirm mode resolves a strict boolean, input mode {confirmed, value} (empty on cancellation), alert mode one acknowledgement button; Cancel, Close, backdrop, Escape, and supersession all resolve as non-confirmation. Native window.prompt/confirm/alert are forbidden in web/modules: inconsistent across shells, event-loop blocking, and window.prompt silently returns null in the macOS PyWebView shell. Critical controls act only on the exact confirmed result — Panic's confirm-and-send is one testable operation.

Confirm and New Project dialogs use bindDialogFocus without sharing result semantics. Callers bind after mounting and dispose before removal; restoration does not take focus from a different active surface. Menus close before their actions open a dialog. model_chooser.js supplies the shared editable suggestions for model roles and actor/reviewer route editors, using popup geometry alone so menu focus behavior cannot interfere with native text editing. The input keeps its draft/caret/composition through discovery updates; Arrow keys highlight and Enter or pointer selection assigns a suggestion, while Escape/blur dismiss without assignment. Fixed source/account choices remain native selects, and unknown saved or arbitrary API model ids stay editable without an inventory allowlist.

Chat and Projects

web/modules/chat.js owns the canonical message timeline, drafts, attachment staging, runtime controls, budget projection, routing annotations, task and child cards, WebSocket subscriptions, thread routing, dedupe/insertion order, unread state, and reconnect reconciliation; its per-instance chat_media.js controller owns delivered-media presentation and resources. A required Project quiz additionally projects one keyed System pointer into Main, without a second durable message or answer form. Its task/quiz identity opens the existing Project question; task detail restores a question outside bounded history using the stored labels and optional aligned option details. Recorded answer/expiry updates the same pointer; resumed work with an unanswered quiz has neutral question wording, and missing source stays unavailable. active_chat_activities.required_question recovers current waits through the existing stat-keyed result memo, shared with finalizing. Every ordinary message has one canonical durable chat row; Project views are lenses over those rows and task bindings, so conversion creates no second message, unread event, or cost record. Routing acknowledgements live in a compact sidecar keyed by client_message_id and update the existing owner message without a synthetic assistant bubble. Successful annotations carry the confirmed destination Project address for the shared open-or-focus action; opening a pointer never toggles an already open room closed.

Messages, media bubbles, and task-card roots order by raw numeric timestamps (ties keep arrival order; the typing indicator stays last). Application-controlled height mutations use a stable-viewport seam: within 48 CSS pixels of the live edge the transcript follows the bottom, otherwise it restores the visible anchor; remote delivery while the reader is away coalesces into one instance-local activity marker cleared at the bottom. Card text is selectable and a selecting drag never toggles the card; band anatomy and the quiet Reviews N count: docs/DESIGN.md. Scroll position is remembered per chat instance.

History reconciliation is two-pass: progress and system records first rebuild timestamped task-card state, messages and cards insert chronologically, and only then may terminal state seal a card — so terminal replay cannot discard earlier progress, and cards whose final summary was missed offline recover. Live echoes and history rows deduplicate on stable message identities. That rebuild is also the memory bound: cards are minted from progress rows, task_summary rows and subagent lineage, so the cap is relative to the population the last rebuild produced — once more than 200 live cards exist beyond it the next history sync replays durable history instead of folding into the existing cards, and that rebuild sets the new floor (web/modules/chat.js).

The Main chat receives ordinary main-thread dialogue, required Project question pointers, and the two host-stamped Project lifecycle rows — project_started and the terminal project_completion_summary, each with the shared "Open Project" action; all other Project traffic stays in the Project thread. Lifecycle rows are plain dashboard text: the producer strips markdown once before durable write and live send, history normalizes older rows on read (chat.jsonl is never rewritten), and the renderer escapes any system row without markdown: true (except skill_review's dedicated renderer). Finality metadata is row-specific. Durable task_summary rows carry the SAME phase the browser paints: project_dialogue.outcome_phase mirrors log_events.js — the terminality gate plus the severity fold over normalized axes — and stamps it as outcome_phase beside the status word taken from OUTCOME_PHASE_HEADLINE, so there is no second status-word family. The pre-finalization authored row carries the phase with outcome_final=false but no host verdict clause; only terminal task_summary rows append the host verdict clause (_completion_verdict, the host twin of taskReasonDetail's acceptance branch) when one exists. The Main project_completion_summary instead puts the shared headline and any verdict directly in its plain text; it does not carry outcome_phase. The start row carries neither outcome finality nor a verdict. A verdict is the execution reason on ordinary terminal summaries, or — when the host's acceptance decision is anything but accepted — the decision status with its stored rationale, markdown-stripped and flattened like every other lifecycle-row text; the latter leads the Main completion row so a host-authored row never presents an unaccepted claim as the whole story. History keeps older status words unrewritten, and the Telegram skill reads the stamped task-summary phase instead of deriving one of its own. Project panels accept only their own registered chat_id; the complete Project chat-id set is separate from the sidebar's bounded summary, and a projects_changed frame adds a new chat id synchronously before the asynchronous state refresh, so an early frame cannot be misclassified as Main.

The Main composer exposes one-shot Swarm planning, the owner Low/Max context choice, file attachment, and Send. Swarm places a structural force_plan fact on the next ordinary message and disarms after that send — never inferred from keywords. Low/Max uses the dedicated owner endpoint; a derived Low fit never masquerades as an owner-selected posture, and selecting Max may invoke the exact-route capability acknowledgement when provider metadata cannot prove the window. Project panels omit global Restart, Panic, evolution, consciousness, review, and budget controls — those belong to the one Ouroboros process, not a Project thread.

Attachments stage via paperclip, paste, or drag-and-drop without the former count/50 MiB rejection; files stream through upload and canonical staging, while the composer and task authority show bounded previews. Staging is local; upload begins immediately before Send, and offline handling is the WebSocket queue contract (§4). Structured attachment metadata lets the gateway expose a native image to a vision-capable route and stage the complete file set through the shared artifact substrate.

Simple text messages may appear immediately as pending local bubbles and reconcile against their echoed client_message_id. Delivered documents capture immutable task-owned bytes and carry a verified file_ref plus canonical download URL (with a compatible Files URL when available), rebuilding from history without persisting base64. Files above the former 50 MiB inline boundary remain file-backed; browser downloads use HEAD followed by a native download request and desktop saves stream through the existing launcher helper; photos and videos keep base64 only in the live frame — supported media is stored under the canonical task artifact root with a content-addressed URL so replay and rebuilds restore the same bubble. Structured links rows render at most twelve independently revalidated HTTP(S) buttons, identically live and replayed. The supervisor transport owns durable media copies so all producers converge on the canonical data root; failed persistence stays an honest caption row and cannot finalize a task card.

Ordinary Main and Project conversations use complete native execution even while other work runs; Project presence and routing metadata do not select the transient ephemeral contract. Each turn owns its agent and creates no supervisor queue record. The thread-safe process-local supervisor.active_activity.DirectActivityRegistry retains its private actor handle from preparation through result delivery; Stop, Hurry, quiz and steering resolve that exact actor, while active_direct_turns and typing frames expose only its presentation fields. The update admission lock covers check-and-registration only; writer drain waits the registered executions, including post-task work. Custody maintenance and restart census include those same live IDs, and Background Consciousness stays paused while any foreground execution remains. Explicit Swarm routing retains its ephemeral contract. The producer carries _is_direct_chat into the durable result and terminal event so late/replayed ordinary Project completions stay in their room and do not emit Main task-completion summaries. Both render their progress, tool telemetry and typed conclusion on the ordinary live card; an ephemeral decision turn's ephemeral_decision marker only withholds the task claims a managed root would get — no Turn into project, and no Cancel because no host cancelable marker ever rides its frames — it never hides the work, its blank-status task_done concludes the activity like the final frame, while the already-computed outcome axes, reason and accounting decide success, warning or failure. The final chat row retains those same facts, tool/round counts and its ephemeral marker so history can replay real tool activity and one keyed terminal note without a durable task result; a ledger-final amount stays final because ephemeral turns run no post-task synthesis. Historical rows missing those facts remain legacy, without an inferred cost or recovered outcome. GET /api/state also exposes active_chat_activities: those rows united with ROOT managed queue tasks as kind="managed_task" with phase queued, budget_paused (PENDING fenced by a budget-pause fact; nothing dispatches it until an explicit resume, so it must not masquerade as queued), working, or finalizing (RUNNING with an open post-task checkpoint), so a chat instance created after the task started hydrates from the queue authority rather than transient typing frames. active_chat_activities_complete authorizes absence-based reconciliation only when every required live-identity source was read successfully; partial snapshots still carry positive activity. Optional question-detail failure does not invalidate the live-id set. The client status reducer derives the chat header only from connection and authoritative activity state; a terminal failure stays a factual task result, never a reasonless header Attention, and a local Sending... submission is retired only by authoritative evidence (typing frame, snapshot turn, durable routing receipt, replayed user row, turn conclusion, or offline-queue eviction) — never by the live user-row echo or a socket write.

Owner-message continuity is journal-backed: each locally sent owner row is kept (bounded, with its routing annotation) until a fetched history response returns the same client_message_id, and a full rebuild re-renders unconfirmed journal rows after the server response, so a stale history snapshot cannot erase a message the owner just sent. Task finalization is honest: a root's early final answer carries task_phase="finalizing" (the same fact replay derives from the open post-task checkpoint), the card holds a sticky Finalizing… phase, and task_cost_finalized is bookkeeping that never resolves a card. If an observed managed root disappears from a queue-authoritative snapshot whose request began after the observation, the state-refresh fan-out removes activity/cancel authority and starts one single-flight durable task-detail read; request generations prevent an older response from undoing a newer projection. Two liveness invariants close the stuck-Working... class: the history window is lineage-closed — subagent lineage older than the window's progress recency floor is kept only while the child still runs or finalizes, or its parent is REPRESENTED by the response (a row that proves the task's card: telemetry, its own message or summary, a folded review group or plan reference it owns — a typed photo/document/quiz delivery does NOT count), or the parent is alive, and an anchored child represents its own children in turn, so a nested swarm is kept or dropped whole; a FINAL row failing that test is emitted without lineage fields (ouroboros/gateway/history.py, one anchor set computed in the quota pass), so replay cannot mint an unfinishable parent card while a visible or live parent keeps its completed children and their executor receipts — and liveness reconciles over the CARD SET, not only the activity registry: an unvouched connected root card is finished ONLY by proven durable terminal detail, and a card without a current result stays unconfirmed until a complete fresh activity snapshot excludes it; it then uses the retained typed terminal fact or a neutral Outcome unavailable anchor, without task phase or controls; a root proven terminal (its task_done or its terminal detail) settles the descendant cards a lost child terminal left open from each child's OWN durable result — one single-flight read per child through the ordinary child-terminal path, never a cascade from the parent's outcome, never from the card-set scan. The header badge has exactly one writer, the status reducer, with the panel-boot "Online" seed as the sole exception.

Task activity collapses into a live task card per root instead of flooding the transcript; log_events.js keeps the live task card and grouped task cards on one reducer across Chat and Dashboard Logs. Non-terminal LLM/tool/checkpoint failures stay inspectable timeline facts but do not promote the card; only authoritative terminal task truth changes terminal status, and unknown Chat event names do not acquire severity from keyword substrings. Dashboard Logs files a row under Errors from the same typed projection that paints its phase pill (categorizeLogEvent reads the summarizeLogEvent phase — a failed task_done, an errored/killed/timed-out tool, an LLM call failure, a review lifecycle error), a replayed tools.jsonl row reads its is_error/signal facts exactly like the live frame, typed ok/level facts outrank a failure-shaped name, the name-substring test survives only as the disclosed non-expanding remainder for an unknown name that carries no typed fact, and a task group keeps Errors once earned. A delegated harness run's swarm_fanout row (host role="delegated_run") is labelled as a delegated run, never as a subagent. A failed child keeps a local Failed chip while its root continues. A card keeps a concise plain-text latest-activity line and an expandable timeline; server-truncated rows fetch the complete typed task record on demand into a bounded viewer. The compact card is a navigation projection, not the result authority; urgent toast/unread behavior stays confined to explicit task_incident facts, never inferred from severity. The ONE explanatory line under a terminal headline is taskReasonDetail: an owner-requested soft stop has none, a hard failure or a cancellation names its execution reason, a non-accepted host acceptance decision names the decision status plus its stored rationale, and everything else keeps the typed reason phrase (an unknown code stays raw); Logs meta additionally names review <status> and acceptance <status>.

Subagents render as distinct child cards keyed by their actual child task ids; parents keep lineage references without duplicating the child's final answer. Nested children collapse by default: the headline names the role (a short task id disambiguates otherwise identical siblings; Logs retain diagnostic model/id detail), with status carried by the chip. Agent model or Coordinator is labelled separately in metadata, beside the executor observation/receipt. A child keeps one useful activity line visible and a root allows three, without an empty activity band or a duplicated headline; full narration and Reviews retain independent disclosure. A collapsed child keeps that identity row at every card width — the narrow-container regime gives its title a 160px demand and lets the side controls wrap under it rather than dropping the title to a third band — and nested frames stay: each child box paints its own border around its whole branch and an opaque secondary background, preserving ancestry without accumulating translucent brightness at greater depth. A full history rebuild keeps those branches inside their parents: a parent represented only by its final row is minted as the nested card its own lineage names, and the replay mount leaves a card that pass 2 nested inside another where the lineage placed it. A harness pin the resolution refused shows {harness} · blocked on the child card (the route the pin named, projected onto the live frame beside the typed subagent_executor_unavailable terminal), never dispatched or no run yet. Reviews are not task lineage: a reviewer run's execution receipt belongs inside the real owning task card, and its harness or neutral API mark identifies the delivery channel without minting a child card or proving execution. Selected subagent_id/configured snapshot, requested route, effective engine route/model/account, and terminal execution evidence stay separate facts — intent must not be redrawn as proof of where the run settled. Legacy lane/executor fields stay readable on historical cards only.

Live executor presentation travels as optional executor_observation on the existing progress metadata: tools/delegate.py passes already-polled run detail and owned custody to delegate_progress.emit, which selects the latest original typed timeline actor (harnessId, attemptId, event type and snapshot lastSeq). The record binds task_id, task_attempt (empty only when unknown), run_id, attempt_id, harness_id, phase and revision; an optional model names its model_source. The current producer supplies only requested, and only when custody's requested harness matches that timeline actor; the serving model is absent from that live timeline, and a final-attempt summary never fills the gap. subagent_messages.executor_observation_meta validates/copies the same task-bound shape at Agent emission, supervisor delivery and history replay. A plain later coordinator note inherits no observation. The gateway/JSDoc mirrors and Chat's shared metadata carry list retain the field through both live and terminal-channel projection paths. log_events.executorChip labels it last update, keeps requested/observed source explicit and rejects older timestamp/revision observations at the existing sticky-chip seam. Settled execution_evidence.harness_models appears separately as observed history. harness_presentation.js::executorIdentityMarkup is the pure markup owner for the projected executor chip, labelled Coordinator/Agent model and observed-model list; chat.js supplies facts without assembling another identity label. This adds no polling, persistent current-actor record, task authority or terminal evidence; execution_evidence and actual_substrate retain their settlement-only meaning.

An Advisory task-author decision is presented once inside its task's acceptance group; an older panel never inherits a later author's hash. The same keyed renderer preserves disclosure and reading position.

The task card's Reviews section is a read-only projection over independent domain authorities (skill history, typed plan-review state, task-acceptance evidence, repository-review records), which keep their own lifecycle, verdict, enforcement, and raw evidence. A row is admitted only with stable review identity, an exact real presentation-owner task, typed domain state, and exact subject/candidate binding where required; incomplete review stays on its domain surface. Admitted today: task-bound Skill Review, plan review, task acceptance; advisory and commit review stay on their domain surfaces until a bounded exact projection exists. When review history is the only retained fact for an exact owner, Chat may render an inert owner anchor with Reviews — no task phase or liveness until canonical task activity arrives. The projection never synthesizes a task-wide verdict and changes no review routing, status, attention, or enforcement. Skill references fold into groups and attempts by the bounded Chat-history reader; plan review and task acceptance hydrate through the task-detail seam (a canonical plan-state write appends one empty typed review_reference for a single-flight refresh; task-result plan_review_state stays authority). A review_reference row is addressed by the task's bound project chat, else by the chat its caller named, else by the hidden partition: Main is never its default, so a run with no owner-visible room keeps its review rows out of the owner's conversation (rows written before a bind are re-classified on read, never rewritten). Truncation uses the quota cause so Load older expands both overlays; duplicate Skill lifecycle acknowledgements are typed lifecycle_pointer rows that may enrich a present exact owner but never mint a card. Opening and closing Reviews belongs only to the user (docs/DESIGN.md). chat_id=0 remains the hidden Skill Review partition, never a Main surface. No review inbox, generic review endpoint, review ledger, or second state machine.

Card cost is sticky task-scope evidence. Only frames carrying task accounting status, finality, subtree, reservation, or unknown-cost fields may update it; an unrelated per-call cost_usd delta is never relabelled as the task total. Compact cards render one amount — the accounted upper bound worded up to while the ledger is open, plain once final — preferring the complete subtree projection; a running root's heartbeat may carry a non-final aggregate from the physical-attempt ledger, so the amount advances without a second timer, endpoint, or client-side sum, and a history window replays the same non-final projection (live_root_cost_projection, the one cost owner) on a running or finalizing root's latest in-window progress row so a reload between heartbeats shows the same ceiling — a subtree with no attributable rows stays absent, never zero. Precedence: unavailable, then pending, then final, newer evidence winning within a class; costless frames cannot erase a known value. Dashboard and task-detail accounting derive from the physical-attempt ledger, not the card; compact review rows copy or sum no money, and Skill wave attempt/slot money appears only inside the lazy exact-job detail, joined from the same ledger via physical_attempt_v1.

A Cancel action appears only on unfinished, unconverted root cards carrying the host-attested cancelable=true; both pooled tasks and addressable native direct turns use the existing custody owner, while card shape alone grants no control. The stop control is the shared dropdown above, sent with cascade:true (root plus live descendant subtree): the hard path answers only after teardown or a typed refusal; the graceful "Wrap up" path records a durable finalize intent and answers immediately with a typed 202 pending acknowledgement. Natural completion wins a race, a missing live task reconciles from the durable record, and a process that cannot be proven dead remains a visible refusal rather than being painted Cancelled.

A Main root card may be turned into a Project: conversion creates or reuses the Project, names it owner-facing, binds the task and its canonical origin message, moves live work onto the Project lane, and replaces the Main action with a calm Project pointer; tasks already bound to a Project get no second conversion button. The pointer is the one shared project chip (ui_helpers.js::renderProjectChip): cards inside a Project panel and nested subagent cards never receive it, and clicking opens the panel or does nothing when already open — a pointer never closes what it points at. Naming reuses an already coined model title or falls back through the server naming path; the UI invents no second name authority. Project-SCOPED is not project-BOUND: a headless/CLI run carries a project_id for lease and memory without a durable binding — addressed to the Project thread at admission, it never mints a Main card and needs no conversion button.

A Project chat has one compact navigation pointer in its existing status bar (project_work_pointer.js). projectWorkTarget selects the last connected non-child card in the existing registry order that is unfinished, or the last represented root when all are finished; it creates no card or execution state. The pointer updates through Chat's existing stable-viewport mutation seam. Clicking scrolls only that chat's messages container to the selected card and records the existing reading intent; it does not select the recipient of the next message. Unless historyWindow.complete is explicitly true, the adjacent note says Loaded messages only; an empty represented set disables the button as No task card in loaded messages, never a claim that the Project has no work. The binding's disposer removes its one listener and button/note with the chat instance. No separate task pane, history fetch, poller or persisted pointer state is introduced.

The New Project dialog supports exactly one source: no folder, a fresh managed genesis workspace, an attached existing folder, or a cloned Git URL. The selected folder remains visible independently of the directory being browsed, and submission uses that selected target. Its shared focus boundary leaves source/creation and non-confirming cancellation semantics with project_create.js. Attach uses a server-side directory browser so the flow works in web and Docker too; a non-git attached folder is rejected unless the owner explicitly requests attach-snapshot initialization — never initialized silently; clone failures distinguish missing credentials. All three disclose that Project tasks receive read, write, and shell access in the chosen folder; provenance remains a durable historical fact, not recomputed from current git state.

Deleting a Project is lifecycle work, not filesystem deletion: the server fences new admission, cancels and quiesces the Project task subtree, then tombstones the registry entry; the UI may acknowledge that deletion started but does not claim completion early. Project id, canonical history, task bindings, memory, provenance, and the working folder remain preserved — deleting the row is not permission to erase the owner's repository or the agent's history.

Project unread state is the durable comparison visible_revision > project_seen_revision. Only owner-visible assistant/result content, delivered media, or a real incident advances the visible revision. Opening a panel does not clear unread by itself: the browser refreshes and paints the exact revision, verifies the panel is still visible and connected, then posts that revision as acknowledgement; the server clamps and max-merges it so a stale tab cannot move the cursor backwards or acknowledge future output.

A chat instance has an explicit resource lifecycle: destroy() marks it dead, disposes subscriptions, listeners, observers, and timers, and removes the DOM last; late async work checks the destroyed flag. app.js keeps at most one live Project chat instance; closing or switching stashes scroll intent and destroys it. The narrow exception is client state the server cannot reconstruct (staged File objects or an upload in flight): such an instance is hidden and marked pending, reused on reopen, and returns to the destroy policy after that work settles; typed but unsent text survives separately in per-thread session storage. This keeps hidden Project rooms from accumulating listeners or acknowledging unseen revisions without discarding data.

All five task-event sources (progress, chat, events, tools and supervisor) read their retained archive chains plus live files. Interactive history uses bounded, archive-aware readers until the requested thread's filtered quota is met; Project history cannot be satisfied by unrelated Main rows, and display reads avoid materializing or rebasing artifacts. Task-event lineage discovery, task-list ordering and Main's newest-result selection share gateway/task_list_scan.py's compact process-local memo: dev/inode/size/mtime/ctime invalidate changed files, failed or torn reads are never cached, and full selected results still pass the schema reader. Directory enumeration remains proportional to the number of results. Each living event follower retains its own per-file proven child bindings only across failed reads; valid parent/role changes, schema refusals and deletion remove them. Nonfatal history_gap frames with lineage_incomplete distinguish retained bindings from unknown membership and do not invent a view change. Fresh connections never treat client cursor paths as proof. These diagnostics do not consume a legacy history rank or acknowledge log bytes. Main's dialogue projection reads twenty nonempty text rows from a live tail and at most two archives, retaining message ids and the 500-character text preview; its unvisited-row count is explicitly unknown.

POST /api/tasks/{id}/events is the read-only SSE v2 transport (TaskEventsRequest): JSON {v:2, wait, cursor} carries {v,seq,view,positions}, with one byte offset per root/source across the complete archive+live chain. Events use physical root/source order, never timestamp ranks; each log row carries a stable root/source/byte-position event_id and the cursor after THAT row. Filter changes (task ids, roots, task_done suppression or a proven creation floor) emit cursor_replay and replay the new view. The CLI advances only after consuming a frame and deduplicates its most recent 4096 log identities; older replay duplicates remain possible. A cursor_checkpoint at wait expiry also advances skipped bytes without incrementing delivery sequence or claiming task progress. Live partial lines wait; immutable malformed/partial lines emit history_gap and advance. Unreadable required segments or shorter chains emit cursor_unavailable without resetting. SSE uses the existing JSONL chain helper with one source-bound metadata snapshot per pass: chain sizes/prefix offsets are collected once, selected segments are found by binary search, and the original live generation is rebound by inode after rotation. Consumed prefixes are stat'd without opening, at most the live file and one archive are held, and each 64 KiB-plus-final-line buffer is closed before network delivery. Each source pins its logical chain EOF for one pass, even through rotation: later appends belong to the next pass, so they cannot starve other sources or result/checkpoint handling. A row cut only by that snapshot boundary stays pending even if its file rotates before the next batch. The default whole-chain helper API remains available to custody/accounting readers. POST structure uses the existing derived validate_ingress contract, with semantic cursor checks afterward. Archives must remain immutable and retained: manual prefix removal masked by subsequent growth is outside the offset guarantee. Only explicit creation facts may skip older archives, and skipped bytes still count in positions. Fresh host UUID allocations in HTTP task creation, tool scheduling and supervisor scheduling stamp created_at before preparation; supplied/restored ids, legacy rows and first-at-final results are not backfilled. Those records conservatively retain a full cold scan. Result-directory failures remain explicit: SSE emits an error without inventing an empty view, Main marks result omissions unknown with the source error, and the legacy list wrapper keeps its fail-soft return with a warning. Each connection emits a fresh task-result projection and materializes the terminal result once; synthetic results have no log identity and are never deduplicated. Legacy GET and iter_task_events keep timestamp-sorted integer ranks, whose retroactive insertions may repeat or omit rows across reconnects; a from-zero replay recovers retained history. The CLI falls back to GET only on a first-connection HTTP 405. Subagent admission/custody dedupe remains the supervisor queue's transition (§5).

The agent-facing chat_history reader uses the same live-plus-rotated timeline and may narrow by exact provider, account, conversation, thread, actor, and inclusive date bounds before the count/offset/text-search window — presence provenance is searchable as structured transport fact, not only flattened prose.

Files

Files is a full gateway-backed file manager, not a chat attachment picker: directory navigation, breadcrumbs, filtering, image and sandboxed PDF preview, text preview/editing, an explicit binary/unsupported state, file and directory creation, save, drag-and-drop upload, download, open in the default OS application, copy, move, paste, and recursive delete. The desktop host bridge and web fallback share one download contract. Unsaved text is guarded on selection, directory change, navigation, and unload; a cancelled leave keeps the current document/selection, and ordinary directory refresh keeps its editor. Save exists only for a writable complete text read — a truncated preview stays read-only — with one in-flight write for pointer and keyboard submission. Save failures and clipboard feedback occupy a sibling status region rather than replacing editable content. The backend is the path authority (root confinement and symlink policy in §4); UI path strings and disabled buttons are presentation only.

Skills and Widgets

Skills has three views: installed skills, ClawHub, and OuroborosHub. Marketplace panes initialize lazily; the header Refresh dispatches to the active view's existing callback. The installed view merges extension truth with the serialized lifecycle queue so install, update, review, dependency work, enable, disable, repair, uninstall, and failure stay visible while queued or running. A failed primary read keeps a labelled previous list or an unavailable state, never a successful empty list. Optional catalog/lifecycle enrichment is independent; late hub badges update only their own nodes. ClawHub cannot derive Install from an unavailable installed-state read. Terminal failed work retains its domain error/actions without the computation animation.

Installation, deterministic preflight, LLM review, owner grants, dependency readiness, extension loading, enablement, and execution are separate lifecycle facts: a fresh review does not imply granted keys or installed dependencies, and enabled=true does not override a blocked review or load error. Owner attestation, where eligible, skips only the expensive LLM review; deterministic preflight and post-pass reconciliation still run. Repair and run creates an ordinary managed development task visible in Chat. The full confirmed request is recorded through normal owner-message ingress and its exact origin rides the task. After review and prerequisites, the model enables and tests the selected installation, leaving it working. A later direct owner disable wins over the older request; automatic repair-and-review carries no such enable authority. Hub publication uses the selected-skill preflight and ordinary managed task; the passive Installed projection neither runs Betterleaks nor claims publication readiness.

Widgets is a separate page because extension UI is an execution surface, not catalogue metadata. It renders only UI tabs registered by reviewed live extensions, read from a dedicated passive projection, GET /api/widgets (live tabs and their owners' revisions under one loader lock; each card carries the owning skill's live payload content_hash as revision, a change-signature fact, not an ETag — the Skills page stays on /api/extensions). On every entry the page paints the last known cards first, then fetches the list and compares an order-independent signature (key, render declaration, span, title/icon, ws_prefix, revision): an unchanged list moves no <article> node; a changed list is patched by key, a running card whose entry changed being stopped in order before its replacement mounts. The same reconcile runs on a visible extension_lifecycle event and on every WebSocket (re)connect. A failed list read shows contextual Retry beside its error; it invokes that same keyed reconcile, preserves unchanged frames and the page-session owner Stop choice, and disappears after success. There is no list polling or global Refresh/remount control, so retrying discovery does not stop a kept-running program behind the owner's back. Three modes share one capability set (sandbox="allow-scripts allow-pointer-lock allow-downloads", allow="autoplay; fullscreen; clipboard-write", never allow-same-origin): an extension-route iframe (kind: iframe) is the skill's own page with no bridge, and must declare the route it loads — the frozen registration contract still accepts an omitted route, which paints a not-supported card rather than a widget; a declarative widget is rendered by host-owned code from a validated schema; a reviewed module widget runs in an opaque-origin srcdoc iframe under a document CSP that web/modules/widget_module.js builds from the page origin — default-src 'none', scripts inline/blob:/the skill's module prefix with 'wasm-unsafe-eval', workers from blob:, images/media/fonts from data:/blob:/the skill's route prefix, no connect-src. Framed declarations may set a bounded height (320–8,192 px); a module without one starts at the floor and reports its #root content height through the nonce-bound bridge, capped by an optional module-only max_height; legacy route iframes stay explicit-height-only because an opaque document cannot be measured; geometry keys are rejected for declarative renders. Card order and the owner's per-card launch-policy override (widget_start_mode) are host UI state via /api/ui/preferences; the masonry (web/modules/masonry.js) packs cards in that key order and writes only --masonry-* custom properties, so a reorder never moves a node or reloads a frame (disclosed residual: Tab order follows the DOM until a reload). Engineering rules: docs/DEVELOPMENT.md "Embedded surfaces declare geometry and refresh semantics", "UI resources carry a disposer", "Declarative widgets".

Declarative widgets support forms and actions, status/data/text/code/markdown, tables, tabs, charts, polls, jobs, streams, subscriptions, progress, media, files, maps, calendars, kanban, and composition through group, metric, and callout. One recursive validator limits the tree to depth 8 and 256 nodes and reports the exact failing path. Nested interactive components use an explicit id or stable tree path as identity. Within one fixed declaration mount, data updates reconcile component/field nodes in place, preserving input identity, selection, composition, native popup state and live password text; retained form-value snapshots exclude passwords. Chart-owned canvas/parent nodes retain their lifecycle while data updates. Forms/actions carry their own keyed pending/result/error feedback, including job settlement, without requiring an authored status component or borrowing a sibling action's outcome. subscription.render remains transitively passive so an incoming event cannot smuggle a new active control tree past validation. Text, attributes, links, media routes, and field values are escaped or constrained for their actual sink.

Module widgets receive one parent-mediated I/O bridge on a per-mount nonce — the frame's only scriptable network path, since connect-src is closed: OuroborosWidget.fetch (also the frame's fetch) posts the request, the parent accepts only the exact owning prefix under /api/extensions/<skill>/..., issues it with same-origin credentials, refuses to follow a redirect, and streams the answer back as header/data/end frames from which the child rebuilds a real Response over a ReadableStream (binary by default; incremental reads work; no default timeout — the author's init.signal or init.timeoutMs aborts; declarative requests and the module source load keep a 25-second bound); the skill's namespaced WebSocket events are forwarded the same way (OuroborosWidget.onEvent, filtered by the card's ws_prefix). The bridge asks for the next body chunk on consumer pull. Out-of-process responses use the same child-to-host framed stream: ordered raw headers, standard HEAD/Range handling, cancellation until completion and separate diagnostics for cleanup failure after delivery; there is no total or pre-header timer. The existing loaded bundle owns cancellation through supervised_futures, and only the dispatched instance is affected by unload. Module download calls and ordinary supported export links reuse the common native/browser save owners; route URLs stream without building a Blob. External links use ui_helpers.openExternalViaHostBridge through the same source/nonce-bound bridge: trusted anchor clicks, OuroborosWidget.openExternal and the frame's no-handle window.open. Current user activation is checked where available; older hosts retain trusted anchor clicks. The host hands off to the native opener, ready Telegram SDK or browser before asynchronous work; noopener null is not proof of failure. Disposal clears link replies and hooks; sandbox/CSP and route-iframe behavior stay unchanged. The out-of-process WS push (POST /ui/ws-message) admits a 60-message burst reserve per skill that refills one message per second, refusing the excess with a typed 429 (§12). The module source endpoint (GET /api/extensions/{skill}/module/{entry:path}) authorizes against the live loader registration only and serves the declared entry or any reviewed sibling .js/.mjs from the texts captured when the bundle registered (an edit after load is not served until the skill reloads), answering every response with Access-Control-Allow-Origin: * because the requesting frame is an opaque origin. This keeps useful route I/O without giving reviewed skill JavaScript the SPA's cookies, DOM, or broad API authority. The same nonce carries the frame's fault channel: an in-frame script error, an unhandled rejection or a CSP violation is posted as one ouro-widget-error message (bounded, deduplicated, clipped) and reaches the card's own status slot, while the lifecycle state stays running because the frame is still mounted; the frame also exposes data-widget-content-height and data-widget-frame-capped so a widget pinned at its ceiling is distinguishable from one that painted nothing. Chart.js is bundled locally; rendering must not depend on a third-party CDN. Framed cards start under a launch policy — the owner's override over the author's validated render.start over the kind default (module and route iframe → manual, declarative → auto), a pure function in web/modules/widget_card.js — with one primary Start/Stop control and a policy menu (Auto / Manual / Keep running); retain keeps a framed card mounted while Widgets is hidden and ends it on Stop, on its skill leaving the live list, on a changed revision (stopped in order and re-mounted) and with the window, never with the server alone. A mounted widget owns its resources through one disposer: for a module widget that is the ordered stop with acknowledgement — dispose message posted, bridge kept answering the child's hooks, then abort/unlisten/remove on ouro-widget-disposed or after WIDGET_DISPOSE_ACK_TIMEOUT_MS (one second) — with one settle promise and one mount in flight per card key, so a remount waits for the pending stop instead of racing it; a route iframe disposes synchronously. Poll and WebSocket writers use monotonic progress per job so an older response cannot rewind a newer event; job polling keeps its job_id across bounded retryable failures while explicit terminal states stay terminal; transient list failure preserves the last good widgets.

Optional author controls use the SAME installed CSS and pure field/status source, not a second declarative renderer. ouroboros.server_web.read_author_kit_assets(request.app.state.repo_dir) reads the fixed web/ui.css and web/modules/ui_primitives.js files through the serving-root web resolver; it creates no endpoint, cache or state. docs/examples/author_ui_kit/ supplies two ordinary extension recipes: a module gets those texts at mount through its own GET route and the existing authenticated OuroborosWidget.fetch bridge, adds CSS and imports the pure module from a frame-owned Blob URL; a route iframe embeds the sources safely in its initial HTML under its own nonce CSP. Temporary Blob URLs are revoked. Neither path needs an opaque /static request or a change to auth, the bridge, sandbox or widget schema. The request-root input survives in-process and out-of-process dispatch. Authors opt controls into the named classes, use .ouro-ui for font/native-control context, and may override or omit the kit; there is no body reset, required conformance, hot-theme protocol or forced remount. New mounts load current source; retained frames keep their loaded copy. A failed kit load is shown inside that author application. Source delivery and real framed-consumer tests are tests/test_author_ui_kit.py and tests/test_author_ui_kit_browser.py.

The extension response handle retains the existing process and bundle contexts through asynchronous startup and teardown. Blocking context entry/exit, process registration, pipe shutdown and termination run off the ASGI loop; cancellation waits for its exact startup worker before closing any spawned child. Process facts retain their existing thread-local publisher. runtime_limits.py, re-exported by config.py, owns the 64 KiB response chunk and two-second post-response cleanup grace; neither is a response deadline. The reader bounds one metadata frame to 512 KiB and one body frame to its chunk plus flag before allocating its payload; this does not bound a whole response. An abnormal child exit retains its bounded, sanitized stderr and exit code after the existing drain, without changing a successfully delivered body. Static module assets still come from the captured registration without a child.

Dashboard

Dashboard groups Logs, Evolution, Costs, Updates, and Activity. It uses the common tab binder while its activation callback owns panel visibility and loading. Sub-tabs activate their own loading and refresh policy instead of running every expensive reader while hidden. Charting uses the bundled Chart.js (§3 Widgets).

Logs merges live WebSocket log frames with bounded REST backfill from events, tools, progress, and supervisor logs: chronological order, dedupe against live and reconnect overlap, bounded grouped task cards, raw record on demand, and shared category/severity/review presentation from log_events.js. Failed backfill names the unavailable history sources in a separate status row while live events continue; a later successful read clears that gap. Clearing the visible panel does not delete the underlying logs.

Activity shows running and pending queue entries, background-consciousness state, and scheduled work, with mechanical controls only where that surface is authoritative: typed cascade cancellation, start/stop for background consciousness, enable/disable/delete for owner-managed schedules. Each Activity section reports its own failed read, retaining last-known content/actions where available; a failed queue read does not erase working background/schedule sections or claim an empty queue. A schedule reconciled from a skill manifest is read-only here — a direct edit would be overwritten by the skill lifecycle and falsely appear durable.

Costs is a projection of the physical-attempt ledger distinguishing confirmed, reserved, unresolved upper-bound, unknown/unmetered usage, open rows, and finality; unavailable data renders as unavailable, not $0. Breakdowns by model, key, model category, and task category are views over the same ledger. Total budget can hot-apply; the per-task value is a hard cost cap over the whole root tree of the next task, not an own-task soft warning. Raising a cap does not resume work that already finalized or paused.

Evolution shows current evolution and background-consciousness state, campaign objective/progress, queue/failure/budget information, and durable history. Starting a campaign uses the shared input dialog: cancellation starts nothing, confirmed empty input selects the backend's default autonomous objective, entered text becomes the objective. Light runtime mode disables self-modifying campaigns rather than presenting a control the backend will refuse.

Updates separates passive status from explicit mutation. Opening the page reads cached/local update and git state without touching the network; the passive read carries the last real check's timestamp and a minimal update_tx projection so a re-opened panel sees an assisted resolution in progress. One exported pure verdict (updates.js::updateVerdict(status, phase), unit-pinned by web/tests/update_verdict.test.js) renders exactly ONE action button whose label is always the real next continuation, from Check for updates through Restart now (the degraded case where the automatic restart callback failed; it posts the ordinary /restart command); "Up to date" is claimed only over an actual check result, a failed check keeps its actionable label, and unknown backend warning classes surface verbatim. The apply flow verifies the preflight (verifiedUpdatePlan) and confirms BEFORE applying, naming the path: a clean update restarts the server, a conflicting one starts the reviewed assisted task (model spend disclosed) whose progress lands in chat. The served-SHA decision remains the only page-reload authority; the boot-owned pending_boot_smoke and applying_replace phases keep the synthetic restarting state until a post-reconnect update_status_ready proves boot finalization returned. Typed apply-failure facts (reason, blockers, rollback/smoke state, stash note, assisted budget floor) reach the owner rather than one string. Recovery holds the separately confirmed replace action, "Save recovery point" (moves only this installation's local ouroboros-stable fallback, never the official QA feed — §8), and ONE restore list where local tags label the commits they point at (/api/git/log tags carry their peeled sha); unavailable, divergent, dirty, unsafe, failed-check, rollback, and restart-required stay visible as states, never as extra buttons. An update letter — one short markdown paragraph Ouroboros writes about the update on offer — rides the same status payload as an additive letter key and refreshes on the same update_status_ready event, so it costs no endpoint, socket type or poll of its own. updates.js::updateLetterView(status, phase) projects it beside the verdict without touching it: a pending letter reads "What's new", an applied one (HEAD equals the target, or the last check that described this HEAD recorded the target as already inside it — a divergent install applies through a merge commit) becomes "What changed in this version" because the letter outlives the update it describes, a superseded or relocated one keeps its text under the range it was written for, and a failed write — including a range git could not read at all, which is never recorded as "nothing to say" — shows the last good letter (related and labelled by ITS range, which may predate the one that failed) with the reason, or the reason alone. It renders through the sanitizing chat markdown pipeline and adds no second action: unmanaged, unreadable and never-checked states and the restart phase hide it rather than grow a control; a passive refresh or a running check keeps the last known paragraph, and unchanged text keeps its DOM. The sidebar update pill (update_status.js) is a pointer, not a second apply surface; visual rules: DESIGN.md §8.

Settings and onboarding

Settings has Accounts, Secrets, Models, Agents, Behavior, Advanced, and About tabs — a sequence from connections to runtime detail. Accounts: managed subscriptions and their shared service banner, API providers, custom compatible endpoints, local runtime entry points, and the optional non-loopback network gate. Secrets: known provider/integration secrets, skill-requested keys, and owner-defined custom keys, without returning stored values. Models: compact source/model/account role rows, ordered fallbacks, context assertions and effort lanes. Agents: task actors and review lanes, with delegation permissions, per-root and depth limits and subagent path roots; their accounts are managed in Accounts. Behavior: context, safety-supervisor coverage, task acceptance, self-evolution, prompt-cache posture. Advanced: process, timeout, local-model, integration, source-control, and cleanup controls (worker count is process capacity, so it lives here). About reports application/runtime identity.

The Settings client collects and validates the whole current draft before sending Save; any local error keeps all rows/values available for correction and sends no partial save. settings_controls.js separates pure custom-key collection from field-error painting and keeps dirty reads passive. Ordinary refresh and failed writes preserve current edits; leaving or explicitly reloading a dirty draft asks before discarding it. The existing write response still distinguishes saved, unsaved and unknown outcomes, with owner-only decisions on their own endpoints. No durable cross-page draft store or secret persistence is introduced.

Each provider card has one compact Test action backed by POST /api/providers/test: request exactly {provider_id, overrides?}, response exactly {ok, error?}. Overrides are request-local — an omitted field reads the saved value, an explicitly edited empty field stays empty (for the compatible card it suppresses the corresponding legacy OpenAI fallback) — and draft credentials never mutate Settings, process environment, or LLMClient caches. Model selection reuses a configured route, then the provider's maintained main default; only the generic OpenAI-compatible route performs bounded catalogue discovery when neither exists. A configured test is one bounded, physically accounted probe (ouroboros/llm_probe.py) with none of the normal-chat retry/fallback/tools/reasoning/web/cache/response-format/capability-learning paths. The card renders Testing…, Works, or Not ready with one controlled short reason; the tooltip notes the request may incur provider charges.

Desktop onboarding and the blocking web overlay are the same served /onboarding page — same steps, same backend normalization; context mode stays a separate owner setting. Startup readiness is structural (§2): a configured managed-model Main route can satisfy the gate without an API key; credential validity, entitlement, model availability, and local-process health remain runtime status, not onboarding admission. Every host completes through the single POST /api/onboarding/complete transaction (§2), so completion is all-or-nothing and install-time agent defaults are part of the same save.

Accounts is the owner-facing projection of Ouroboros's owned Claudexor daemon; the browser never receives its control token or interprets credentials. Status combines daemon/runtime readiness, login-capable harness discovery, credential profiles, honest vendor-live versus local-session verification, fresh quota windows, and optional model discovery. One /v2/quota envelope supplies both windows and typed per-subject absences (quota, quota_absences), so a missing usage reading stays separate from login truth; only fresh snapshots are eligible for percentages or exhaustion; a fully-used ratio without a valid future reset stays non-blocking and visibly unproven, while an explicit active cooldown remains effective, a malformed absence list becomes empty rather than breaking the response, and redacted absence detail never selects semantics. API-key-only adapters gain no fake Login buttons merely by appearing in a broader execution catalogue.

Accounts group into one card per agent family: family name, an aggregate status counting the accounts rotation can actually use (signed-in AND enabled), a fail-safe "Next up" badge naming who an unpinned run would take (the store's one dual-wire reader: unified accountPools first, legacy per-harness next_up second), and the card's add action. Rows are ONE type on both engine generations: on a UNIFIED engine (the server-stamped unified_accounts feature fact) every account — migrated default logins included, under reserved <harness>-default registry ids — is a named row with Enabled toggle (the engine's per-profile PATCH) and Remove; on a LEGACY engine the native pseudo-row keeps the same layout and only its ACTIONS differ — no Remove, no toggle, because that engine has no route for either and a dead button would claim an effect this process cannot have. Status wording is honest: verification=not_run is neutral unknown, and explicit availability=unknown re-runs the shared Refresh rather than a new sign-in — an auth probe failure is never proof of logout. The legacy ''-keyed quota alias is granted only to the literal reserved <harness>-default id (exact-keyed readings always win); pinned route health applies no alias, failing open to the engine's own typed refusal. Removing a named account preserves the complete deletion receipt: a refusal remains a refusal, and a vendor-owned / left-unchanged / OS-user disposition becomes a retained-credential warning, not a false sign-out claim. One service banner explains a daemon or runtime problem once, per facet, instead of decorating rows. Per-facet independence is real on both sides of the wire: claudexor_accounts.py fans the catalog, account, and quota reads out independently and stamps each classification into the payload's reads block, so one refusal does not collapse its siblings into a global unreachable verdict; the client's shared status store reads that stamp, and only a legacy payload without it is read coarsely.

Connect is link-first and harness-agnostic. A typed disclosure renders the sign-in URL and any one-time code; flows that may need a pasted callback code keep that field visible while active. Engines publish optional-without-default setupLogin per exact harness row (in_app → omitted setup transport, external_terminal → client_pty; malformed present data is a capability gap; explicit null delegates to the exact pinned engine's typed setup/profile admission). Typed facts select every continuation, prose never does: credential_profile_required plus add_named_account selects the name-the-account face, a typed duplicate profile is idempotent, typed pre-job refusals prove setup custody absent/released, and the exact terminal_transport_* code (or durable job.nativeCommand.errorCode) offers the explicit external-terminal continuation. Only the engine's structural missing-vendor-binary first-create job lets the same owner action also consent to one hidden local install through the exact managed Claudexor CLI and one retry; success requires Claudexor's strict post-install proof (absolute installedBinary plus bounded installedVersion) — exit zero without proof is refused, and a second login refusal is returned rather than looped. A new explicit client_pty job exposes its labelled attach command immediately; there is no embedded terminal login surface. Terminal job state and the current account row reconcile so a stale verification read during login cannot claim failure after the account connected.

Account status refresh runs immediately and on visible page/tab activation; hidden pages do not pay for daemon round-trips. Entering Agents is an explicit owner action: after the fresh read, an already-provisioned stale home restarts through the wake endpoint, while not_provisioned, foreign-owned, and repair states stay behind Connect; background polling is read-only and never wakes the daemon. Job polling uses one request at a time, backs off on consecutive failures, and after ten stops with an honest unconfirmed state — lost contact proves neither failure nor settlement. One transition lock covers Start, Retry, and Dismiss, and a new login begins only once release of the prior job is proven (loginReleaseProven): a 2xx cancel alone is not proof, a network/server failure retains the card and job id because dropping it could orphan a live server job, and after each await the handler rechecks whether polling settled the job so a stale cancel continuation cannot overwrite a terminal result.

Review lanes edits one structured reviewer configuration. Each triad, scope, optional advisory, or deep self-review row picks its reviewer from ONE flat select: Available-subagents roster rows lead as references, then the inline channels — API delivery or a coding-agent session — followed by model, optional credential profile, and effort. The two multi-row categories are driven by one CATEGORIES table and the two single-row categories (advisory, deep self-review) by one renderer/binder parameterized by a SINGLETONS table; the deep self-review block states its one difference from the advisory where the owner picks — an API model there is ONE packed review (Atlas + memory), not an inspection episode — and, while the row is only synthesized from OUROBOROS_MODEL_DEEP_SELF_REVIEW, says so — an untouched synthesized (or empty) placeholder is OMITTED from the save payload, so an unrelated save never writes the key's value into the setting and a repair save beside a config_error (the endpoint still attaches the synthesized row there) succeeds without it; editing the row materializes it, and a blanked model box is hinted client-side and refused typed at save (owner fork 3 = A). The former Models-tab field is gone (R7). Saved choices that disappear from discovery stay visible as unavailable rather than silently changing; capability labels configure nothing, and server-returned limits plus last effective execution disclose what a saved row actually ran as. An unloaded view authors no replacement. The pure serializer preserves a successfully loaded empty triad/scope; the Settings local validation gate reports that invalid draft before POST, while the backend independently enforces the same configuration contract. A row pinned to an account discovery no longer lists keeps its pin, disclosed once above the rows — and only on the word of a facet actually read (account pins answer to accounts, models to catalog; an unread or failed facet yields "not checked", not "not in discovery"). The standing note states the rule, never the situation: every review surface — commit, scope, plan, advisory, skill review and task acceptance — follows its configured rows and waits for subscription capacity rather than falling back to API spend (owner R2 retired the task-acceptance API pin and its default-panel disclosure).

Available subagents is the single task-actor editor. Its list-level Enabled flag and at most ten stable rows are the saved OUROBOROS_SUBAGENTS intent. The owner sees numbered compact cards (docs/DESIGN.md §6 row anatomy) and authors one prose field, Description (recommended_use), beside the structured API-model or Agent-session route, optional effort, and optional managed-model/session account pin (empty pin = Claudexor's compatible-account rotation). Each card's status dot composes two axes — intent (Saved / Draft / Generated) and availability (Available / Not checked / Unavailable / No account / Limit reached, or Checked at start for an API-model route) — with the dot's tone the worse of the two (subagent_status_primitives.sessionRouteVerdict / rowStatus). Identity is the stable internal subagent_id plus live route facts; parse accepts and drops a legacy display name, and a visual ordinal never becomes durable identity. The editor shares only neutral route/model/account/status primitives with Review lanes. Add and Duplicate reveal the new entry through the shared ui_helpers.revealNewRow. A fresh entry is not an error: its meta line carries a neutral hint until a save attempt (noteSaveAttempt; validate() stays pure), after which the section line summarises and the offending card is tinted and names its error, reconciled in place by one painter; in the wizard the Finish error shows on the summary step with the card already tinted.

Saved intent and live evidence are separate axes. Row status distinguishes saved/generated/draft intent from current availability and may show the last requested→effective run evidence; a status/catalog/accounts failure annotates rows, never erases them. Settings GET may offer an unsaved migration/default candidate when no canonical value exists; the editor materializes it only on Save. A late bounded status or preview response may replace a still-clean generated baseline, never an owner-edited draft, and unchanged repaint preserves focus/caret. Generic Settings save strictly validates and canonicalizes the materialized value before the serialized off-event-loop owner transaction. A running task keeps its immutable start snapshot; the save response says changes apply from the next task.

Mutative subagents use Off, Auto, and On; an explicit value applies to every acting surface, while Auto delegates the default to runtime mode — surface-aware in Light, which allows only children building outside the Ouroboros runtime (external workspace, genesis); why: §6 Safety and runtime mode, exact matrix ouroboros/config.py. Read-only children stay available. Owner-controlled, applies from the next task, independent of API vs coding-agent route.

Prompt Cache TTL is one global owner choice — provider default, five minutes, or one hour — applied to every lane so task, review, and safety builders cannot drift into conflicting cache horizons. The provider-send finalizer applies it only to existing cache markers on compatible Anthropic-family payloads; the UI does not promise cache behavior on providers that manage it implicitly. The shipped one-hour posture favors reuse across long waits and review cycles.

Settings save classifies effects into three classes rather than claiming everything became live at once (exact key membership: gateway/settings.py and config defaults). Hot-apply — e.g. total budget, tool timeout, MCP configuration (a failed hot-reconfigure surfaces as a save warning); retained soft/hard timeout keys are accepted only as deprecated audited no-ops and reported as retired. From the next task — e.g. models, credentials, reviewer/subagent configuration; a running task keeps its starting snapshot and the response says so. Restart-required — e.g. worker count, bind host, the skills repo path (pooled workers load the extension registry once at spawn); such a save offers Restart now over the owner /restart command. A failed task-start settings reload is disclosed as a persisted, chat-visible task_start_settings_reload_failed event. Ordinary runtime/context controls keep dedicated owner paths. Cyber generic saves may configure runtime access and Supervisor through the same writer with audit; access remains pending until restart. Context mode remains on its owner-only path, and the agent cannot self-switch review scope/enforcement. Every in-process owner-settings writer holds the shared settings_document_mutation() lock across its read/merge/write; the file lock is a write precondition, not a substitute. Loading reviewer and delegation settings waits at most the boundedStatusRefresh two-second foreground beat for Claudexor; a cold refresh continues and repaints when it lands. Backend failure and browser transport failure stay distinct; absence of a successful status read is never evidence a runtime is healthy.

Visual verification policy

A visible change is exercised in at least one relevant real consumer flow and the rendered result is inspected with vision; a saved screenshot alone is not verification. Mobile, WebKit, additional browsers, and special viewports are selected from the actual interaction risk — not a universal matrix (reviewer items: docs/CHECKLISTS.md). No visual-QA runner, endpoint, ledger, or mandatory device matrix is introduced by this policy.

4. Server API Endpoints

If OUROBOROS_NETWORK_PASSWORD is configured, non-loopback HTTP and WebSocket access requires authentication; loopback clients bypass the gate, and /api/health plus the middleware-owned login/logout paths stay reachable. Browser sessions use a server-keyed, expiring HttpOnly HMAC cookie; Secure is set only under TLS so a plain-HTTP LAN session does not enter a login loop. An unauthenticated WebSocket is closed with code 4401 before ws_endpoint accepts it. With no configured password, non-loopback access remains open by explicit operator choice.

The executable browser/CLI route SSOT is ouroboros/gateway/router.py; file-browser routes are contributed by gateway/files.py::file_browser_routes(). gateway/contracts.py is the frozen descriptive envelope and endpoint index mirrored by web/modules/api_types.js and parity tests; its TypedDict classes perform no runtime JSON validation. The loopback Host Service is a separate token-authenticated app assembled by gateway/host_service.py::create_host_service_app, not another public owner API.

Every /api/files/* operation resolves its requested path and refuses the operation when that resolution leaves the configured file root. In-root symlinks remain usable; out-of-root symlinks may be listed with is_symlink: true but cannot be read, written, downloaded, deleted, or traversed. The backend check is authoritative regardless of browser path presentation.

Method Path Handler
GET / server.index_page
GET /api/health gateway.state.api_health
GET /api/state gateway.state.api_state
GET /api/extensions gateway.extensions.api_extensions_index (unique rows additionally carry content_hash, published (validated receipt object or null), published_malformed; identity-collision rows carry identity_collision: true and omit the receipt fields)
POST /api/skills/{skill}/publish-preflight gateway.skill_publish.api_skill_publish_preflight
GET /api/extensions/{skill}/manifest gateway.extensions.api_extension_manifest
GET /api/extensions/{skill}/module/{entry:path} gateway.extensions.api_extension_module (live-registration authorization, reviewed .js/.mjs siblings from captured texts, Access-Control-Allow-Origin: * on every answer)
GET /api/widgets gateway.widgets.api_widgets (passive projection of the loader's live UI tabs via extension_loader.live_widget_projection; Cache-Control: no-store)
GET /api/extensions/{skill}/settings_section gateway.extensions.api_extension_settings_section
ANY /api/extensions/{skill}/{rest:path} gateway.extensions.api_extension_dispatch
GET /api/skills/daemons gateway.extensions.api_skill_daemons
POST /api/skills/{skill}/toggle gateway.extensions.api_skill_toggle
POST /api/skills/{skill}/delete gateway.extensions.api_skill_delete
GET /api/skills/lifecycle-queue gateway.extensions.api_skill_lifecycle_queue
POST /api/skills/{skill}/review gateway.extensions.api_skill_review
GET /api/skills/{skill}/review-history/{job_id} gateway.extensions.api_skill_review_history_detail (bounded lazy detail from a fixed tail window of review_history.jsonl; missing job 404, outside-window honestly unavailable; slot/attempt usage joins from the physical-attempt ledger)
POST /api/owner/skills/{skill}/attest-review gateway.extensions.api_owner_skill_attest_review (OWNER-ONLY skip of the LLM review; the deterministic preflight floor still runs, 409 on failure; routes through run_skill_review_lifecycle for the post-pass reconcile)
POST /api/skills/{skill}/grants gateway.extensions.api_skill_grants
POST /api/skills/{skill}/reconcile gateway.extensions.api_skill_reconcile
GET /api/marketplace/clawhub/search gateway.marketplace.api_marketplace_search
GET /api/marketplace/clawhub/installed gateway.marketplace.api_marketplace_installed
GET /api/marketplace/clawhub/info/{slug:path} gateway.marketplace.api_marketplace_info
GET /api/marketplace/clawhub/preview/{slug:path} gateway.marketplace.api_marketplace_preview
POST /api/marketplace/clawhub/install gateway.marketplace.api_marketplace_install
POST /api/marketplace/clawhub/update/{name} gateway.marketplace.api_marketplace_update
POST /api/marketplace/clawhub/uninstall/{name} gateway.marketplace.api_marketplace_uninstall
GET /api/marketplace/ouroboroshub/catalog gateway.marketplace.api_ouroboroshub_catalog
GET /api/marketplace/ouroboroshub/installed gateway.marketplace.api_ouroboroshub_installed
GET /api/marketplace/ouroboroshub/preview/{slug:path} gateway.marketplace.api_ouroboroshub_preview
POST /api/marketplace/ouroboroshub/install gateway.marketplace.api_ouroboroshub_install (also the adopt transport: {adopt: true, expected_content_hash} replaces an external same-name occupant with the sha256-verified catalog payload; adopt forces auto_review, conflicts with overwrite, typed 400/409/502 codes ride the lifecycle payload)
POST /api/marketplace/ouroboroshub/update/{name} gateway.marketplace.api_ouroboroshub_update
POST /api/marketplace/ouroboroshub/uninstall/{name} gateway.marketplace.api_ouroboroshub_uninstall
POST /api/marketplace/ouroboroshub/publication/{name}/clear gateway.marketplace.api_ouroboroshub_clear_publication (compares the displayed receipt and clears only local waiting state)
GET /api/files/list gateway.files.api_files_list
GET /api/files/read gateway.files.api_files_read
GET /api/files/content gateway.files.api_files_content
GET /api/files/download gateway.files.api_files_download
POST /api/files/upload gateway.files.api_files_upload
POST /api/files/mkdir gateway.files.api_files_mkdir
POST /api/files/write gateway.files.api_files_write
POST /api/files/delete gateway.files.api_files_delete
POST /api/files/transfer gateway.files.api_files_transfer
GET /onboarding gateway.onboarding_host.onboarding_page
GET /api/onboarding gateway.settings.api_onboarding
POST /api/onboarding/complete gateway.onboarding.api_onboarding_complete
POST /api/onboarding/subagents/preview gateway.onboarding.api_onboarding_subagents_preview
GET /api/settings gateway.settings.api_settings_get
POST /api/settings gateway.settings.api_settings_post
GET /api/reviewer-slots gateway.settings.api_reviewer_slots
GET /api/claudexor/status gateway.claudexor_accounts.api_claudexor_status
POST /api/claudexor/quota/refresh gateway.claudexor_quota.api_claudexor_quota_refresh
POST /api/claudexor/wake gateway.claudexor_accounts.api_claudexor_wake
POST /api/claudexor/login gateway.claudexor_accounts.api_claudexor_login
GET /api/claudexor/login/{job_id} gateway.claudexor_accounts.api_claudexor_login_job
DELETE /api/claudexor/login/{job_id} gateway.claudexor_accounts.api_claudexor_login_job
POST /api/claudexor/login/{job_id}/input gateway.claudexor_accounts.api_claudexor_login_job
POST /api/claudexor/login/{job_id}/reconcile gateway.claudexor_accounts.api_claudexor_login_job_reconcile
DELETE /api/claudexor/credential-profiles/{harness}/{profile_id} gateway.claudexor_accounts.api_claudexor_credential_profile
PATCH /api/claudexor/credential-profiles/{harness}/{profile_id} gateway.claudexor_accounts.api_claudexor_credential_profile
POST /api/owner/runtime-mode gateway.settings.api_owner_runtime_mode
POST /api/owner/auto-grant gateway.settings.api_owner_auto_grant
POST /api/owner/context-mode gateway.settings.api_owner_context_mode
POST /api/owner/safety-mode gateway.settings.api_owner_safety_mode
POST /api/owner/skills/{skill}/presence-runtime gateway.presence_settings.api_owner_skill_presence_runtime
POST /api/owner/capability-ack gateway.settings.api_acknowledge_capability
GET /api/ui/preferences gateway.ui_preferences.api_ui_preferences_get
POST /api/ui/preferences gateway.ui_preferences.api_ui_preferences_post
GET /api/model-catalog gateway.models.api_model_catalog
POST /api/openai-compatible/models gateway.models.api_openai_compatible_models
POST /api/providers/test gateway.models.api_provider_test
POST /api/tasks gateway.tasks.api_tasks_create
GET /api/tasks gateway.tasks.api_tasks_list
GET /api/tasks/{task_id} gateway.tasks.api_task_get
GET /api/tasks/{task_id}/events gateway.tasks.api_task_events (legacy integer rank)
POST /api/tasks/{task_id}/events gateway.tasks.api_task_events (read-only v2 cursor)
GET /api/tasks/{task_id}/artifacts/{name} gateway.tasks.api_task_artifact
POST /api/tasks/{task_id}/cancel gateway.tasks.api_task_cancel
POST /api/tasks/{task_id}/hurry gateway.tasks.api_task_hurry
POST /api/tasks/{task_id}/resume gateway.tasks.api_task_resume
POST /api/decisions gateway.tasks.api_decision_answer
GET /api/schedules gateway.schedules.api_schedules_list
POST /api/schedules gateway.schedules.api_schedules_upsert
DELETE /api/schedules/{schedule_id} gateway.schedules.api_schedules_delete
POST /api/command gateway.control.api_command
POST /api/reset gateway.control.api_reset
GET /api/git/log gateway.control.api_git_log
POST /api/git/rollback gateway.control.api_git_rollback
POST /api/git/promote gateway.control.api_git_promote
GET /api/update/status gateway.control.api_update_status
POST /api/update/check gateway.control.api_update_check
POST /api/update/preflight gateway.control.api_update_preflight
POST /api/update/apply gateway.control.api_update_apply
GET /api/cost-breakdown gateway.history.make_cost_breakdown_endpoint (the router's import path; the factory body and _ACCOUNTING_SUMMARY_FIELDS live in gateway.cost_breakdown)
GET /api/evolution-data gateway.control.api_evolution_data
GET /api/projects gateway.projects.api_projects_list
POST /api/projects gateway.projects.api_projects_create
POST /api/projects/from-task gateway.projects.api_project_from_task
POST /api/projects/{project_id}/update gateway.projects.api_project_update
POST /api/projects/{project_id}/delete gateway.projects.api_project_delete
GET /api/fs/dirs gateway.projects.api_fs_dirs
GET /api/chat/history gateway.history.make_chat_history_endpoint
GET /api/logs/{name} gateway.logs.api_logs_tail
POST /api/chat/upload gateway.files.api_chat_upload
DELETE /api/chat/upload gateway.files.api_chat_upload_delete
POST /api/local-model/start gateway.models.api_local_model_start
POST /api/local-model/stop gateway.models.api_local_model_stop
GET /api/local-model/status gateway.models.api_local_model_status
POST /api/local-model/test gateway.models.api_local_model_test
POST /api/local-model/install-runtime gateway.models.api_local_model_install_runtime
GET /api/mcp/status gateway.mcp.api_mcp_status
POST /api/mcp/refresh gateway.mcp.api_mcp_refresh
POST /api/mcp/test gateway.mcp.api_mcp_test
WS /ws gateway.ws.ws_endpoint
STATIC /static/* server.NoCacheStaticFiles
GET 127.0.0.1:${OUROBOROS_HOST_SERVICE_PORT:-8767}/identity gateway.host_service._api_identity
GET 127.0.0.1:${OUROBOROS_HOST_SERVICE_PORT:-8767}/tools/schemas gateway.host_service._api_tool_schemas
POST 127.0.0.1:${OUROBOROS_HOST_SERVICE_PORT:-8767}/chat/allocate-internal gateway.host_service._api_allocate_internal
POST 127.0.0.1:${OUROBOROS_HOST_SERVICE_PORT:-8767}/chat/inject gateway.host_service._api_chat_inject
GET 127.0.0.1:${OUROBOROS_HOST_SERVICE_PORT:-8767}/chat/operations/{operation_ref:path} gateway.host_service._api_chat_operation (the calling skill's own accepted message: pending, running with its task or turn, the durable answer, or the terminal task status)
POST 127.0.0.1:${OUROBOROS_HOST_SERVICE_PORT:-8767}/chat/cancel gateway.host_service._api_chat_cancel (the existing cancellation owner on work that message started; a typed outcome, never a cancellation that did not happen)
POST 127.0.0.1:${OUROBOROS_HOST_SERVICE_PORT:-8767}/chat/decision gateway.host_service._api_chat_decision (the task_decision.answer_decision ingress relayed for a transport skill)
POST 127.0.0.1:${OUROBOROS_HOST_SERVICE_PORT:-8767}/presence/turn gateway.host_service._api_presence_turn
GET 127.0.0.1:${OUROBOROS_HOST_SERVICE_PORT:-8767}/presence/work/{work_ref} gateway.host_service._api_presence_work
POST 127.0.0.1:${OUROBOROS_HOST_SERVICE_PORT:-8767}/ui/ws-message gateway.host_service._api_ws_message
WS 127.0.0.1:${OUROBOROS_HOST_SERVICE_PORT:-8767}/events gateway.host_service._ws_events

Rationale: server.py owns process startup/lifespan/static mounting, while gateway/* owns browser-facing HTTP/WS contracts; this keeps UI and runtime coupling explicit and testable.

WebSocket protocol

/ws is the live browser delivery channel, not a durable state owner: queue, task, Project, review, skill, settings, cost, and update modules persist their own truth, and REST/history endpoints reconstruct it after reload or disconnection. gateway/contracts.py describes the frozen envelope shapes and message-type index for Python/JavaScript parity; gateway/ws.py performs the actual transport checks — incoming text must decode to a JSON object, extension types must parse as an owned namespace, and built-in chat or command frames must carry a non-empty payload before entering the message bridge. The authentication middleware above governs socket admission; the public socket never receives the Host Service token or the owned Claudexor daemon token.

The browser constructs one socket for the whole SPA. Feature modules subscribe before connection, and the initial complete Project chat-id set is fetched before the first open so an early Project frame cannot be mistaken for Main traffic. ws.on(type, listener) stores listeners in insertion-ordered sets and returns a disposer; emission uses a listener snapshot, so a listener added during dispatch does not receive the current frame and disposing one listener cannot skip its neighbor. Every decoded frame first reaches the generic message event and then its type-specific event, which lets Widgets consume reviewed namespaced events without duplicating the socket.

A browser chat frame contains the owner text and may add sender_session_id, client_message_id, force_plan, uploaded attachment references, chat_id, project_id, and client_surface — raw sending-surface observables measured at SEND time because the pywebview bridge appears asynchronously after load. The gateway normalizes that payload through client_surface.normalize_client_surface, stamps host received_at, and persists it on the canonical inbound row — distinct from the transport dict (transport is chat-scoped reply routing; the surface fact is per-message provenance). The fact is assembled at its PRODUCER, never inferred at render (the per-producer stamp catalog and closed-key bound: ouroboros/client_surface.py); synthetic A2A chats stamp no owner surface (machine traffic never wears one); machine producers stamp nothing (client_surface is a reserved schedule-template key rejected at admission); promotion/steering CARRY the originating owner turn's fact. The loop injects a surface note only when sending-surface identity changes within an attempt (viewport excluded — a resize is not a device change). Absence is an honest gap. The client generates a message id when absent and uses it to reconcile its pending bubble, the echoed canonical user row, routing annotations, and mailbox retries; a successful browser send() means only that the current socket accepted the frame, not that a task was durably admitted.

Ordinary frames sent while disconnected enter a process-local queue capped at 100 entries (oldest dropped), flushed in order after reconnect and lost on page reload — not a second durable outbox. Attachment messages deliberately set queue:false: uploads occur immediately before send, so retaining only the socket frame would leave unowned temporary files; on socket loss Chat refuses the message, cleans uploaded temporaries best-effort, and retains the staged files for explicit retry.

For a chat frame, the gateway validates uploaded filenames as basenames confined under the upload root, exposes the first eligible bounded image as native image content, forwards the complete validated attachment set as task-staging metadata, and calls the local message bridge with the exact thread, Project, sender-session, client-message, and planning facts. The web owner identity is fixed: chat_id selects a thread and cannot mint an external owner identity. If the bridge is not initialized, the socket returns a visible assistant warning rather than accepting the message silently.

A built-in command frame carries a slash command and enters the same bridge with rebroadcast disabled; runtime command routing, owner authorization, queue authority, and typed outcomes remain outside the socket module. Main header controls therefore reuse the ordinary command contract for Restart, Panic, review, evolution, and background consciousness; Panic is sent only after the shared dialog returns the strict confirmed boolean. The socket does not infer intent from command-looking prose.

Built-in outbound envelopes: chat, photo, video, document, typing, log, heartbeat, extension_lifecycle, message_annotation, projects_changed, task_named, and update_status_ready. Chat progress may carry task lineage, role, requested/effective model lane, delegated route, terminal execution evidence, review projection, cancellation eligibility, outcome axes, artifact references, and nullable cost/finality fields — additive presentation facts; consumers must not infer a missing execution receipt, cost, or task result from the absence of one optional field.

Thread routing is explicit. Project chat, typing, media, and log frames carry chat_id; a Project panel consumes its own thread, while Main admits required Project question pointers and the two host-stamped Project lifecycle rows (project_started, project_completion_summary). projects_changed carries a new chat id so every tab can extend its fan-out set before fetching the registry; when even that ordering loses the race, the server-stamped project_thread marker on the frame itself keeps Main from adopting it — set once at the message-bus broadcast choke from the registry (a membership lens, never a numeric range, so external transport ids such as Telegram stay unstamped) and enforced by Main's fan-out gate (chat_activity.mainThreadAccepts). Task-scoped LOG events acquire their final chat id at supervisor ingress: worker diagnostics carry only their own task_id, and supervisor/log_addressing.py::address_task_event stamps the audience from host-attested truth (the precedence chain lives in its docstring; an explicit event chat_id of 0 is the hidden partition, HIDDEN_CHAT_ID, never "missing"); direct/ephemeral turns carry their chat BY VALUE, stamped at the producer, because the registry entry dies with the turn while queued events drain later. Addressing is honest — an A2A row keeps its true audience, suppressed only at the broadcast choke (push_log) so machine traffic never reaches the browser; the same addressing runs in the server-process append sink and at every supervisor handler owning a suppressed type's explicit push, and a genuinely unaddressable event keeps the legacy chat-0 frame. message_annotation updates one canonical owner message without creating another bubble; task_named updates a card only where that task already exists. Media/document consumers validate MIME, base64, and download-route shapes before building browser URLs.

Extension WebSocket traffic is structurally namespaced by extension_loader.extension_surface_name() so an extension cannot shadow a built-in type. On each incoming extension frame the gateway resolves the owning skill and reconciles whether its extension is still desired, reviewed, granted, enabled, and live. A missing or failed handler returns a visible log frame. Out-of-process handlers execute in their extension child off the event loop; in-process handlers first record the required execution/cost disclosure. A non-None result returns as <request-type>.reply; exceptions become typed error log frames rather than terminating the socket loop.

Server broadcasts snapshot the connected-client list and send to all clients concurrently, so one slow or half-open browser cannot head-of-line-block delivery to every other tab. Failed sends remove only the dead clients and append a durable broadcast_partial_failure event; the original domain event stays owned by its durable producer. Restart shutdown closes remaining clients best-effort with code 1012 so they enter the ordinary reconnect path.

The browser reconnects with bounded exponential delay, shows the reconnect overlay, and resets the delay after a successful open; a watchdog closes an apparently open connection after 45 seconds without any inbound frame, so heartbeat traffic proves stream liveness rather than task progress. One served-SHA decision (ws.js decide(): keep / reload-changed / reload-unknown) governs both recovery paths so a transient drop cannot destroy in-page state: a changed or no-longer-provable SHA reloads (a restarted server must not keep old JavaScript or CSS alive in PyWebView), an unchanged SHA keeps the page and its queued outbound messages, and an unversioned /api/state stays on keep when no non-empty SHA was ever remembered — the owner-selected default under uncertainty, accepting possibly-stale assets as the disclosed tradeoff. While the socket stays down, delayed recovery probes consult /api/state without adopting the served SHA; a 200 whose body is not a parseable object counts as a failed probe, probes are single-flight and generation-scoped per disconnect episode, and after several consecutive healthy probes with the socket still down, one forced reload per episode remains as the fuse for a stale browser runtime.

Each Chat instance handles open by resynchronizing archive-aware durable history and close by withdrawing online/accounting presentation; reconnect deduplication covers overlap between live frames and REST replay, Logs merges the same way, and large history parsing runs off the server event loop. Delivery is live plus replay, not a promise that every transient frame is persisted: durable chat rows, task results, queue snapshots, Project revisions, review ledgers, cost ledgers, and lifecycle state remain the recovery authorities.

5. Supervisor Loop

server.py::_run_supervisor() is the single scheduler for pooled tasks. A healthy tick publishes liveness, rotates the runtime logs, checks worker health, drains worker, direct-chat, and consciousness events, accepts owner bridge input, enforces deadlines and schedules, runs throttled reconciliation and evolution admission, assigns eligible work, and persists state/queue_snapshot.json. Bridge intake precedes timeout, maintenance, evolution, and assignment work so a slow control-plane step cannot make a new owner message invisible. Three consecutive loop failures clear supervisor readiness, stop its watchdog generation, and notify the owner instead of leaving a healthy-looking server that no longer assigns work; a failure raised while a shutdown or restart is already in progress (the lifespan teardown sets a process-local stop event first and joins the loop for a bounded window before workers, bridge and event bus go down) is not a crash — the loop exits quietly, without the counter, the error, or the alarm — and the crash backoff waits on that stop event so a shutdown is never held by it.

PENDING and RUNNING, guarded by supervisor.queue._queue_lock, are the live task-lifecycle authority. Admission reserves identity before project, workspace, attachment, or routing side effects can create a duplicate; refuses a disabled pool, duplicate task, project deletion, accepted or sealed root, or exhausted root budget; attaches the task contract; and preserves stable priority order. Assignment runs against the same locked state and skips reaping slots, budget-paused work, closed project roots, conflicting project writers, and tasks exceeding the root's subagent capacity or depth reservation; evolution tasks are dropped there when evolution_block_reason() is set (Light runtime mode, supervisor/workers.py). That is the last of three evolution-only runtime-mode fences: owner and post-task entry points refuse a campaign start, enqueue_evolution_task_if_needed() independently pauses and disables a carried campaign before queueing it, and assignment drops what still slipped through (supervisor/evolution_lifecycle.py). Generic supervisor.queue.enqueue_task() has no runtime-mode predicate at all. Configured worker count is therefore not available capacity: the truthful value is the currently assignable idle count after custody, reaping, and admission fences.

A headless task is ADDRESSED when it is admitted, not when it is displayed (log_addressing.ingress_chat_id). A registered project's run has exactly ONE destination: an explicit chat_id may only agree with that thread, and any other value — the hidden partition included — is refused with a typed 400 rather than honoured or silently overridden, because a run addressed away from its room is the one shape that puts a card in Main whose project holds none of its work. Without a Project, ordinary API tasks default to HIDDEN_CHAT_ID (0). The confirmed browser Publish flow explicitly carries source="web" and WEB_UI_CHAT_ID to request Main; source is caller-declared addressing on the existing owner API, not a new authentication proof. Other non-Project conversation addresses remain refused. A run scoped to a REGISTERED, active project is admitted into that project's thread (dialogue, children, attachments and answer in the room the owner already has; Main still receives the one host-stamped completion row), and Main is told it finished only when its work is actually in that room — addressed there at admission or BOUND to the project; Registration alone does not qualify. Every other run stays in the hidden partition, silent in every chat, read back through the terminal, --result-json-out, the chat-blind Logs panel and GET /api/tasks/<id>; a reserved but inactive project keeps its chat acceptable so the queue's lifecycle fence refuses with its own typed reason, and a project deleted mid-run keeps its reserved chat. The run is also NAMED at admission and chat promotion, without a new model call: a caller-supplied title (ouroboros run --title or the top-level contract field; metadata.title is refused with a 400 like metadata.project_id) is authorship and fills both title and suggested_name; otherwise the request's first line, stripped of markdown and capped at the project-name length, fills suggested_name ALONE, so a truncated prompt never outranks a real name coined later, and a task_named frame is broadcast on admission so the live card is never born showing its status phrase as a title — the client buffers a task_named that arrives before the card's record exists (web/modules/chat.js), so frame order does not matter.

queue_snapshot.json is an atomic recovery and diagnostic projection, not a second scheduler. It carries pending and running rows, acceptance and root-budget fences, resident/active/parked worker counts, assignable capacity, and any pool-disabled reason. Startup restores a recent snapshot into an otherwise-empty pending queue and never resurrects ordinary RUNNING work: instead it FENCES every surviving RUNNING row with a durable cancel intent (reason='server_shutdown') and lets the ordinary cancellation custody path terminalize it, expire its open quiz and close the paired owner wait, so a window closed on live work ends honestly instead of leaving a ghost; the restore ledger row names them as terminalized_running and the boot notice states the intent, because custody writes each terminal result a watchdog window later. A selected native owner-wait handoff remains eligible beyond snapshot age only through its current waiting source and acknowledged planned-restart transaction. Terminal tasks stay terminal, a task with an active durable cancel intent (or a legacy cancel-requested latch file) is left for cancellation custody, descendants below an accepted or sealed root finalize as cancelled, and malformed durable fence evidence fails closed. Snapshot capture copies the live containers under the queue lock because concurrent HTTP mutation can otherwise crash the supervisor mid-iteration.

Pooled completion separates a finished file-save attempt from publishable terminal truth. worker_process.worker_main calls headless.prepare_terminal_task_files before its own non-ephemeral buffered task_done, after blocking post-task work; earlier answer/metrics frames retain their order. The worker stays occupied until that first attempt ends. Its private integer _files_prepared_attempt identifies the attempt, not save success, and never becomes a durable result or public event field. events_task_done re-reads CURRENT and uses headless.terminal_task_files_ready: a split drive requires the existing child-bound copyback projection and body, not merely an early terminal post-task checkpoint; pending refs may remain, but workspace artifact finalization cannot still be pending. Queue removal, slot release, accounting and project/evolution hooks stay with the normal event owner.

Legacy or faulted completions use enqueue_terminal_file_recovery: worker_health owns file-job preparation/recovery, while task_reaper owns execution queues and deferred-job replay. Existing lazy facades preserve caller names; this is one D08 custody flow. Missing/unreadable CURRENT publication retains the same RUNNING ownership and retries on the health cadence; it does not emit false Done or replay model work. The helper's transient terminal_source_present distinguishes confirmed terminal input, confirmed absence, and an unknown read. Confirmed absence reaches the existing lifecycle-fault owner: an early sticky completed status is preserved while execution becomes infra_failed; authored answer, review/objective and cost survive. A CURRENT publication/cancellation that won the race is not faulted. Pooled mailbox cleanup follows file preparation through settled_mailbox_cleanup_allowed; accepted pending attachment refs and open post-task work also protect startup cleanup. Direct/ephemeral paths keep their existing ownership.

A required owner wait keeps its task RUNNING and retains the same worker, command queue, live browser and services. The completed-tool checkpoint and task-result owner_wait projection reach durable storage before the snapshot confirms the park. The worker lends only its active_capacity; the ordinary assignment tick replenishes the configured active capacity through the existing spawn/readiness owner. Addressed input requests a wake, which waits for active capacity and retires an idle replacement or atomically transfers a confirmed-dead exhausted replacement's reservation to the original worker. A failed grant restores both capacity marks; temporary boot/reap reservations still count against the cap, so maintenance cannot buy an extra replacement. Resumption preserves the task attempt and start timestamp and marks its continuation authority consumed before dispatch; no completion, new admission or replay of completed tools occurs. Waiting spares only the idle rail. Stop, deadline and absolute ceiling keep their authority, and cancellation, timeout or crash retires an inactive worker or replaces it when active capacity is missing, without exceeding the configured active limit. A process crash or automatic timeout retry preserves the current-attempt checkpoint as evidence and does not blindly replay the task; cold continuation is limited to confirmed planned restart, which cannot preserve an OS browser session. Every shutdown cleanup recognizes the same restart transaction. An owner-requested MANAGED UPDATE is such a planned restart: its writer fence prepares the owner-wait/planned-restart handoff before stopping the pool and passes those exact task ids as preserve_running_task_ids, so a parked wait is requeued instead of interrupted, and a preparation failure blocks the update (repo-writer admission re-opens) rather than terminalizing the wait. At the common direct re-exec seam, a valid managed update in the same restartable phases allowed by _safe_restart_serialized arms the already-prepared transaction through the existing one-shot environment handoff. This includes assisted updates whose waiting tasks are already PENDING; the same-PID successor acknowledges it (launcher mode observes exit code 42). No-resume flags suppress this token, and an aborted update's leftover recovery record alone cannot authorize a later manual Restart or rollback. A failed restart callback retains the existing disarm behavior. A manual ROLLBACK returns the tree to an older runtime, so it deliberately parks nothing and keeps the ordinary interrupt semantics. Cold preparation restores the original CostCeiling before constructing either context projection and retains hard clocks, refreshes genuine assignment progress, and rebinds ContextFit to the saved model. Cold continuation completes the saved round's pending budget decision before advancing to another model round or applying queued route overrides. TaskModelWait retains its per-role choices (including Auto), auto-continue choices and completed quota union; both native execution and supervisor timing retain the original execution basis and clock revision. Calendar deadlines and time spent waiting for the owner are unchanged. Successful grant consumes the queue task's resume locator; source evidence remains available through a later crash or timeout.

Ordinary Main/Project roots retain their in-process actor through the same completed-tool owner-wait boundary. owner_wait.direct_owner_wait uses that actor's existing mailbox and TaskModelWait controls; no pooled slot is held, lent or synthesized. Pending owner events are flushed before sleeping, and the loop handles controls/deadlines before its post-tool budget decision after wake. Its saved source is evidence, not an automatic cold-restart grant for an ordinary conversation. An addressed text or existing quiz answer resumes the same live stack and browser.

Cancellation is intent-then-custody; intent and outcome are separate fields, because one field carrying both wedges a task forever. The skeleton:

  1. Intent. Every cancel ingress — the agent cancel_task tool, the HTTP single and cascade endpoints, evolution stop, project deletion, the per-descendant mints of a cascade sweep, and the boot migration of legacy cancel_requested files — writes one durable row through ouroboros/cancel_intents.request_cancel into the locked projection state/cancel_intents.json (active intents only; every transition also appends a forensic cancel_intent supervisor-ledger row). Every ingress fails closed: a failed intent write refuses the cancel with a typed error (tool CANCEL_INTENT_WRITE_FAILED, HTTP 503; a corrupt projection file gets its own refusal, CANCEL_INTENT_PROJECTION_CORRUPT, naming the preserved file), so no teardown ever runs without a durable, watchdog-replayable fence. Evolution stop keeps any task whose intent write failed and reports the stop INCOMPLETE with typed per-task outcomes.
  2. Scope. A cascade ingress mints its intent with scope: cascade; recorded scope is widen-only (single→cascade, never narrowed, so Stop-now cannot shrink a cascade). A cascade over an already-settled root with live descendants still mints the durable coordination intent (allow_settled_target) — that intent is the watchdog's replay trigger for the subtree — and the sweep mints a per-descendant intent so a crash leaves no live descendant unfenced. Timeout reaping is deliberately NOT a cancel ingress: the reaper keeps its own custody protocol over the reaping slot marker and never mints intents, because cancellation is reserved for explicit intent.
  3. Claim. supervisor.task_lifecycle.cancel_task_custody is the ONE settle owner. It claims the intent before any custody mutation (owner + generation, exclusive while alive); a refused claim exits having touched nothing, so two racing custodies can never interleave into a double settle. The one secondary settle site — the pre-assignment pending drop — holds the same claim/generation fence before settling; a queued task whose budget is exhausted now PAUSES instead of terminalizing, so there is no batch terminalizer beside custody. Stopping a LIVE direct-chat turn is addressable through the same custody: there is no worker process to kill, so the chat lane writes the typed finalize_now control into the canonical drive's owner mailbox — the one the turn's loop drains at every round boundary, where it ends with zero further model calls (the paid post-task synthesis included: the hard stop records the existing _skip_post_task_synthesis marker on the tool context and emit_task_results carries it onto the task record, so no summary, reflection or consolidation bills after the stop and no open checkpoint is left for the boot reconciler; a Stop-now that lands AFTER the loop returned, while that paid synthesis is already running on its in-process thread, is still addressable — the pipeline's in-flight key counts as live ownership for task_has_live_ownership, custody keeps the durable immediate intent open (typed "still live") instead of settling it, and the synthesis worker consults that intent (or the live marker) before EACH paid stage, skipping the rest and disclosing them as post_task_stop_reason owner_stopped:skipped=<stages> on a degraded checkpoint that rides the result row and task_cost_finalized) — arms it exactly once per turn (the turn's own stamp is the latch, atomic against the turn's completion so a turn that already ended gets no control and no owner toast), then waits the short OUROBOROS_DIRECT_TURN_STOP_WAIT_SEC bound. The typed outcome is gone / ended / live; live releases the claim so the supervisor sweep retries rather than publishing a fabricated cancelled row over a turn that is still running.
  4. Kill and re-check. Custody confirms process death, then re-reads the child's real settled result. Natural completion WINS: a child that finished before the kill keeps its completed result, artifacts, and cost, and the cancel settles as already-settled.
  5. Reconcile and capture. The task's open delegated runs are reconciled from durable custody rows and always re-audited (open runs and pending invocations are disclosed regardless of the reconcile outcome's shape). Workspace artifacts are captured from the real tree; a failed or owed-but-unrunnable capture is failed, never missing, and a shared-tree capture carries attribution: shared_unproven.
  6. Settle. The settled result is written with reconstructed-or-honestly-unknown cost, never a fabricated final $0. parent_decision is stamped only at this outcome.
  7. Owe, then publish. The owner's terminal answer is registered as OWED in the durable outbox (or a typed no-chat handoff row) BEFORE the intent settles and before task_done publishes, so a crash between settle and send replays the answer instead of losing both it and the watchdog trigger. A cascade delivers one root message with a children digest under the deterministic delivery id cascade:<root_tid>:<request_id> (replays dedup; a later separate cancel delivers its own); digest membership merges the root's durable descendants with the sweep's outcomes, each child's line rebuilt from its current durable status.
  8. Watchdog. The supervisor tick dispatches the existing cancel/delivery/ref sweep off drain under one nonblocking in-flight lock (server_maintenance._run_cancel_delivery_ref_sweep, unchanged ~20 s cadence). Its sweep_cancel_intents re-feeds unclaimed or abandoned-claim intents into custody — replaying a cascade as a cascade, not as a single cancel — so a lost control event or a custody attempt that died mid-teardown cannot wedge a cancellation. The cascade postcondition is judged on physical queue/durable liveness (the intent itself excluded); only that no-live check settles the coordination intent, after the tree's summary is registered as owed.

Readers see the typed projection cancel_state: "pending" (with cancel_reason beside it) on effective results until the settle; the UI shows an interim "Cancelling…" and restores the Cancel button only when a fetched live non-pending task detail proves the intent is gone. Steering writes — steer_task, mailbox follow-ups, forward_to_worker — are refused typed while a cancel is pending; the check runs before attachment staging and a refusal removes just-staged inputs. Queue restore and pre-assignment consult the projection under the queue lock, so a cancelled pending task never starts. task_done is validated through the DURABLE result unconditionally for every non-ephemeral event: a non-settled event status, a settled claim over a non-settled durable row, or a blank status over a non-settled/absent row is refused as a durable lifecycle fault — left to custody when a cancellation is pending, otherwise published through the existing lifecycle-fault owner with a typed infrastructure-failure axis, preserving an existing sticky terminal status; unreadable publication keeps an owned file retry instead of releasing a false terminal; that synthetic terminal rides the normal dispatch seam, including the assisted-update orphan watchdog and the cooperative-checkpoint hooks. Terminal answers ride one durable delivery seam (supervisor/terminal_delivery.py): restart-surviving delivery_id dedupe shared with the natural final-answer path, a loud UNREVIEWED salvage message (bounded preview, exact omitted count, full-copy receipt) for cancelled and non-retry-reaped tasks, one root message with a children digest for a cascade, nothing for a retryable reap; routing follows the task's lineage chat. The already-settled fast path and the finalize-on-miss lane run the same delegated-run audit as the kill path, so a cancel over a dead task with live delegated runs never reads as a clean completion. An agent-requested cancel (cancel_task tool → cancel_task control event) publishes nothing of its own for the cancelled, already-settled and not-found outcomes — the custody seam's terminal rows in the task's own thread and the typed tool result are the truth, and a host acknowledgement in the global owner chat was duplicate, untyped speech — while a FAILED settle still speaks as the typed cancellation_fault progress incident on the requested task. _handle_cancel_task dispatches _drive_cancel_task_event off drain with a per-root/task in-flight key; the existing durable claim/generation remains custody authority, and failure releases the local key for watchdog recovery. HTTP cancellation keeps its existing worker dispatch and response contract; unrelated Stop work does not queue behind a new shared file executor. The custody, completion-wins, and owed-before-published invariants are restated in §10.

Stop POLICY is an axis on the same durable intent, independent of cascade scope. An omitted/empty-body cancellation is the synchronous IMMEDIATE teardown, keeping programmatic callers' bounded budgets. An explicit stop_policy=finalize_then_cancel answers 202 with the intent OPEN and runs one bounded owner-stop finalization episode (supervisor/owner_stop.py): live descendants settle first and feed a bounded child-result projection into the root's final turn; the root receives a deterministic finalize_now control whose typed first line (owner_requested_finalization) routes to its own loop rail — zero or one tool-less model turn, retained-candidate reuse, terminalizing completed/best-effort under the honest owner reason rather than a false deadline reason. Two anchors are immutable: the grace budget starts at the durable control_drained_at (first drain wins, so a task inside a long tool call still gets its final turn when the hard bounds allow), under the request's outer OWNER_STOP_OUTER_CAP_SEC cap. Held tasks bypass only generic idle/finalization-grace rails; the task's explicit deadline and absolute ceiling remain independent hard axes and are never widened. Expiry, a hard-bound hit, a pending root, or an already-settled root feeds ordinary custody. Policy transitions are monotonic: an immediate request HARDENS a pending graceful intent (preserving any cascade scope, revoking an unread finalization control, revalidated at drain); graceful can never soften an accepted immediate. A successful graceful root suppresses the redundant cascade summary; Panic bypasses both. The UI projects the pending soft stop through cancel_state+stop_policy ("Finalizing…"). Beside stopping sits the owner "hurry" control: a typed task-local kind=hurry owner-mailbox control (ouroboros/owner_hurry.py, gateway/task_hurry.py) that skips the next otherwise-eligible acceptance panel with a typed reason, zeroes remaining improvement passes, and makes force-plan projection task-locally advisory — never a chat message, never a settings mutation, never a P3/commit/review-gate weakening. Its effect is attempt-scoped (task["_attempt"]); a shared retry_reset strips it on every same-id requeue, reaper timeout and crash requeue alike. These invariants hold for every install configuration class.

The event bus is process-lifetime rather than worker-generation-lifetime: full-pool and single-slot respawns reuse one manager-backed queue shared by workers, direct chat, and consciousness. A force-killed producer can corrupt a raw multiprocessing feeder frame, while rebuilding the queue on pool rotation strands surviving producers on the old endpoint; synchronous manager serialization and isolated producer connections avoid both. Live-frame publication of persisted rows is exactly-once and process-symmetric: ouroboros/utils.py::append_jsonl streams only runtime logs/*.jsonl rows into the process log sink (never chat.jsonl, which has its own live channel, and never state/memory/receipt stores), and each process suppresses the types whose live delivery has a dedicated owner — workers via WORKER_LOG_SINK_SUPPRESSED_TYPES, the server via the superset SERVER_LOG_SINK_SUPPRESSED_TYPES in make_server_log_sink. One persisted event produces exactly one live frame, pinned by tests/test_log_forwarding.py; an LLM call failure is one durable llm_api_error row and nothing else (the live-only llm_round_error sibling the same producer once emitted was the same failure twice — Background Consciousness keeps its own live llm_round_error).

Heartbeat and progress are different evidence: a heartbeat proves a process or loop is alive; owner-visible progress and model-usage events prove the task advanced. Fresh descendant progress or queued descendants can keep an orchestrator alive, while an explicit deadline, absolute ceiling, cancellation, and budget stop remain hard. After the typed finalization episode (§6), timeout handling freezes its decision under the queue lock, removes the task from ordinary assignment, marks the worker reaping, and hands kill, join, salvage, retry, and respawn to the single off-loop reaper. An orchestrator with live descendants is not blindly retried, because a retry would replay its plan and spawn a competing tree.

No retry or new assignment may occupy a timed-out slot until the original process is provably dead. If kill and join cannot establish death, the reaper preserves a low-rank RUNNING result, keeps the slot marked reaping, emits a visible wedged receipt and restart hint, and performs no terminal write, task_done, retry, or respawn. One slot is sacrificed rather than letting a still-running process race a replacement and overwrite its result; the next supervisor generation reconciles the durable record after old-generation process custody has run.

A spawned or respawned slot is not assignable until its child's PID-bound worker_ready row arrives; supervisor/worker_pool_lifecycle.py keeps both spawn paths under the existing readiness watcher and lifecycle→queue lock order. Temporary reaping reserves active capacity even when not assignable. After WORKER_READY_MAX_ATTEMPTS windows of WORKER_READY_WINDOW_SEC (runtime_limits.py, re-exported by config.py), Worker.readiness_exhausted records the final outcome for that exact slot; late ready/watcher events cannot reopen it or alter a replacement generation. worker_pool_lifecycle owns _worker_pool_execution_state and exhaustion disablement behind the existing workers facade; it distinguishes busy/booting/reaping capacity and a valid live owner-wait stack from total exhaustion. Reserve, final enqueue, snapshot and public admission share that fact; repository-writer admission remains a separate outer policy. Total exhaustion closes pooled ingress, including owner /review, without blocking direct chat/control or internal boot/update recovery. Once existing RUNNING completion custody is settled, disable_exhausted_worker_pool uses the ordinary pool stop to fail unstarted PENDING work honestly with a Restart hint; failed publication retains existing terminalization-retry custody. A new incoming task cannot clear the latch or restart the exhausted attempt budget. Readiness remains separate from liveness and task idle time. Missing events are an empty read; watcher errors release only still-booting, non-exhausted slots to the crash detector with worker_ready_released. Linux workers use forkserver; macOS and Windows use spawn.

Unexpected worker death reserves exact worker/meta/task/attempt/root custody under the queue lock and enqueues confirmed_dead_worker on the existing reaper. worker_health.recover_confirmed_dead_worker performs archive/file work outside that lock and checks strict CURRENT readiness after the shared helper. A saved terminal source wins even after signal death; unknown or incomplete file publication retains the same job via TerminalFileRecoveryPending and health-driven deferred replay. Only confirmed absence of a usable terminal source reaches the existing signal/non-signal crash policy: a signal is an infrastructure failure; an otherwise eligible non-signal crash retries within QUEUE_MAX_RETRIES, preserving owner-wait replay restrictions and cost. Generation checks and atomic retry rollback prevent stale jobs from affecting replacements. Crash-storm jobs follow discovered dead-worker jobs, with respawn suppressed while their terminal sources settle; the existing storm fence then stops pooled admission. Direct chat remains independently available.

Startup terminal-file recovery runs in _run_supervisor after process custody and before _startup_prune_sweeps; without providers the lifespan uses the same recovery without spawning a worker. Prior worker/server PIDs are captured before spawn replaces the old PID record. Unknown or still-live ownership defers recovery rather than racing a writer. The existing recovery owner scans only known headless/task-drive roots, protects live/pending work and acknowledged owner-wait continuations, and invokes the shared file helper before orphan-result/post-task reconciliation. Its transient report carries recovered, unresolved, protected and error facts. Any unresolved/protected source or ownership/read error sets preserve_task_sources for that startup pass, skipping task-drive/tree/temp-source deletion while unrelated maintenance retains its existing owners. No initial adoption anchor or new durable recovery store is needed.

Startup and throttled maintenance reconcile three residue classes. Process custody checks strict PID, start-time, command fingerprint, owner task, session, and generation evidence before reaping an owned process. Delegated-run reconciliation applies the same owner-gone reasoning to external harness rows; the write-side healing seams (boot-backfill reverse join, cursor pass, sweep refresh) and the counters-as-snapshot rule are specified once under §6 Delegated subagents. Task, review, and project reconciliation repair durable records whose producer no longer exists. None of these are command-line-class kill sweeps, and one development or runtime instance never reaps another. The dedicated watchdog separately observes supervisor-loop liveness and every registered native actor; it alerts and recommends /restart but cannot kill an individual in-process thread. Other owner conversations run on independent native actors, without a second scheduler.

Cooperative project checkpointing has two equivalent quiescence triggers: a host-minted genesis or cooperative tree is checked when its root settles with no live descendants, and again when the last child settles beneath an already-terminal root. The second trigger exists because a root-scope budget stop terminalizes the root before its children reach their own dispatch boundaries; a root-only trigger would see a live tree once and never return. Event dispatch detects the condition only after removing the finishing task from RUNNING. The bounded git chain runs on a daemon thread, revalidates quiescence under the queue lock immediately before mutation, and uses a per-root latch that replays a trigger arriving during an in-flight check. Only host-minted project roots are eligible; owner-attached folders are never auto-committed, credential-shaped files stay excluded and disclosed, and every material success, skip, or error receives a durable receipt.

The bridge recognizes /panic, /restart, /review, /evolve [on|off], /bg [start|stop|status], and /status; all other text enters ordinary agent routing. External transports may invoke these commands only with positive owner identity and a transport-specific owner-chat binding. The commands reuse runtime-mode, queue, cancellation, and typed-result authority rather than implementing parallel control paths. Runtime logs rotate on the same supervisor tick and archive readers preserve their retained timelines. Only explicitly isolated devtool roots may use the narrow rotation sentinel from §1; normal runtime roots never inherit it.

6. Agent Core

Task lifecycle

model_execution is a compact projection of the last ordinary solve response the loop accepted (text or tool calls), selected from marked llm_call_refs. It preserves the initial requested model/local route separately from the used route and an independently reported provider label. Empty responses, forced finalization, review and post-task calls cannot replace that observation. The existing owner-wait usage checkpoint preserves both facts; result, terminal frame, history and card share the projection without changing dispatch or claiming final-answer authorship. The canonical terminal task_summary retains the same projection when the full task result ages out; a current result or live card observation remains authoritative over that historical fallback.

A queued user task enters through a reviewed transport, is admitted by the supervisor queue, and runs in OuroborosAgent; each direct-chat turn runs on its own in-process agent, while an explicit Swarm routing turn retains the ephemeral contract; both are tracked by the process-local DirectActivityRegistry and create no PENDING/RUNNING queue record (§3). The root pipeline captures the task contract and immutable context core, executes the LLM/tool loop, preserves a delivery candidate, stores the result and artifacts, emits lifecycle and usage evidence, performs the root-only post-task work, and publishes the typed outcome. Queue admission proves only that asynchronous work was durably accepted; completion, objective satisfaction, artifact finality, verification, and review acceptance remain separate facts.

DeliveryCandidate is retained before verification or review so a later notice, reviewer failure, deadline, or provider outage cannot erase a useful answer. outcomes.py combines execution, objective, review, artifact, and child-absorption axes without converting one axis into another — the terminal custody overlay (outcomes.custody_debt_axes) is an instance of that rule, not an exception to it; verify-before-done receipts and exact artifact references are host-attested evidence — declarations and answer prose are not substitutes. A forced exit may publish the best current candidate only with its typed rail and evidence-freshness disclosure, and lifecycle may remain completed while the objective or review axis records a best-effort or unaccepted result.

Host plan/orphan disclosures live beside the model answer as terminal_host_notice, through the existing result record and terminal outbox. Delivery splits them into an assistant answer and a System notice, preserving each delivery's replay/dedupe; CLI and single-body Presence retain an explicitly host-labelled status without changing stored answer bytes. The same current answer keeps its hash and real PASS or FAIL when only this notice changes. _current_delivery_candidate still checks owner generation, current panel/binding and evidence: equal text cannot revive a superseded verdict. Synthesis receives the notice separately too. Automatic parent handoff and ordinary get_task_result/wait_task also retain that separate host status. Their whole-result-or-pointer budget includes it. The compact wait_tasks projection retains the same semantic notice separately from result; accepting its current child hash cannot hide an unseen limitation. The existing child-result hash includes the new notice when present, so a changed child limitation invalidates an old parent disposition; legacy rows without the field and the model answer's own hash stay unchanged. The task:<id> plan-evidence reader carries the notice as a separate field too; the existing evidence resolver hashes it before attachment budgeting/redaction, so a changed source limitation cannot replay a review of the old evidence.

The delivery-control protocol is resolved here and only here (DEVELOPMENT keeps the rule and points here). The candidate carries sticky loop-local provenance that its lineage has seen a host-issued delivery-control episode (_arm_delivery_control); with no such episode, exact JSON is ordinary text. In a marked lineage, both the ordinary and the forced resolver intercept recognizable whole-body envelopes and balanced trailing protocol attempts — valid keep resolves to the retained candidate, valid replace to full_answer, and anything malformed preserves the retained candidate. Both resolvers strip one whole-body fence and treat a balanced protocol object at the very END of prose as a protocol attempt (utils.extract_trailing_json_object + loop_delivery._parse_delivery_control_body) — the trailing-object rule deliberately refuses substring scanning, because quoted protocol literals mid-prose are legitimate text. The ordinary resolver takes one repair round then degraded-preserve; the forced resolver resolves purely and never re-loops — malformation preserves the retained candidate with the typed delivery_control_degraded reason, which is this forced rail's own code, while the ordinary repair path records invalid_delivery_control_after_repair. Every degradation carries the cause it computed: outcomes.derive_loop_outcome falls back to delivery_control_degraded only for a degradation that reports no cause and publishes the loop_outcome.degraded/degraded_reason pair the benchmark ledgers read. A balanced trailing object remains a malformed protocol attempt after the transient latch clears, provided the retained lineage saw the host control episode; it preserves the candidate and never leaks JSON. A control object quoted mid-prose stays prose, and without a host control episode protocol-shaped JSON remains ordinary text. On the forced rail a truncated trailing fragment remains prose under the existing parser contract.

The child-absorption gate is an action gate: while undispositioned direct children remain, the loop HOLDS the candidate (child_absorption_or_revision_required) instead of arming the JSON-only control instruction — the hold-vs-arm split exists so the model never receives two contradictory instructions in one round. A typed keep cannot close the gate, and after the one bounded reminder the gate forces the best-effort children_unabsorbed rail with a current id [status] sha256 listing. The absorption digest's ## child header carries the child's typed custody debt (delegated_runs_unreconciled, the first ten ids plus an omitted count and a get_task_result pointer) — visibility only: the parent's authority over that patch is exactly the orphan rule, and a child's debt never relabels the root card (DESIGN §4, a failed child keeps a compact factual marker inside its parent while the root continues under its own authoritative status). The acceptance-subtree copy of that debt is a disclosed gap, deferred with the pre-finalization reminder. Finalizing over an UNDISPOSED OWN delegated patch is deliberately NOT gated: the consequence is disclosed where the decision is made — in the integrate_delegated_patch schema and in the apply receipts — and lands as the additive Done-with-warnings custody overlay rather than a hold. A pre-finalization reminder for that case is deferred.

Provider death is the one forced rail that is NOT a best-effort completion: _handle_provider_unavailable salvages the best available text but stamps infra_failed, so the task terminalizes failed with the typed provider_unavailable reason and an immediate "provider outage — NOT completed" owner notification; a waited-out transport outage reaches the same rail through its deterministic no-resend branch (transport_unavailable_no_resend, or provider_outcome_unknown_no_resend when the round still holds a transport-death repeat record). The rail makes its one forced model call only while a call can still land: a transport that spent its same-model retry wall stamps _llm_retry_wall_exhausted, which the rail reads as a no-call gate beside context_overflow and provider_outcome_unknown (_provider_unavailable_result alone decides the terminal precedence, in this order: an unresolved round record, then the latched wait cause, then the overflow salvage, then the deadline, then the retry wall (an exhausted deadline suppresses the wall source)), shipping the salvage with no request — so the forced rail does not re-pay a second retry window over a proven-dead provider; the marker is a last-invocation bool on the shared usage dict, a disclosed residual. Terminal delivery preserves producer authorship. The provider rail records terminal_provider_notice as host presentation beside the raw result. Managed/direct receipts and secondary System incidents carry the known wait duration, finalization cause and unknown-outcome warning; secondary notices do not recommend a blind rerun. Ephemeral replies and message/deferred Presence bodies carry a separate host-labelled status section, retaining raw source independently and making no task-details promise for a transient turn without a task record. Cached Presence metadata replays that completed projection without adding another notice; silent/tool-delivered outcomes retain their delivery authority. Every forced rail stamps one closed-vocabulary producer word at the single forced-finalization sink: model_final for complete model text (host-authored plan/deferred-child notices stay separate, as on the ordinary no-tool final), host_notice for a terminal text the host wrote alone (kept verbatim on every transport, with its markdown, and projected as a System row without a system_type), and host_salvage only on the provider-death rail, where managed/direct delivery replaces the text with the enriched outage receipt and the full bytes stay in task details; the provider-death wrapper upgrades a bare notice to salvage so a no-call/no-resend arm keeps its receipt. A missing origin identifies only rows written before that stamp existed.

Host-enforced task acceptance is a root-owned completion coach, not the P3 commit gate. off disables it. Both auto and required review observable effects and typed deliverables/criteria. In auto, an explicit task_acceptance_review request also qualifies, including read-only research; queue membership alone does not. required retains its non-direct-root criterion. Ordinary conversation, exploration or cognitive-memory updates alone do not qualify in auto; no prose or tool-count classifier decides their meaning. Ephemeral control turns are excluded, and child reviews remain advisory evidence superseded by the root decision.

Before an eligible panel is called, supervisor/task_lifecycle.py closes subtask admission under the queue lock and task_status.find_child_tasks proves the recursive subtree terminal and quiescent; revision reopens the fence, terminal or degraded completion seals it. Fence acknowledgement, subtree lookup, the timing telemetry stream, and the packet's mutation-attribution read use the canonical budget_drive_root; the one-shot state/acceptance_fence_acks/ IPC sidecar is not a lifecycle authority. The reviewer packet preserves verbatim owner directives, the full task contract and criteria, canonical deliverable identity, terminal child state, verification receipts, artifact references, the host-attested lifecycle facts (review status, staleness, readiness, enablement) of every skill the task touched — visibility only, never a gate — and an explicit omissions manifest; a required component that cannot be assembled makes the affected actor DEGRADED, never a silently smaller prompt. The packet is SIZED against the review quorum's real windows — the same reviewer_window/review_synthesis.quorum_input_token_limit seam the triad and plan review use — resolved once per task and memoised on the acceptance context, so the packet bytes cannot drift between the binding build and the staleness rebuild. The same cached per-slot caps drive the pre-send fit check; dispatch never recalibrates them. Non-core sections then shed through a DISCLOSED ladder: the predecessor authority envelope first, then the trajectory tail and its results, artifact previews, agent-supplied evidence, and last a diff preview that keeps the durable repo_diff_source_ref. Each shed is a row in omissions_manifest. A slot whose own window cannot hold the rendered prompt is a typed $0 not_dispatched row while the rest of the panel reviews; a packet that still overflows after every shed stamps __immutable_core_overflow__ naming the oversized sections and refuses the panel without spending anything. __unresolved_partial_artifacts__ withholds the panel's packet rows only for a tool result whose exact source is genuinely source_unavailable (a retrieving row reads the exact source itself) — a budget shed with a durable, actor-resolvable source ref is an omission, never an unresolved partial.

The packet is delivery-conditional (owner R4/R15, 2026-09-01); the FULL packet is not. Every triad row reaches the panel as configured (reviewer_slot_config.triad_delivery_slots, R2). A packet (api_chat) row receives the assembled packet as before. A retrieving row receives the route-owned work order that loop_acceptance_review.acceptance_retrieving_work_order writes onto ReviewRequest.slot_session_tasks: the same task-stable contract and the same output contract (review_execution.review_output_contract, handed over as policy["output_contract"] so the retrieving executors never fall back to the generic object form; route-owned policy keys are not rendered into the api pack's Policy JSON), absolute retrieval pointers (the task's ACTIVE workspace — review_repo_dirs_for's subject root, never the governance repo — as session_root, the task result record, the artifact directory, the verification-receipt file, the tool-trajectory log), and the packet in the form its delivery can use: an agent-session row gets the FULL packet — its run is unobserved by the host, so the packet is its only attested view — with the disclosure that access outside the workspace is not guaranteed and a refused read is absence of evidence, not absence of the artifact; a native inspection row gets the packet WITHOUT its freely degradable tail (tool-trajectory rows and artifact previews, manifested as retrieving_delivery omissions) plus the real data root (policy["native_data_root"], R5), because its episode reads those sources itself under host_file_read_attestation: host_observed. Reviewer evidence_refs resolve against the FULL packet on every delivery — never against the rendered projection — so a retrieving row citing a real receipt is clean exactly like a packet row. The immutable-core overflow refuses every delivery; a partial tool-result projection refuses only packet rows (a retrieving row reads the exact source). The owner deadline rides the request (R23). The gates are route-aware: the wave budget gate prices API money only and DECIDES on one work-order send per paid row (a packet row by its real message pair, a native row by one send of its work order; a session row rides the owner's subscription and is not priced); that admission is the whole money rule — there is no rounds multiplier and no second, read-only pricing pass (owner R52), so a panel's cost is bounded at dispatch by the per-send wallet binding rather than predicted before it. only a floor that does not fit is refused review_wave_budget_insufficient, and the packet rows' second physical send for format repair never buys a retrieving row a second episode or session — its executor canonicalizes its own answer. Child-task and off-mode acceptance stay advisory and run packet rows only.

A paid panel is the scarce resource, so it is priced by what the agent actually CHANGED: the host buys one panel per PAID IDENTITY — sha256(candidate_hash + the sorted set of nonempty (obligation_id, disposition, sha256(reason)) tuples) — so exactly two things mint a new panel: a changed candidate answer, or a new nonempty obligation disposition. An empty disposition reason hashes to nothing and buys nothing — an empty rebuttal is not an argument. The evidence revision deliberately does NOT price: every cosmetic tool call shifts it; it stays stale-packet detection. A resubmit with an unchanged paid identity replays the recorded verdict for free and terminalizes with the typed identical_acceptance_refused reason; a replayed CLEAN pass terminalizes accepted on its own branch.

The configured slots are independent actors with adaptive quorum (config.adaptive_quorum: 2-of-N for N≥3, both for N=2, single reviewer as loud single_reviewer_no_diversity; a fewer-responded shortfall stays a loud infra quorum failure). Each actor receives one substantive interaction on its bound route: at most two physical sends for a packet row, one bounded episode for a native row, one delegated session for a session row (owner R0/R1, 2026-09-01: all three deliveries follow the triad rows and a retrieving verdict is equally authoritative). Transport status, parse status, semantic verdict, criterion support, route, quorum contribution, and binding hashes remain distinct, so an unavailable or malformed response cannot masquerade as a negative judgment. An acceptance panel that refuses before any transport projects not_dispatched on every row and on the panel; not_dispatched is a transport state distinct from success, timeout and provider_transport_error, and it is never a verdict — the refusal cause rides the row’s error text into the degraded reasons and the owner line. PASS/FAIL/DEGRADED are reviewer verdicts; the host-owned completion decision is separately accepted, revision_requested, or finalized_unaccepted, written only by loop_acceptance._set_acceptance_decision. Only a clean quorum may authorize accepted; the agent may add its own disposition but cannot overwrite the host decision. A task-acceptance FAIL contributes only with the required outcome tier and a bounded correction rail; transport/unparseable no-quorum records terminally as finalized_unaccepted with reason=review_degraded — never PASS, never revision authority.

A clean criterion is evidence-resolved, not merely well argued: reviewer evidence_refs must be exact members of the host packet's enumerable reference vocabulary, and a claim id resolves only through acceptance_support_refs linked to a passing host receipt for that claim. Agent-supplied, declared-intent, unattested, and non-resolving sections never certify success; an OPEN plan wave binds nothing: its reviewed claims are disclosed as acceptance_claims_source='none_open_plan_wave' beside a non-binding plan_claims_exhibit that sits in DECLARED_INTENT_SECTIONS, so citing it never resolves and it is no longer indistinguishable from a task that never had claims. An unresolved reference keeps the actor's record for audit but removes its clean contribution (criteria_refs_unresolved). This total, fail-closed resolver is why the task cannot certify itself by echoing its expected outcome.

Actionable findings enter the durable obligation dialogue with stable identity: fix, rebut with an evidence-bearing disposition, or ask the reviewers to declare the issue unreachable or a stable disagreement. Re-raises must name an existing obligation id or are disclosed as new; a valid rebuttal retires the row, an invalid one reopens it with both positions preserved, and the reviewer's reviewer_rebuttal_response rides into the next panel's catalog so it can tell "already answered" from "never answered". Each panel receives the bounded acceptance_dialogue_history held OUTSIDE the hashed evidence material — precisely so reading the history cannot mint a fresh paid binding.

Under Blocking, termination beyond a clean pass belongs to the reviewers, not a host counter. Advisory also accepts an explicit author finish after the first feedback was delivered: loop_acceptance.merge_agent_acceptance_stance binds the new stance to that feedback and the current tool/owner-directive state; loop_acceptance_review._finish_advisory_author applies it before another panel, under the existing quiescence and owner-generation checks. A revised answer receives its own author hash while the earlier critic keeps its verdict and hash. Termination beyond a clean pass belongs to the reviewers, not a host counter. Their typed dialogue_status votes reduce over contributing actors gated by contract validity; a PASS/FAIL beside a malformed parse never votes. One contributing reviewer may hold the loop open, but only WITH MATERIAL — a continue vote without a concrete finding or completion coach is disclosed as continue_without_findings and abstains, because a bare "keep going" bought a panel every round. One well-formed terminal vote ends the dialogue; zero well-formed votes reduce to the typed inconclusive — deliberately outside the reviewer vocabulary, granting no authority in either direction. Required+Blocking continues until clean acceptance or a real deadline, budget, round, lifecycle, or configured improvement-pass rail; Required+Advisory may publish an honest non-clean result. An explicit max_improvement_passes binds every policy, otherwise the shared cap (§6 Review stack). Pacing predicts no review duration (owner R52, 2026-09-03): a task has three host-owned rails — a deadline, a paid-cycle cap and a wallet — and a panel starts iff the cycle cap has room, the wallet can buy one work-order send per paid row, and MORE than the configured floor (OUROBOROS_ACCEPTANCE_REVIEW_EST_SEC, never below 200 s) remains above the finalization reserve; an improvement pass additionally needs that same floor scaled by task_pacing._window_scale (×2 under the adaptive policy, ×1 otherwise), and a spendable window at or below the applicable line is refused review_skipped_deadline_reserve / improvement_window_inside_reserve. task_pacing owns both predicates (review_launch_allowed, improvement_pass_allowed); the launch rule is evaluated ONCE per panel, at loop admission before a panel is built (owner R55, 2026-09-03) — the paid dispatch claim (review_dispatch.task_acceptance_paid_dispatch_stamp) checks cancellation and the paid-cycle wallet only, the acceptance dialogue asks the separate post-panel improvement rule after one, and no other surface evaluates time. Once launched, a review is an ordinary operation clamped to the owner deadline and the task ceiling (R23) with the per-send money fence: a panel the deadline actually cuts is a typed DEGRADED outcome, a panel that finishes keeps its normal verdict, never a free skip. Disclosed residual: a panel whose evidence build consumed the margin after admission dispatches and may be cut by the deadline — one ADMITTED panel: admission prices one work-order send per paid row, but packet rows may use the permitted repair/retry send and native rows may run several rounds, every send remaining deadline- and wallet-fenced where pricing exists, so the total is NOT bounded to one floor wave. Panel durations are recorded as task_acceptance_review_timing telemetry (with delivery, deliveries, native_rounds and native_rows) that no gate reads back; the improvement capsule reports the actual verdict, open obligation ids, remaining rails, and concrete next moves. The structured review axis is mirrored as top-level review_status for compatibility. A delivery warning preserves the current bound task-acceptance objective and its source; execution degradation and the separate host notice remain visible alongside PASS or FAIL.

A forced turn sends the round's exact tool envelope — the same schemas and the same server-web flag — so the provider prefix stays a cache hit; "tool-less" describes the instruction and the host, which executes no call, so a reply that still asks for a tool — with or without prose beside it — is incomplete on every rail and degrades through the host fallback path instead of publishing as the model's final. Every forced deadline, budget, or round rail uses the common terminal recorder; if the task was eligible but no panel ran, the review axis records an eligible bypass with zero runs plus the rail-specific trigger. No forced rail can take another model round, so a dangling revision_requested terminalizes as finalized_unaccepted (revision_unavailable_on_forced_rail) on EVERY forced rail through one shared helper, naming the prior reason; accepted and finalized_unaccepted decisions are never overwritten, no bypass reason is stamped over a panel that ran, and the pair stays outside the blocked-terminal set so the objective remains best_effort. The forced children_unabsorbed rail additionally still runs the acceptance panel for an acceptance-eligible root with a quiescent subtree, with the undispositioned-children debt in its evidence — keeping a forced delivery distinct from both clean acceptance and no-panel-warranted. A superseded panel remains an audit row with its run_count intact, and pending_delivery_acceptance is a transient eligibility, never a terminal state. The root's post-task phases use the minimal root_phase_checkpoint: startup replays only a durable pending_once phase, while an indeterminate running phase is disclosed as degraded rather than replaying paid work. The agent-callable task_acceptance_review does not call the panel for an eligible root: it validates and stores claims and evidence, then returns deferred_to_host_acceptance and authoritative=false.

Finalization controls are typed owner-mailbox entries rather than injected owner prose. The supervisor may request one bounded tool-less answer, salvage the last persisted assistant text, and retain a full canonical copy when a preview would truncate it. A grace episode has one durable control and can be revoked atomically when the task itself resumes; descendant activity does not count as the task's own progress. A process that cannot be killed remains visibly running, and custody checks prevent another runtime instance from reaping work it does not own.

Disclosed cancel-lifecycle residuals (deliberate): a cascade over a tree with no resolvable lineage chat whose typed handoff-row append ALSO fails still settles; an empty-intent release can add liveness noise to a foreign claim's forensic trail (bounded by the generation fences); cascade postcondition timing can flake under heavy load (the watchdog re-feeds — one retry, never a lost teardown); and the cost projection of a task whose delegated runs stayed open may read cost_usd=0/cost_final=true while a run is still live — the disclosure line names the open runs.

Tool capability and execution

tool_capabilities.py is the SSOT for core, meta, parallel-safe, stateful-browser, untruncated, capped-result, and reviewed-mutative tool classes. tool_policy.py chooses the initial capability set; ToolRegistry remains the execution authority; loop_tool_execution.py owns timeouts, concurrency, live evidence, result handling, and mutative ceilings. Ordinary top-level presets share one built-in name surface: project focus changes the default target, while root policy, runtime mode, task-contract disables, credentials, resources, ephemeral rules, selected-resource bindings, and delegated-child profiles narrow independently — a tool being registered or discoverable is therefore not the same as being callable for a particular target. Lazy capability discovery returns an explicit capability omission or CAPABILITY_UNAVAILABLE fact, never a silent disappearance. enable_tools/discovery answer a REGISTERED tool filtered by real policy with a typed "hidden by policy: " (ToolRegistry.policy_hidden_reason), never the same "Not found" as a nonexistent name; the contract-disabled check precedes registration so a disabled extension/MCP name also reports its reason. The swarm-router's promoted_task_toolset is one live top_level_tools projection with typed unavailable_builtin_tools; child allowlists remain deliberate narrower principals. Review output and cognitive artifacts remain outside ordinary result truncation. Outcome classification keeps policy refusal separate from execution failure: user_files_path_blocked, cwd_blocked, and artifact_output_undeclared are typed non-failure surfaces, while a declared output that cannot be registered remains the genuine artifact_output_error — an expected authority boundary never falsely becomes the task's headline failure.

Tool API v2 exposes neutral canonical names directly (read_file, list_files, search_code, write_file, edit_text, edit_batch, apply_patch, run_command, run_script, verify_and_record, service tools, commit_reviewed, vcs_*, schedule_subagent, schedule_followup, wait_task, wait_tasks); legacy public names are neither exposed nor translated. The file tools share a path-based public ABI, and because payload-borne paths (edit_batch entries, apply_patch targets) miss the dispatch seam that rewrites a path ARG, both ends canonicalize explicitly through tool_access.canonical_repo_relative_path — one normalization contract keeps a guard from judging repo/BIBLE.md while the write lands on BIBLE.md; _ROOT_ARG_REPO_WRITE_TOOLS is the single set every repo-write fence keys on.

Filesystem tool output is self-locating: results use canonical root:path labels and run_command/run_script echo the resolved cwd. A direct Project room selects one active physical folder for reads, writes, editing, process cwd, VCS and delegation; governance remains at system_repo. A selected missing folder keeps its address and warning rather than falling back to Ouroboros source. Plain folders support ordinary file/process work without Git; mutating delegation uses the existing snapshot capability and returns its typed target-specific failure when Git is unavailable. File bindings preserve physical identity: absolute paths inside any selected base normalize to that base; an outside absolute path is refused before safe_relpath can turn it into a similarly named file. Repo basename-prefix and canonical delegated-artifact read redirects retain their existing contracts. ToolContext.repo_path/drive_path use the same physical resolver, and unsupported roots of repo-only batch/patch editing reach the existing typed handler refusal before payload selectors are resolved. user_files is the first-class root for user-visible files under the owner's home (the Ouroboros repo and runtime control-plane are rejected); task_drive is task-scoped scratch; artifact_store is task-scoped under data/task_results/artifacts/<task_id>/, and external deliverables written through user_files or declared process outputs are copied there for audit — declared directory outputs as complete manifest+zip pairs written and hashed in chunks, and a rewritten user-visible file retains its previous copy under task_results/artifact_versions/<task_id>/ (see the §1 tree). Two READ-ONLY orchestrator roots complete the set: subagent_projects and deliverables grant read/list/search only to orchestrator profiles, so a parent can inspect a child's tree or a finished deliverable when synthesizing; a top-level task may still write the physical Deliverables container through the existing user_files-authorized paths. For argv-visible targets the shell guard checks the lexical Deliverables origin before generic roots, then the symlink-resolved destination, so hidden, credential-like, protected, and symlink-escaping descendants do not inherit a broader root's admission; the same target-first rule applies to declared-output custody and Presence ceilings, whose logical user_files-relative prefix survives a remapped physical binding. Pre-execution target extraction is segment-aware but conservative: shell -c recursion is bounded at three levels; a visible heredoc is interpreter program text only when there is no inline program or script-file operand, and shell stdin bodies recurse like -c. Python UNKNOWN remains unprovable even beside recovered targets. cd/pushd and env -C/--chdir update a sequential, symlink-resolved effective cwd for later/wrapped relative writes; find/xargs replacement words are templates, not concrete targets. Uncertainty widens only the owning row's tokens and inline/heredoc body, never independent provably read-only segments. The raw mention view stays separate for Windows drive/UNC spellings, and the light fence keeps its unfiltered inline-body view. This is not shell interpretation: computed destinations and unsupported wrapper grammars remain fail-closed only when write shape or uncertainty is visible. Runtime Light applies the same per-row writer targets and effective cwd to runtime-data access, expanding the existing known runtime/home spellings only on those targets and preserving separate secret/project-store read boundaries; a second substring write scan cannot turn comparisons or prose into writes. Root process writes may select any already-authorized user_files target independently of cwd, while acting children retain their own write root. shell_parse.local_shell_subject removes SSH remote command arguments only from filesystem writer-target inspection, preserving options, the local -E log sink and outer input/output redirects through the shared grammar; every other guard and execution receives the original argv. Child, external-task and Light read policies consume the same physical paths from shell_guards.shell_inspection_paths, preserving sequential and wrapper cwd without giving source names credential authority. Positional GitHub policy still classifies direct gh and shell-wrapper segments only; remote ssh ... gh auth remains an inherited residual. A remote body can still reach local files through an existing SSH trust relationship, including loopback; local deterministic inspection does not interpret remote effects or provide an SSH sandbox. Unknown option/heredoc forms keep the existing conservative inspection. Ordinary .config, Library, settings.json and exact ~/.ssh/config mutations use the existing resource authority; the enumerated owner locations in credential_shapes.owner_credential_locations, credential-leaf rules and VCS control directories stay protected. This is not blanket protection for all credential stores: unlisted locations such as .cargo/credentials.toml, .terraform.d/credentials.tfrc.json and .kaggle/kaggle.json retain ordinary root configuration access. Child read/list/search/query share source visibility for ordinary auth/ and tokens/ paths and public PEM certificates. Task and artifact files are not repository credential stores. core_secret_paths.restricted_data_roots anchors owner/control checks on the child, canonical parent and configured admission roots, shared with vision/media. Restricted file reads mask complete private-key blocks before selecting line/character windows, preserving positions and line breaks; known credential bytes remain masked on delivered views. The post-execution shell audit is still best-effort and cannot reconstruct post-cd relative writes, variable/indirect destinations, inline-code path construction, unwalked recursive copies, or inode aliases (shell_guards.py). Additional disclosed residuals: a cp/mv/ln invocation carrying an unknown value-taking long option keeps its operands mention-only; a Deliverables path used as a cp SOURCE is a read and takes no Deliverables target policy; write_shape's safe-stdio redirect set is matched by exact string, so a glued 2>/dev/null; still classifies a read-only line write-shaped — no block, but the refusal a protected-root read still gets is worded as a write; a quoted standalone '>' token is stripped as a redirect. The protected Bible/identity deletion predicate compares physical source operands per command segment: atomic replacement and unrelated names stay usable; working-tree restore/checkout/ordinary commit are not history rewrites. Computed deletion targets remain a disclosed parser limit covered by the chosen review/Supervisor contract. The redirect grammar is duplicated in shell_audit.py and spelled a third way inside `light_shell_repo_mutation.

Structured tools also accept a Docker workspace's explicitly mapped backend absolute address through workspace_executor.map_backend_path; the resolved host target must still lie inside active_workspace. Other roots and unmapped absolute paths retain their confinement. Query scope and read/search labels use the physical binding. Restricted repository views mask known token formats and PEM private keys while preserving ordinary long identifiers, hashes and source bodies; the opaque-run fallback remains enabled for owner-home output. Runtime secret/control entries keep their runtime-data predicate when that physical data root lies inside the project, across read, list, search and code queries.

Web access mechanisms (three distinct paths — do not conflate)

The three web paths differ in who chooses the query, which model reasons, which authority performs the fetch, and where evidence is recorded:

path reasoning model fetch authority evidence
Main-loop native search (provider server tool, attached only when the main-loop setting allows) the same solve model decides whether to search; no second model enters the scaffold provider-side citations and request counts fold into llm_usage and the task's host-attested retrieval fact — context for acceptance, never a criterion by itself; absence means only that this path recorded no search; the provider-side query is not available to the host and is not claimed as logged
web_search function tool a provider-backed call can introduce a second model and its own cost the configured web-search route or a keyless retrieval backend, executed by ToolRegistry arguments and bounded results in tools.jsonl; never native-search usage on the answering call
Browser tools (browse_page, browser_action) the main loop a local stateful Playwright session that can fetch or act on arbitrary pages arguments and result previews as tool evidence; browser state and local action semantics are not equivalent to either retrieval path

For ordinary delegated profiles, the URL, route, private-range, and control-plane guards apply to both local-readonly and acting children; Cyber-effective acting children inherit the selected access while explicit task restrictions remain; a concrete private origin is reachable only through host-established resource_policy.allowed_origins — ordinary Main/root scheduling names it in schedule_subagent, delegated and external Presence callers can only inherit or narrow it (tools/control_subagent_spec.py), and the origin normalizer rejects paths, credentials and wildcards. JavaScript evaluate is exposed only to a valid acting child on its current page; acting children retain the pre-existing shell-to-loopback /ws route without WebSocket authentication. This separation is methodological authority: an evaluation or acceptance claim must name the path actually used.

Context fitting, retry, and compaction

config.get_context_mode() is the effective Main sizing/rendering source, while config.get_owner_context_mode() is the persistent owner-intent/P3 source; they differ only during the auto-Low compatibility window. In Max, ARCHITECTURE.md is full-resident for every task class because it is Ouroboros's capability/tools/access map; in Low it is replaced by its lossless navigation map. DEVELOPMENT.md is mode-independent: full when the active repository binding says the work targets Ouroboros's own body; a bound external workspace, any subagent, or an API/CLI/scheduled external surface receives the visible on-demand pointer (context_requires_development/context_requires_self_body_docs override) — context economy comes from dropping the self-engineering handbook for external work, never from hiding the capability map in Max. prompts/SYSTEM.md and BIBLE.md are tier-0 and full in both projections (context_layout.TIER0_ALWAYS_FULL); the one path that omits SYSTEM.md sections is the local-model overflow compactor (llm.py::_compact_local_text), hence the prompt's load-bearing floor lives in its preamble. A tool's contract lives in its get_tools() schema, sent every round; mechanism documentation lives in this file — a prompt sentence restating either is a second copy that drifts.

context_fit.py renders Max and Low projections from one immutable core and measures each ordinary Main candidate on one labelled density basis against the selected route capacity. Owner Low adds an elastic 200K total-context economy target; crossing it is never a synthetic failure. With unknown capacity, owner Max gets one honest Max call while owner Low may reclaim toward its known target. Predicted Max pressure retains the Max document projection and may request one deficit-sized mutable-history pass; only an actual provider overflow authorizes task-local Low, and one same-route semantic recovery is permitted only when the final candidate has the same route, round, and response reserve and strictly fewer context-bearing bytes. Owner mode and P3 applicability never change.

context_compaction.py is a requested materializer, not a second threshold, timer, or retry authority. Pure selection chooses a positive-reclaim prefix of completed assistant-call plus contiguous matching-result units; owner turns and malformed, interrupted, or visually opaque units remain verbatim. A non-empty selection first writes an exact private checkpoint, then summarizes complete gap-free hashed map/fold input; independent covered units may apply while a failed unit stays raw. No eligible positive reclaim means no checkpoint, summarizer call, or transcript mutation; the one route+round latch prevents repeating the same automatic pass.

OpenRouter omits the neutral tool_choice="auto" while building the physical request, before request-wire identity is bound; explicit tool choices remain on the wire. This keeps an implicit API default from becoming an unsupported parameter under strict routing.

Ouroboros stores one provider-neutral, function-shaped conversation. Direct OpenAI agent/tool traffic remains on Chat Completions: openai_chat_custom.py projects only the physical copy into Chat custom tools and normalizes returned custom calls back into canonical function calls. The dialect ladder stays fixed at custom/function/none ordinals 1/2/3: generic parameter repair composes on the current rung before a dialect change, only an exact dialect rejection advances a rung, and the custom→function control-flow fallback is never learned as a durable dialect action; the compatibility layer never migrates the conversation to Responses or changes the model/provider/API surface.

request_wire_recovery.py is the single provider-neutral adaptation driver for both exception and HTTP-200 body-error shapes. It identifies one credential-free exact route/profile, applies fresh typed evidence before send, and composes only bounded set_value, drop_field, and registered replace_dialect actions; a learnable reactive action becomes durable (14-day cross-process store) only when a semantically valid normalized response is bound to the exact settled physical-attempt capture. Provider prose cannot switch routes. Every settled candidate is disclosed as usage.request_wire, aggregated in order as bounded request_wire_history — not the physical-attempt ledger; state/usage_attempts.jsonl remains the complete monetary authority. Private custom-argument receipts never enter stored history or public observability.

Direct Anthropic tool turns retain a private route-bound receipt containing the complete native assistant content list in original order, including thinking/redacted-thinking and signatures; the immediately matching same-route continuation replays it byte-for-byte before its tool results, and a provider, endpoint, API-surface, or model change scrubs the receipt. An unfinished native unit cannot be compacted; summarizer and public projections omit opaque values. Owner-requested none is sent as thinking.type=disabled; no guessed legacy budget_tokens mapping is invented.

Retry budgets are failure-class specific: empty/incomplete responses and transient 429/5xx failures may retry the same model with deadline-bounded backoff; auth, quota, permanent bad requests, and confirmed oversize fail fast; exhausted compatibility returns to the configured model fallback chain rather than becoming a second router. LLMClient keeps leading system messages authoritative and demotes later notices to visibly marked user notices.

Transport failures are classified by physical-attempt custody, and each class has its own owner and rails. A REMOTE pre-dispatch transport failure is transport_unavailable (loop_llm_call.classify_llm_exception): released custody ($0 — no request bytes left the host) plus the typed pre-dispatch predicate plus a non-local provider. It takes exactly one physical attempt per call, because pacing belongs to the round-level wait episode (ouroboros/loop_transport.py): the round gate latches an episode, optionally walks the existing fallback chain once when USE_LOCAL_FALLBACK makes it local (remote candidates never dial over a proven dead egress), then waits — durable network_wait events, owner progress notes that keep the idle rail alive, an owner-interruptible backoff sleep, and a free redial of the SAME round. A managed task waits as long as its existing rails allow (owner deadline minus the dispatch-admission reserve, budget, Stop, absolute ceiling), never a new setting, so a dead egress no longer dies after the transient burst (OUROBOROS_TRANSIENT_RETRY_MAX still bounds transient PROVIDER failures). Every turn stamped direct-chat (owner chat and Presence turns) or ephemeral — the interactive class — waits the same way but carries no queue rails and ordinarily no owner deadline, so its episode is bounded by the raw configured task idle timeout (OUROBOROS_TASK_IDLE_TIMEOUT_SEC): the bound limits idle WAITING, measured from each outage episode's entry (a flapping egress starts a new episode); a granted redial runs to its own connect timeout and a dispatched response is always accepted, so the bound never cancels in-flight work. When an explicit deadline window also exists the shorter one binds, and the durable ended detail names the rail that expired — the bound's own interactive_wait_window_exhausted, or the deadline's detail when the owner window closed first. Interactive progress notes omit cancellation promises; direct-turn Stop still follows its existing typed control. Closing network_wait rows describe the worker's cooperative exits, not every possible end of its process: external kill/Panic/crash or a direct-turn hard stop can leave no closing row. Correlate a confirmed task_done or durable terminal result for the same task; silence alone never proves completion, and immediate Stop never waits for a matching log row. Recovery is an owner note for every episode; local adoption and error-kind change are notes for interactive turns only (a managed task keeps its durable network_wait ended row and its ordinary progress); exhaustion is a note for an interactive turn, while a managed task's exhaustion is its terminal result; only an ephemeral turn's episode-boundary notes (entry, recovery/closure, exhaustion) carry the typed task_incident toast pair — the turn's one-shot urgency channel beside its live-card rows — each stamped with its valence as the optional toast_tone (entry and local adoption warn, recovery is ok, exhaustion is the error; the browser keeps the alarm tone for a frame that carries none); periodic notes stay silent. The mid-flight class is provider_outcome_unknown: a dispatched request without a terminal provider fact retains its unresolved monetary bound and is never resent. Ordinary managed cognition uses the upstream-observation continuation described under Review delivery: the existing wait admits a marked new physical attempt after connectivity returns. Other caller classes retain their existing rules. The interactive primary-round rail remains narrow and typed: when the death is a typed transport death (transport_custody.is_retryable_transport_death — httpx ReadError/WriteError/RemoteProtocolError reached through the explicit __cause__ chain the SDK sets, or a requests wrapper carrying ProtocolError/RemoteDisconnected in its own arguments; never a timeout, a provider status/body error, a pre-dispatch failure, a local provider or a loopback route), the PRIMARY main-loop round dispatch alone (_dispatch_round_model passes transport_death_retries) may repeat the SAME logical request at most twice per round, each repeat a NEW physical attempt with its own ledger row, re-prepared at send time — so a non-deterministic projection (for example a vision caption that failed on the first attempt) may differ between the attempts and may cost its own preparation call (the earlier rows stay unresolved at their upper bound, so a fully dead round reserves up to three upper bounds from the typed-death rail, on top of any unresolved rows the transient burst already left — honest accounting chosen over a cheaper rail), with a 4 s then 8 s deadline-aware backoff and a round record (execution_id:round:round_idx, persisted on the usage dict) that counts the repeats and carries the class its latest granted repeat failed with, once that failure is classified as an exception — so the terminal names that class rather than a later free redial's, and where no class is stamped (that repeat returned an empty response, or was refused before it was sent) the terminal falls back to the sticky kind: the empty response's own kind, or the unknown outcome itself — and that the wait episode's free redial of the same round neither clears nor re-arms; every llm_api_error row is decided before it is written and says whether a repeat follows, and llm_non_retryable_same_request marks only the exhaustion. The paid-repeat wait also observes current unseen typed finalize_now controls through the execution mailbox; ordinary dialogue, hurry and revoked/stale owner controls do not cancel it. A control or deadline refusal removes only the never-sent grant, records its own refusal row (finalize_control_pending distinguishes the control wake from deadline refusal), and exits through the existing unknown no-resend terminal. Prior dispatched/unresolved attempts remain accounted. The bounded repeat rail belongs to interactive primary rounds (direct-chat, Presence and ephemeral turns). Ordinary managed tasks and native API children use the upstream-observation continuation, while exact configured-session nanny routes retain their separate hold contract. Every other surface — forced-final, fallback-chain candidates, review actors, safety, probes, web search, consolidation/summary/reflection, Background Consciousness, external-harness delegated runs — keeps no-resend (their budget is the default 0, or they never enter call_llm_with_retry), and an unknown outcome that is not a typed transport death is never resent anywhere. A round that holds a transport-death repeat record (a granted repeat, no usable response since) sends nothing further except those typed-death repeats: a repeat that fails with any other class (a provider status, a transient, an empty response, a context overflow) ends the round on the same unknown no-resend terminal — no transient burst, no empty-response retry, no compaction retry, no forced-final dial, no fallback chain, and the wait episode's local-only pass is blocked as well — because the earlier request may still be live; only a repeat that never left the host (released custody) stays the free wait episode's to redial, under the same round budget, and when that episode's window closes the terminal is still the record's — source provider_outcome_unknown_no_resend, worded as both the wait and the unresolved attempt. The fence keys on that record, not on every unresolved attempt: a wait episode's local-only fallback pass that itself ends provider_outcome_unknown writes no record (the pass is not the primary dispatch), so the episode keeps its latched remote cause and its free redials of the round, while that dispatched local attempt is itself never resent. A configured-session nanny with exactly one live delegated leaf has its own way to start a new logical request after an unknown outcome: instead of dying — and cancelling a healthy leaf through cause-blind terminal cleanup — the round gate latches a durable unknown-provider hold (ouroboros/delegate_hold.py) and parks the task in the ordinary supervised_wait; a meaningful leaf wake resumes with a NEW round whose transcript carries the wake receipt — the unique host-attested input that authorizes a new logical request without resending the unknown one — while control wakes exit through the no-resend terminal and budget admission stays fail-closed against the unresolved upper bound.

Retry rails nest, each with its own owner and bound; the table exists so the multiplication is visible in one place (a physical send is always its own execute_physical_attempt lifecycle, whichever rail asked for it):

Rail Owner Bound What it repeats
SDK max_retries client construction (net_transport.py, llm.py) 0 everywhere nothing: every physical send is visible to the ledger
request-wire recovery ladder request_wire_recovery.py inside one LLMClient.chat at most eight typed compatibility actions per call a rejected request in a corrected wire shape
transport-death repeat call_llm_with_retry, primary round dispatch only ≤2 per round, 4 s then 8 s backoff a dispatched request that died with a typed transport death, as a new attempt
transient burst call_llm_with_retry OUROBOROS_TRANSIENT_RETRY_MAX attempts per call (the outer ceiling of the same loop) typed 408/429/5xx and empty/incomplete responses
round redial loop_transport.py wait episode a managed task: owner deadline, budget, Stop, absolute ceiling; an interactive turn: the task idle timeout measured from episode entry (OUROBOROS_TASK_IDLE_TIMEOUT_SEC), or an explicit deadline when that is shorter a released ($0) pre-dispatch failure of the same round
review physical rail review_substrate.py, review_native_episode.py 2 sends per packet or session actor on the P3/acceptance surfaces; a native-retrieval slot carries no send count (its bounds are the transcript bound, the owner deadline and the paid ledger); no rail elsewhere a released review send, never an unknown one

The cached OpenAI-compatible clients, the no-proxy per-call clients, and the web-search clients share one transport factory (net_transport.py) that sets platform-guarded TCP keepalive socket options — on Linux and Darwin the idle threshold, probe interval and probe count from config.py under hasattr guards (Darwin spells the idle threshold TCP_KEEPALIVE and otherwise takes the same interval/count constants CPython exports there, instead of XNU's 75 s × 8), on every other platform (Windows included) SO_KEEPALIVE alone — so a NAT/VPN mapping silently dropped during a long silent reasoning stretch is detected by kernel probes instead of hanging until the read timeout. When any proxy httpx would honor is configured, the cached and web-search clients skip the explicit transport (httpx env-proxy mounts require it absent). Proxy-routed installs, the Anthropic-native requests lane, every non-Linux/non-Darwin platform and a handful of library clients run without keepalive tuning — a disclosed residual (net_transport.py).

Main-loop model cognition is a separate typed in-flight fact: immediately before each exact provider-call seam the worker sends a direct supervisor started event bound to task attempt, execution, round, call id, and retry attempt, and every terminal sends the matching fact; a stale terminal from an earlier retry or attempt cannot clear the current row. The supervisor keeps only that process-local active row, with no elapsed-time expiry, consulted only by the idle predicate. OUROBOROS_LLM_TRANSPORT_READ_TIMEOUT_SEC (default 2700 s) remains a configurable dead-socket bound, not a cognition deadline; deadline, budget, cancellation, and the absolute ceiling remain independent hard axes.

OpenAI-compatible response choices keep their outer finish_reason as the bounded observational usage fact response_finish_reason; it never enters canonical history and changes neither the empty-response classifier nor retry policy. The trusted provider canary accepts a schema-valid native tool call even when the assistant also returns text, recording only length/hash as warning telemetry; a malformed native argument, invalid schema, or missing call remains a red contract failure — no prose parser, salvage, provider hop, or unbounded retry.

Prompt caching is stable-first: governance and task-stable contracts precede mutable evidence, review builders disclose the stable/dynamic boundary, and untrusted payloads stay outside the governance cache block. Provider-specific cache hints are sent only where supported and receive one exact retry without the rejected hint; rejection evidence is durable and route-specific. Gateway response-cache recovery is narrower and reactive: only a main-loop provider_incomplete_response arms a fresh-response request for later attempts, and only the generic openai-compatible route renders LiteLLM's extra_body.cache.no-cache control — direct providers and OpenRouter never receive it, and no URL/model heuristic exists.

Vision and local image evidence

analyze_screenshot and vlm_query are bounded secondary-model calls through LLMClient.vision_query; view_image attaches a local image natively. Send-time image routing works on a copy of the transcript, so captioning or placeholder conversion never mutates canonical history; payloads are validated, capped/downscaled, and confined to readable roots derived from the Tool API policy matrix plus the protected-artifact rule. vision.attach_local_image_to_context is the single attachment seam for explicit view_image and the typed auto_attach_image opt-in — same durable copy, trust boundary, and live-image eviction budget; auto-attachment failure is non-fatal. The loader's own refusals — missing file, outside the readable roots, oversized, unreadable, not a supported image — are published as typed TOOL_ARG_ERROR results before their text returns, so view_image and vlm_query record a known failure as one; a refusal from a policy owner keeps that owner's typed marker (blocked), and the host's same-round auto-attach, which runs outside any builtin invocation, meets no sidecar and stays non-fatal. Vision/local-media tools are not web tools and may be withheld by the task contract.

The existing uploads and skill-output roots resolve through the task's canonical data owner, matching the skill producer even on an isolated child drive. They count as independent image admission before user-home confinement; per-path secret, owner-state, project-store and protected-artifact checks still apply. The common loader copies an admitted image into the task's uploads/views before same-round attachment, so both the original output and its retained copy remain available to explicit readers.

Background consciousness and Evolution

Background Consciousness is the high-horizon awareness loop. It can inspect code through search_code/query_code, groom memory, identity candidates, knowledge, and the improvement backlog, message the owner, and initiate an already configured reviewed Presence binding; it does not acquire the binding's tools or transport authority, and it does not run shell/code work, subagents, reviews, commits, or evolution toggles — awareness proposes work without bypassing task, budget, and review authority. Awareness of the body's own version is ambient: the Runtime context of every task and every consciousness cycle carries official_update — running version versus the official target as of the last fetching check, plus the update letter (update_letter.py); no wake is forced by a check, and whether to mention an update to the owner is the mind's judgment (prompts/CONSCIOUSNESS.md).

Its observation inbox is the append-only state/consciousness_observations.jsonl store. Producers write a stable-ID enqueue row before the wake notification returns; a caller-supplied retryable ID requires a complete deduplication read, while newly minted IDs can still preserve fresh observations during an older source gap. _ack_observations appends ACK rows only after the thought receipt, tool receipts, budget settlement, and other durable writes for that snapshot have succeeded, so any failure leaves the snapshot pending with an actor-readable source reference. The status surface reports only pending count, oldest timestamp, source, and gap count. When the same cycle can call update_identity, an observation gap joins the identity-completeness envelope and blocks that destructive write until the actor resolves the named source.

An active campaign owns an explicit objective, campaign id, transaction, and task claim; evolution_mode_enabled is only its scheduling projection. Dispatch, review, commit, publication, and restart revalidate that exact authority; a restored row without a live uncommitted claim is cancelled, and a reviewed commit binds to the claim by exact SHA before publication. If authority changes after commit, the commit moves to a private inspection ref and any attempt-created tag leaves the normal namespace; concurrent index/worktree edits are not reset. Restart verification and boot reconciliation decide whether a cycle is absorbed, abandoned, or still pending, and exact terminal replay resumes only incomplete effects without double-counting. The exact restart-verify claim (state/pending_restart_verify.json; one writer helper serves both the supervisor's evolution restart and the agent's restart tool) is written whether or not the supervisor restarts automatically — OUROBOROS_EVOLUTION_AUTO_RESTART off skips only the restart — so the owner's manual restart verifies the cycle by exact claim rather than by the weaker markerless reconcile.

Campaign cleanup is deterministic and custody-aware: a no-op or abandoned cycle may restore the transaction base while preserving dirty/ahead work in recorded stash or local refs, and cleanup is skipped when another task, a live test, or the operator kill-switch makes reset unsafe. A byte-identical diff whose last terminal is a review-verdict block is refused for free from the first block; changing the diff, or a rebuttal whose content hash is new to the streak, creates a new paid reviewable case, and the shared cycle cap bounds paid triad+scope cycles per root task. Checkpoint and outcome rows preserve git/memory identity, cost, rounds, and explicit omissions so the promotion loop learns from failed as well as absorbed cycles. An agent-requested restart first drains heartbeat-fresh running tasks up to its configured bound, then fails closed rather than cutting another task silently.

Post-task evolution is owner-gated and default-off. The worker may recommend one backlog item by writing a request, but only the supervisor may convert that request into one normal campaign cycle. An owner stop closes the campaign, clears queued requests, persists a sentinel, and cannot be autonomously reversed. Evolution is hard-blocked in light runtime mode at the owner and post-task start entry points, again at idle campaign enqueue, and again at assignment; it uses the normal task/review path in advanced or pro.

Loop self-checkpoints remain plain user-message reminders: a second structured, tool-less reflection protocol in the hot loop produced unusable records and destroyed cache continuity. Durable learning belongs to the post-task reflection flow below.

Safety and runtime mode

Every tool call first crosses deterministic ToolRegistry and resource-root guards; policy-based LLM safety is added where OUROBOROS_SAFETY_MODE requires it, and the deterministic layers run in every safety mode. When the safety model is rate-limited past its one bounded retry, the guarded call is refused with the typed non-verdict ⚠️ SAFETY_UNAVAILABLE outcome plus a durable audit event — an unchecked guarded call is never executed and never accused (the local-FALLBACK lane instead fails open with an audited SAFETY_WARNING). The LLM check degrades to a visible SAFETY_WARNING in exactly three cases — no reachable safety backend; a remote key not covering OUROBOROS_MODEL_LIGHT's provider with no local lane; a local fallback lane whose runtime raised — and the deterministic layer stays in force in all three. runtime_mode_policy.py owns protected self-repo paths, frozen contracts, and release/managed-repo invariants: Light blocks Ouroboros self-repo and control-plane mutation, not normal user deliverables; Advanced may evolve ordinary app code; Pro and Cyber Pro permit protected editor/shell rewrites with the same notice; publication still uses the reviewed commit path. Cyber Pro additionally grants owner host/configuration access, including supplied credentials, model and Supervisor configuration and technical setup attempts. Review scope/enforcement and context Low remain owner-controlled in every mode. Explicit readonly/offline/benchmark contracts and named commit/cognitive child exclusions remain. Access is restart-bound; ordinary settings use the next-task snapshot. This is a self-modification boundary, not an OS sandbox.

Every deterministic write/owner-control guard consumes ONE mode-aware write-shape seam (ouroboros/tools/write_shape.py) — the registry does not even import the coarse legacy scan, so a guard structurally cannot judge on a coarser fact. Interpreter argv takes interpreter_write_shape, non-interpreter argv non_interpreter_write_shape; a provably read-only inspection reads (_is_pure_read_inspection, applied family-wide, including the protected-core mention branch), an unprovable case stays fail-closed, and write-mode opens, library save-APIs, and opaque subprocess/exec escapes classify as writes. The per-utility spelling inventory and its disclosed fail-open/fail-closed residuals live in tools/write_shape.py itself — chasing every alias with more spellings is the arms race BIBLE P5/P13 forbids, and the residuals are covered for external workspaces by the runtime/secret read guard plus the LLM safety supervisor. Workspace write-guard block messages name the resolved offending path and the sanctioned route. Write candidates are computed per command segment (writer_target_rows): a segment's parsed targets are write candidates and its remaining tokens are mentions, so the Deliverables decision and the outside-root refusal apply to write candidates only, while ordinary-mode protected runtime roots refuse on MENTION for every candidate — a writer naming a runtime path as a source operand still cannot launder a read through shell.

Read-only shell git is allowed everywhere; mutating shell git only when its resolved target is outside the Ouroboros system repository and runtime data drives (git_shell_policy; network-disabled tasks still fence network git). Acting self_worktree children remain read-only because patch capture requires an unmoved HEAD. git init/git clone are judged by destination, including relative destinations and path-valued retargeting flags; in external-workspace mode the runtime/secret read guard exempts only an all-read-only git command, and resolve_shell_cwd canonicalizes the cwd once for every guard. The generic Tool API VCS family (vcs_status, vcs_diff, vcs_pull_ff, vcs_restore, vcs_revert) defaults to root=active_workspace and accepts explicit root=system_repo; protected Ouroboros path names constrain generic restore/revert only on the explicit system target, so a project's own BIBLE.md or contracts/ remains ordinary project content, while preflight_review, commit_reviewed/vcs_commit_reviewed, vcs_rollback, and promotion remain system-repository lifecycles.

A task contract may declare resource_policy.protected_artifacts[] as execute-only black boxes: registry guards allow the declared execution but refuse reads, copies, hashes, static inspection, and trace/debug wrappers over those paths. Light-mode cognitive writes are redirected to update_identity, update_scratchpad, or knowledge_write instead of raw memory-file edits; a corrected cognitive redirect is advisory, while an ignored user-file root correction remains a blocking deliverable failure.

Review delivery

The Claude Agent SDK gateway is retired (owner-consented). review_native_episode.py is the generic native tool-round delivery for api_model configured-subagent reviewer rows — the read-only api-route advisory included, but not an advisory-only relocation and not a third public route kind: an episode is chat(tools=…) calls against a fresh instance-local inspection registry (read/list/search/query/vcs-status/diff only, local_readonly_subagent constraint, network off) until the reviewer answers — no round cap (P13; OUROBOROS_REVIEW_NATIVE_MAX_ROUNDS is retired): the bounds are the transcript bound derived from the reviewer's own window (never above the owner ceiling, except for a surface-declared mandatory reading — a window-capped floor, disclosed typed when short; measured on the serialized messages the next send would carry — every appended element is charged as its envelope plus the list separator, so the counter equals the wire size — a mandatory envelope that cannot fit ends the episode typed without another send), the owner deadline together with the coordinator's logical window for the slot (each send's transport timeout is clamped to the window remainder and a spent window refuses before dispatch; the LLM client's physical recovery ladder re-reads the caller's calendar/execution deadline before every recovery send), and the paid ledger (every physical send is a ledger attempt checked against the budget before it is made), plus two floors — a bound that leaves no room below the first send is a typed refusal (native_bound_below_first_send), and a round that carries tool calls but no well-formed one and no prose is a typed malformed end (native_round_without_progress — a post-send, settled end; an empty answer stays the episode's honest end on the empty-response rail); the pre-send refusals (native_bound_below_first_send, native_inspection_unavailable) are not_dispatched ($0) in review custody, a pre-send end leaves the receipt keys (resolved model/provider) empty so the public execution wire never mints a native run that never sent, and an episode that ran paid rounds is dispatched (settled) in custody whatever exception ended it — deadline, transport or the paid ledger (budget_exhausted); the host posts one [EPISODE_BUDGET] landing notice at 80% of the bound (each tool result is clamped to the room left below the bound so no single read can jump past it), exhaustion is a typed refusal (native_transcript_cap_exceeded) for verdict shapes and a disclosed native_incomplete product for the report shape, and every episode end — delivered, refused or errored — leaves one review_native_episode custody row and its facts on the actor usage (failure_custody for the error actor). The default output contract of a retrieving row follows the surface's output shape (triad_review.default_output_contract) whenever the surface does not hand over its own policy["output_contract"]. Mutating external coding work uses the delegated subagent path. Two live residues of the retirement remain on purpose: reviewer_slot_config.py migrates legacy Claude-SDK advisory targets at parse time (_migrate_sdk_advisory_target; an unmapped target is typed legacy_claude_sdk_target_unmapped), and the legacy filename tools/claude_advisory_review.py hosts the live generic advisory gate.

Late review completion is retained in the existing operation-addressed prompt/response CAS. The response manifest stamps a complete producer outcome and the original task/root, attempt, slot/route, operation, subject, contract and roster/epoch binding. Exact reconciliation reads and verifies that complete source and runs the existing surface reducer; a new context needs no second paid attempt or surface writer. Pending, unreadable, mismatched and unknown outcomes retain custody and their full source references. Plan waves carry their original dispatched set, health epoch and operation IDs through the same reconciliation path; free replay spends no new cycle. Skill Review reserves every chunk digest/retry key and slot operation ID in the existing review_job.review_wave before paid dispatch, then retains that binding in terminal history. The next authorized public call may record review_resume_of only for the immediately preceding unsuperseded wave with identical task/root/attempt, skill/state roots, group, content, contract, rebuttal and chunks. Its new lifecycle reaggregates complete CAS at the original wave ID without another paid stamp, even at the cycle cap; old timeout/history rows remain unchanged. Missing or partial producers stay pending and unstarted chunks do not acquire verdict authority from their reserved IDs. A new task/root, changed material/contract/rebuttal, superseding lifecycle or explicit cancellation cannot inherit the old wave. Review-state persistence, current grants/dependencies and enablement still follow the current lifecycle guards; late completion alone never revives a job or enables a skill.

Main remote completions use transport streaming, assembled inside the physical send closure before accounting settles. Compatible choices/tool-call fragments and native blocks, reasoning/signatures, usage snapshots and terminal framing produce the same normalized response shape as JSON. Partial streams never yield usable tool calls or answers; their exact wire bytes and partial assembly remain in private CAS through the existing physical_stream manifest and attempt ID. A complete response with absent final usage still has unknown money. Comments/pings do not define cognitive deadlines. Every recovery candidate checks inherited calendar and quota-adjusted execution bounds before reservation, after preparation and at dispatch; HTTP phase bounds remain distinct from an overall logical wait.

An ordinary managed task whose provider outcome becomes unknown now stays in the existing transport-wait episode. The old attempt and unreported cost remain unknown. After non-generating upstream observation, a user-role [SYSTEM NOTICE] supplies explicit recovery input for one new physical attempt; existing budget, cancellation, owner deadline and absolute ceiling still apply. Finished tools are retained. Direct/ephemeral turns and configured session-nanny custody retain their own contracts; manual Restart/Panic gains no resume authority. Direct remote endpoints are observed through HEAD with the same no-proxy policy. The metadata HEAD reuses the ordinary connection allowance for every socket phase, narrowed by the owner remainder; it holds no cognitive in-flight lease. Subscription metadata qualifies only when the existing catalog reports generic provenance="provider_http" and an original observedAt after wait entry for the selected source/model and effective profile/account fingerprint. Cached reuse never advances that timestamp; a local handshake, static or pre-outage catalog, or timestamp without provider provenance cannot prove recovery. The pinned 3.10.4 raw-model adapter is Codex; capability discovery remains authoritative rather than a new core provider table. Loss of the Claudexor control connection first keeps reading the same accepted model operation, including across endpoint rediscovery, without creating another operation.

Terminal task delivery and later inspection derive a bounded host notice from the existing delegated-custody audit after cleanup. Confirmed terminal cancellations, live runs, unresolved invocation IDs and undisposed patches stay distinct. The host notice accompanies unchanged model text and continuation narrative, so a later cancellation does not rewrite historical authorship or the answer hash. A valid acceptance PASS with partial/missing/rejected criteria remains a parseable contributing verdict; only the existing clean/applied host decision can authorize objective completion.

Commit triad, scope, advisory and task acceptance pass the owner's deadline_at to the native episode; the slot's logical window remains an independent bound. The native episode owns an empty scratch data plane unless the surface opts into its real data root through policy["native_data_root"]; only the scratch is removed at episode end. review_native_transcript_bound applies the window and owner ceiling, while policy["native_mandatory_read_chars"] can lift the bound to the window-capped native_mandatory_read_bound. Actor usage records the declared chars and any native_mandatory_read_disclosure as native_mandatory_read_exceeds_bound, so mandatory reading never silently contradicts the episode budget.

Native read evidence is recorded on actor usage as at most 200 native_tool_receipts, with outcome = executed/refused/error/withheld and host_file_read_attestation: host_observed. An executed read_file receipt also carries start_line, end_line, total_lines, eof, opened_path and opened_root: the complete lines actually delivered, on the normalized file the reader opened, from ctx.last_read_view in tools/core_file_tools.py. The reader resets/stamps that view and the episode clears it before synchronous dispatch; tests/test_native_tool_round_executor.py pins these three writers. _read_extent owns the edge rules and cuts coverage to the delivered prefix; a missing stamp proves no extent. This distinguishes a bounded read from full coverage without trusting the model's path spelling.

Every native end publishes native_rounds, native_tool_calls, native_transcript_chars (last physical send), native_transcript_bound, native_transcript_refused_chars when a next send did not fit, native_landing_notified (posted) versus native_landing_sent (physically sent), native_end_reason and native_custody_row. A non-delivering or incomplete end after any assistant round also retains native_terminal_round, a bounded, redacted, structurally valid JSON document; receipts alone cannot reconstruct the assistant envelope and tool results that ended the episode. These facts remain on failed actors through failure_custody, including deadline, ledger and transport ends regardless of whether the landing notice was sent.

Review delivery has two closed route kinds in review_execution.py: api_chat and agent_session; vendor and harness names are route targets, not new kinds. A slot is bound to one immutable route before its first send and never falls back to another transport. The API executor lazily memoizes the assembled review messages, so its durable prompt record and its at-most-two physical sends use the same bytes; one logical interaction may use one bounded second send for transport or empty-output recovery, never while a dispatched outcome is unknown (task acceptance may also spend it on malformed-format repair). That two-send rail belongs to the PACKET api row: a native tool-round episode (an api row bound to a configured subagent) carries no send count — its bounds are the ones named above — and neither retrieving delivery is sent a format-repair resend: their executors canonicalize their own answers. The hosted-agent executor instead starts one read-only delegated session through the shared Claudexor nanny: route-owned instructions and retrieval pointers, its own tools, no assembled API pack; extraction canonicalizes the already-collected transcript and never launches a second hosted session. Custody, cancellation, full-artifact recovery, and settlement stay on the delegated transport contract.

Advisory availability is evaluated from the current configured slot and route, never inferred from a stale stored verdict. A disabled advisory slot is an audited bypass; an api_chat row requires provider credentials for its RESOLVED model, an agent_session row a resolvable session route. If the commit advisory is unavailable, the commit gate runs its compensating hermetic preflight only when tests remain independently applicable (not explicitly skipped, diff not documentation-only). Malformed structured slot configuration is refused at save and becomes a typed loud review-time failure for commit triad, scope, advisory, plan, skill review — and deep self-review (deep_review_slot() raises on the malformed value and run_deep_self_review returns the typed deep_self_review_unavailable result instead of a report). Task acceptance refuses the same way (a typed DEGRADED panel, reviewer_slot_config_invalid; owner R3). No surface silently chooses the opposite route or a default panel.

The reviewed-commit cycle checks authorization and unresolved prior work before mechanical preparation and staging, then binds the exact candidate before its existing free-cycle/budget admission. Needed preflight runs inline with the full rebuttal and applicable tests; the index and worktree snapshot are revalidated before triad/scope. git_review_cycle owns entry-point resets and pending/blocked finalization. Its custody check joins current reviewer facts with the strict durable advisory record before index cleanup, including interruption before local metadata was updated. AdvisoryRunRecord.execution_pending preserves physical custody for history retention and the external wrapper; blocks_preflight separately tracks logical admission. An explicit audited bypass releases that admission while retaining the original task, invocation, source and unknown physical outcome. Late results update their original history row without superseding the newer bypass. Free advisory replay still checks freshness and buys no automatic preflight; stale coverage requires an explicit audited skip, with applicable compensating tests preserved.

AdvisoryRunRecord.execution binds delegated preflight intent and candidate to the existing invocation token before POST. The same row retains terminal usage/receipt evidence on failure. delegate_custody.invocation_record owns the immutable request, including the full prompt; no second prompt store is created. An exact pending rejoin restores that request through the shared executor. A delivery-kind change cannot replace an unresolved session; same-session model/profile changes still rejoin the canonical request. Changed/foreign/lost identity is refused before preparation, naming the mismatched fields; review_status projects the recorded rejoin intent and execution identifiers through the existing public secret redactor, without exposing the private reviewer prompt. Failed checkpoint writes cannot dispatch. Unresolved rows survive history trimming. Only a corroborated failed_definite invocation discharges a stranded logical checkpoint automatically; age, missing identifiers and unreadable state cannot prove never-started work. An audited skip never posts a replacement preflight. Released history remains eligible for an exact delegated rejoin, but does not lock a later explicit standalone request for different intent or evidence. An unchanged audited bypass can satisfy the existing freshness shortcut without a new model call; that response names the actual bypassed status. Native preflight uses its existing executor's end-event and monetary custody, without a separate advisory operation checkpoint or native resume protocol. Returned unknown physical outcomes remain failed evidence and preserve the existing refusal and audited-skip paths; a hard process death has no native exact-rejoin handle. The external wrapper runs the same cycle, exports the real advisory record to advisory.txt, and preserves a pending checkout without restaging.

Technical failure and commit permission are separate facts. Producers retain failure origin (context, delivery, format or window authority) alongside the original status, source and findings; the shared physical exception projection keeps budget/admission and authority refusals distinct. Under advisory enforcement the commit aggregate can continue on a known bound candidate; under blocking it still refuses. Neither path turns missing review into PASS or changes quorum. Custody, Stop/deadline, ownership and candidate binding remain independent. Readiness projections receive only an exact repo/hash-matched advisory record and keep its failed status/freshness. Contributor slot evidence retains configured-subagent identity while its stored wire uses the existing mutually exclusive reference-or-route form, preserving native retrieval and its API budget admission.

Session reviewer identity comes from gateways.claudexor.final_attempt_facts: the unique final_attempt_id row in the engine-owned final/telemetry.yaml, bound to the requested run id. Model, harness and credential profile come from that SAME attempt. summary.model/harnesses echo requests, and summary route/auth projections may borrow earlier-attempt facts; none supplies a missing observation. Reviewer usage, last-execution views, delegated terminal payloads and settlement preserve known facts or explicit absence. The original custody/billing route remains its chosen authority, separate from the observed actor. No new quorum rule or model-name mapping is implied by this disclosure.

Advisory row parsing owns its PASS|FAIL and critical|advisory values at preflight_review_run._is_checklist_array. Case and surrounding whitespace are canonicalized once for every consumer. An unknown verdict or unknown/missing FAIL severity rejects the whole array into the existing bounded extraction rail; unresolved output stays parse_failure with its full source retained. PASS without a severity remains compatible, and the separate genuine-empty-clean predicate is unchanged. An optional array validator lets the shared canonicalizer honor this surface contract without changing triad, object-verdict or report semantics. Canonicalization never changes the reviewer's judgment by searching the repository for words or identifiers.

Usage ledger substrate vs. accounting policy

usage_ledger.py owns the durable append-only physical-attempt ledger (cross-process locking, sequence and transition validation, append+fsync, replay, loud tail quarantine); usage_accounting.py is the one-way policy layer above it (pricing, reservations, settlement, scopes, budget fences, imports, projections, admission), and the substrate never imports policy. The boundary means a pricing or budget-policy change cannot redefine valid ledger storage, and a locking or repair change cannot silently change what an attempt costs; compatibility events, state mirrors, task fields, and UI projections may carry attempt ids and derived totals but never become a second charge source.

Caller-owned subscription model calls

claudexor::<source>=<model> selects a model operation in the Ouroboros-owned engine, not an Agent Run. LLMClient dispatches sync and async calls through llm_claudexor.py before OpenAI-compatible filtering/retries. Ouroboros retains its SYSTEM, BIBLE, canonical messages, tool selection and execution. The initial raw-model adapter is Codex; connected Claude/Cursor and other harnesses keep their existing Agent capabilities. Direct API keys keep their existing routes.

Subscription image capability comes from the selected role/account's model catalog, not the global model-id overlay. Main send preparation, browser image attachment, captions and registered VLM tools use that same capability reader. Missing metadata stays unknown: image input is retained for the real call rather than silently omitted on a cold engine. A confirmed text-only model keeps the existing caption/capability-gap behavior. Text-only messages and image-off mode do not read the image catalog. Metadata helpers use read_owned_gateway (owned-only discovery and handshake); they never install or start the engine. Actual model calls retain ensure_owned_gateway, as do explicit lifecycle actions.

The send copy omits foreign provider envelope metadata such as stop_reason; canonical history retains it. A nonempty refusal is preserved as assistant text, including when the prior provider supplied no ordinary content. The exact model_request_invalid create refusal proves validation failed before command admission; a generic HTTP error or failed status read cannot prove non-dispatch. The send copy also normalizes all-text tool results, so moving the message-side cache boundary does not rewrite bytes already sent; non-text blocks remain intact.

ClaudexorModelError.display_message adds sanitized typed vendorCode and parameter details to the existing error event and terminal preview only after classification. Exception text and machine fields retain their retry, compaction and custody semantics; unknown outcomes do not expose underlying provider details.

ClaudexorGateway uses purpose-bound byte uploads, one idempotent operation ID, status/result reads, cancellation and explicit result acknowledgement. A lost local HTTP reply rejoins the same operation; an unknown provider outcome cannot authorize another generation. Claudexor performs at most one generation per operation. It shares its managed profile with Agent sessions, while its adapter reads current access credentials and the official CLI retains refresh ownership. There is no shared external engine, relay process, copied operator auth store or public OpenAI-compatible endpoint.

Physical attempts remain in the existing usage ledger. Exact zero cash, known charges, estimates and unknown cash stay distinct; a subscription is not evidence of zero cost. The existing private observability CAS retains exact request/result bytes before ACK. Claudexor removes terminal request bytes and acknowledged result bytes; an unacknowledged ready result expires after 30 days. Compact operation receipts survive body deletion, so an expired read cannot start a new generation. Model-purpose resources are not valid Agent attachments. Private continuation envelopes remain in canonical history but are excluded from public/summarizer projections; reuse is bound to the actual source/model/profile/account identity.

Beside that assistant-level history the engine keeps ONE live turn per model client session, and its opaque token is transport rather than content: it belongs to the caller that is running, not to any stored message. So the caller owns a single mutable slot (llm_claudexor.ModelTurnState) that rides the ordinary LLMClient.chat parameters to the engine boundary, and that boundary is the only writer — a dispatched durable result replaces the slot's value, while a not-dispatched or unknown outcome, or a legacy-shaped exchange that never carried the field, leaves it untouched, because holding state never licenses another generation and silence about a turn never ends one. WHY a slot rather than the transcript: the live turn survives compaction that drops the messages and outlives a body that fails after its headers, so reading the last stored assistant envelope would rebuild the wrong lifetime. One run_llm_loop invocation is one turn (rounds, owner steering, acceptance follow-up, forced finalization, quota waits and reprepares all continue it; a next loop and a cold restart start empty), and one Background Consciousness wake is another, keyed by that wake's existing model-wait owner id. A dispatch that leaves this transport for a direct API or local route ends the turn at the caller and never revives it on return, while the engine alone compares route identity. Opting in at all requires a serving engine at config.CLAUDEXOR_MODEL_TURN_STATE_MIN_VERSION, read from the last SUCCESSFUL handshake (a failed probe never un-proves it, so concurrent status polling of that singleton cannot change the shape a running caller sends), so an older or not-yet-observed engine keeps the legacy request shape instead of risking a schema refusal. The token is never usage, an event, a progress note or a task card; mechanism in the llm_claudexor.py docstring.

Model roles own account pins and context assertions through OUROBOROS_MODEL_ACCOUNTS and OUROBOROS_MODEL_CONTEXT_WINDOWS; ordered fallback entries use their original ordinal. Reviewers and configured actors reuse their existing route credential fields. Identical Main/Light or reviewer model strings never imply identical roles or accounts. Auto uses the largest advertised window of the exact account route, not CLI compaction thresholds. A manual value is a sizing assertion, not a provider unlock or a scope-review acknowledgement. Unknown capacity stays unknown; input size, response reservation and capacity are separate. Each model operation records its submitted options beside the engine's applied options; an absent applied-options report remains explicitly unknown, while the first changed reasoning effort of each model in a task also produces one typed owner-visible notice. An account change rebinds preparation before another physical send. Ordinary sends, prospective wrap-up payloads and forced final replies share the same acting-role/account binding. Prospective subscription accounting uses the same model-request serializer as dispatch, not an OpenAI-compatible approximation. Packed deep self-review rechecks the same full packet's window and input cap when its waiting role changes. The final header identifies the actual responding model/account window. An observed-only sub-floor result retains its complete paid report and custody with the existing incomplete/failure classification; it never silently switches to retrieving delivery or inherits the previous route's window.

Host-authored sampling defaults remain separate from explicit parameters. Review requests carry default_temperature as a hint; each effective API route resolves it at dispatch, while raw model operations use their provider default. An explicit temperature, including zero, is preserved and may be refused. Routes without structured output use the existing safety text-JSON parser, repair and refusal path; they do not acquire a new safety bypass. Ordinary confirmed provider errors retain each helper's previous handling, while quota, owner interruption and unknown physical custody propagate without becoming a helper result or verdict.

model_wait.py holds confirmed quota/auth waits inside the live call. Typed owner actions use the existing mailbox and /api/decisions family model_wait:<task>:<wait>; task-result rows are projections, not restartable stack checkpoints or a second attempt ledger. Metadata polling makes no generation. On Auto, the host may send the last successful same-route account as a preference, but omits that preference for the next request after a status-null or typed per-subject refusal on that account in the same execution. The engine remains the account chooser: without engine refusal evidence it may select the same top-headroom account again, so rotation is possible rather than guaranteed. Pin never rotates. The existing fallback budget remains unchanged; its defaults are one attempt per model and a 120-second cross-model cooldown. Even a configured API fallback waits for an explicit owner switch after a quota refusal. A temporary switch changes only the waiting role; optional persistence uses the ordinary settings writer. Pending acceptance, settings saved and worker application are separate revisioned facts. Fallback Local remains a group-wide setting: a one-role persistence request may change model/account only when Local is unchanged. A different Local choice is task-only; permanent group changes remain in Models. Current task-attempt identity selects actionable wait rows independently of their retained history. A failed mailbox delivery can retry the same accepted request without authoring a new one. Terminal mailbox cleanup waits for both post-task synthesis and pending accepted attachment promotion to settle, retaining the same mailbox while either is owed.

ouroboros/gateway/task_model_wait.py owns model-wait decision effects and the existing bounded settings-writer receipts. task_decision.answer_decision returns (status, payload) to both the Web and authenticated Host Service transports. They supply the installation root and the same live background-owner getter; no synthetic HTTP request or parallel owner registry is needed. supervisor/task_model_wait.py projects worker events and quota clocks through the existing supervisor. Shared wire types remain exported by gateway contracts; ouroboros/gateway/decision_contracts.py owns their decision-family vocabulary. An ephemeral turn binds its live wait to its existing DirectActivityRegistry entry. Both event forwarding and decision ingress resolve that same owner; wait rows remain in memory, with no durable task record. Scope close fences late decisions and cleans its control mailbox. The activity snapshot and retained wait events preserve ephemeral identity, so a fresh or reloaded card offers model-wait controls without Cancel or Turn into project.

Quota pauses use union duration across simultaneous waits: they do not consume internal execution time, while explicit calendar deadlines remain fixed. The existing worker and project-write ownership stay held; a one-worker queue waits. Completed tools and reviewer results remain on the same live stack. Closing a browser does not stop it; stopping the Ouroboros process ends this continuation guarantee. Image tools retain the tracked VLM child: vision_process.py carries typed errors, physical capture, cancellation and operation checkpoints across that existing process boundary. A lost child result remains unknown, never free.

Background wakeups reuse the same controller with their existing live owner and locked memory instead of a queue entry or task-result archive. Foreground pause defers a healed call until resume; stop uses the background stop event. No new absolute background-cycle deadline is imposed; lexical tool deadlines remain. The existing background card hydrates from the current cycle snapshot rather than a historical completion marker. Temporary changes last for that wakeup cycle.

Delegated subagents (Claudexor transport + the nanny)

Ordinary delegation requests no extra engine review panel; new ordinary runs on Claudexor 3.9.8+ default to no panel. This is the behavior's version boundary, not a second release pin. The start receipt's engine_version is the handshaken serving version, distinct from the release pin. Older serving engines and already-recorded runs can retain historical review behavior. Engine review outcome, execution success, the parent's integration decision and Ouroboros review gates remain separate. Timeline projection retains each known participant's harnessId/attemptId, including non-text reviewer events, and names them on live progress without guessing absent identities. Delegated snapshot capture remains relative to its recorded baseline: committed bytes can be captured with a disclosed head_moved; the instruction still forbids committing, and the distinct self_worktree unchanged-HEAD check is preserved.

Read-only children can read/list the existing project-scoped knowledge store without writing its index. Children coordinate through tree_note and tree_read; only the parent may use override_delegation_constraint, and a review_requested note carries an exact evidence reference/hash and wakes the parent without starting a paid cycle. Both read-only and acting children hold the descendant-scoped forward_to_worker, peek_task, cancel_task, and discard_child_result controls; recursive delegation never widens filesystem, budget, depth, deadline, commit, or owner authority. delegation_budget governs descendants (may_delegate, may_fan_out, additive depth provenance; a free-form intent note is never authority); persisted admission facts outrank later Settings changes, and a lower permitted depth is reported capability_reduced, never a silent flat tree.

Registry. OUROBOROS_SUBAGENTS is the active task-actor SSOT: a strict {enabled, items} value of at most ten ConfiguredSubagent rows — stable subagent_id, owner-authored English recommended_use, one normalized route (api_model or agent_session), optional effort and session credential pin (configured_subagents.py; legacy env keys are fail-closed migration inputs). The description is selection context only — host code never parses, ranks, or maps its words to task text — so API models and session harnesses occupy one LLM-selectable list without pretending they share a topology; an exact owner selection either starts that route or reports why not, never a keyword router or automatic API substitution.

Scheduling. schedule_subagent requires subagent_id, a focused objective, and expected_output; the remaining public fields describe child-local context, constraints, memory, capability needs, write surface, narrower deadline, delegation budget, and acceptance claims. There is no model-visible lane/executor axis and no public effort override: the selected row is the complete execution choice, lineage/bounds/route/budget are host-derived, and omission never inherits the parent's acceptance claims. subagent_runtime.select_subagent_snapshot copies an immutable snapshot of the exact enabled row into the child task; an api_model row becomes an ordinary recursive API child on that exact model/effort, an agent_session row an ordinary recursive Ouroboros nanny on that exact external route. The cash side of a burst — each sibling launched before the first sibling's first response pays its own prefix write on cache-write-priced routes — is disclosed to the mind in the tool description as an affordance and deliberately not scheduled by the host. The old lane/executor resolver serves only old durable records; for those historical lane envelopes, schedule_subagent reports the requested lane only — effective facts remain on the dispatched child record — and a task carrying configured_subagent goes straight to subagent_runtime, so legacy policy cannot reinterpret an active selection.

Waiting on children. wait_task/get_task_result return the full single-child handoff (verification receipts red/masked-first, exact omitted count); wait_tasks stays batch-compact: task_id, status, child_result_sha256, outcome_axes, result, terminal_host_notice when present, trace_summary, capability_delta when the child has something to disclose, duplicate_of, plus the nullable cost-finality pair accounted_upper_bound_usd/cost_final and, when the child's envelope carries one, execution_evidence (§11.1). The retired cost_usd spelling is tolerated only when reading stored rows, never emitted. Both use task_status.SETTLED_STATUSES; a pending cancellation is the typed cancel_state: "pending" projection, never completion. A batch wait that expires with children still running discloses what it could not finish: the typed wait_expired_with_live_children block names the live child ids, the window that was requested and the clamp ceiling the schema already states. Facts only, with no advisory text in the payload and no host floor on the next window, because how long to wait is the mind's call (BIBLE P13); an id this tree never minted stays an unknown_task_id and is never counted as a live child, and the pinned per-child field list is unchanged.

What a delegated run costs. Claudexor reports the amount in summary.spendUsd and its exactness in summary.spendEstimated; delegate_custody.disclosed_spend is the single reader of the pair, so the ledger row and the payload the nanny relays cannot tell different stories. Runs ask authPreference: subscription explicitly, because the engine default falls back to a paid key invisibly. Four cases:

disclosure ledger finality
disclosed settled zero 0.0 (subscription_session row) cost_final=true — the proven free session
disclosed settled charge the amount final
estimated amount the amount cost_final=false — an estimated zero is not a proven free session
undisclosed cost_usd: null, increments unknown_unmetered drops cost_final for the projection

An undisclosed spend contributes 0.0 to accounted_usd — inventing a conservative bound would fabricate a number the harness never gave (BIBLE P1) — so a TOTAL_BUDGET fence cannot stop spend it was never told about; the honest consequence is loss of finality, not a guessed charge.

What a delegated run READ. Harnesses count input tokens with incompatible conventions, so settle_run carries the engine's own normalized split — summary.inputTokenUsage — onto the subscription_session row as the optional input_token_usage object: total_tokens, cache_read_tokens, cache_write_tokens, each a nonnegative integer or null for unknown. record_subscription_session is the one validator and the one writer: all three keys are required together, and a partial, extra-keyed, negative or fractional object is unknown as a WHOLE rather than repaired field by field, because a repaired counter reads exactly like a measured one. An engine that reports nothing leaves the row as it was, and the object stays OUT of the idempotent row identity, so an engine that starts reporting it never rewrites or duplicates a session already settled without it. The legacy prompt_tokens/cached_tokens axes keep their own meanings and their own harness-specific semantics, and compaction never folds these idempotency-bearing rows, so the object survives verbatim. Direct physical attempts already carry canonical prompt/read/write counters and take no copy of this one.

The nanny model. An agent_session subagent is an ordinary recursive task-tree child acting as a nanny supervising at most one active bounded external leaf. The task node keeps lineage, authority, deadline, budget, acceptance, cancellation, and descendants; the harness process remains a non-recursive tool leaf — session rows never flatten the task tree into harness processes. The nanny is the host: verification receipts stay host-authored, and harness output is a claim to check, never proof. A nanny-to-nanny chain through schedule_subagent is the host-attested realization of a nested subscription swarm, each level one metered supervisor task plus one free harness run; schedule_subagent.requested_depth is the typed ABSOLUTE request counted from the root, recorded as telemetry that never narrows the configured caps, while the legacy depth_remaining envelope keeps its narrowing semantics — two semantics, both disclosed on the contract, and the root's swarm_efficiency.depth block plus its terminal summary row report requested, permitted and achieved with the typed status. A nanny terminal names BOTH routes as separate facts on projections that already exist: the host's own model_execution carries its model, provider and last typed error (last_llm_error_kind), while each terminal_runs row carries the leaf's model, profile and selected actor. A nanny that died on its own lane is therefore never read as a verdict about the leaf, and a leaf half whose reconciliation was never persisted is omitted rather than guessed. Harness-agnostic by construction: the row holds an opaque Claudexor target, Ouroboros asks for an access profile derived from task authority and lets Claudexor choose the mechanism; no harness-name branch selects a capability or fallback in core dispatch, and login-wire asymmetries stay presentation adapters in gateway/claudexor_accounts.py.

Transport. gateways/claudexor.py is pure transport (descriptor read, /v2 handshake, config.CLAUDEXOR_MIN_VERSION 3.2.0 floor). The daemon bearer token grants the entire /v2 surface, so it never leaves this module, and the HTTP client runs trust_env=False so an ambient proxy cannot intercept the loopback control plane. Production starts obtain a handshaken owned gateway from claudexor_daemon.ensure_owned_gateway (exact reviewed engine/Node pins; never a PATH install); keeping lifecycle above transport keeps account status side-effect-free and harness mechanisms out of Ouroboros (claudexor_daemon.py/claudexor_runtime.py docstrings).

Custody is durable, because the run is not ours to kill. A delegated run lives inside the daemon and survives our worker, so the AUTHORITY is the durable custody rows on the canonical/budget root (delegate_run_started and friends; ouroboros/delegate_custody.py); the module dict is pure memoization. A lookup answers OWNED, FOREIGN, or UNKNOWN — collapsing UNKNOWN into "not yours" made a restarted owner indistinguishable from an intruder. Every INTENDED start mints a fresh per-intention invocation UUID as the wire Idempotency-Key; the content hash is only the LOOKUP identity — a content-stable wire key would hand a deliberate re-run the finished old run — and reuse happens only by explicit token (pending_invocation_id/retry_of, replaying the STORED canonical body under the SAME key). reconcile_orphaned_runs visits every open run whose owning task left the live set (the reaper's owner-is-gone predicate: a delegated run has no pid) and settles the terminal ones, but it CANCELS only behind a deliberate verdict. The periodic sweep reads that verdict from the durable owner result: readable, truly terminal, and an execution axis saying the task finished by its own decision. The owner-cancel kill boundary instead STATES the verdict it is itself about to write, because it audits custody before that write (the A4 ordering), so an owner cancellation stops the paid run at the boundary rather than at the next sweep; a task that had already settled when the cancel arrived writes no cancelled verdict of its own and falls back to the durable rule. A provider or transport death, a worker crash, a missing or unreadable result all SPARE the run, which is left live with a durable left_live reconciliation row: an undignified nanny death must not kill a healthy paid run (owner answer B1-A), and "unknown" is exactly the case that answer is about. maxSeconds is the damage limitation for a spared orphan, never custody. SETTLED is published before registration retirement, after the ledger row lands; settlement and registration retirement are separate durable duties, and a failed retirement remains replayable on project_owned for the later sweep. Writing settled over a suppressed ledger failure turns a lock timeout into a permanent leak, and a start whose row did not land reports started_uncustodied: no supervision, no replacement, until the original run is proven absent or terminal.

No terminal or cancel claim without a verified receipt. delegate_cancel returns confirmed (read back terminal), requested, failed, or containment_fault_run_may_still_be_live; the last two hold a durable CRITICAL containment fault until a receipt or settlement clears it — an overpowered run that may still be alive is an incident, not a reassuring string — and the state read decides, so a refused control is never a verdict about the RUN. A 404 is scoped to the daemon that answered; custody closes such a run delegate_run_closed_absent (unreachable, not settled), inventing no terminal detail, usage, or spend. Results are delivered, not severed: delegate_wait stages the whole terminal detail under task_drive/delegated_runs/<run>.json with a typed output_delivery block, and cut fields are renamed *_preview so a partial read of head-truncated JSON fails loudly instead of looking like an answer.

Four nanny verbs (tools/delegate.py): delegate_start, delegate_wait, delegate_cancel, delegate_answer. There is deliberately no fake hurry: Claudexor truthfully exposes cancel and answers, not in-place steering. delegate_start takes an exact agent_session subagent_id (or recovery-only retry_of); API actor ids are refused. A scheduled configured nanny's bootstrap (subagent_bootstrap, mechanics in the §1 row) starts the exact snapshotted leaf before the first model round — after recovery adoption first, then the durable zero-run/unknown-evidence fences, because a fence may hide a live prior run and a fence-wake outranks every terminal — through the same delegate_start(prompt="") wrapper the model itself uses: one start path, one set of refusal shapes. delegate_start(subagent_id=..., prompt=..., root="skill_payload", bucket=..., skill_name=...) is the orthogonal exact-resource selector: the id chooses transport; the selector NAMES a resource, authorized through a fresh ResolvedResourceBinding for skill_payload.write — it never grants one. A DEFINITE start refusal (typed, no custody handle, closed _DEFINITE_UNRUN_REASONS set) ends the child UNRUN at $0; everything ambiguous wakes the model, because a false "spent nothing" terminal over a possibly-live run is the one direction classification must never fail toward. An unregistered project root is one such typed refusal: the nanny registers first, and a registration WE created is retired at settlement (delegate_registration_policy). The replacement and zero-run fences count the ACTOR's own runs: an unsettled run whose durable custody source is the review substrate is that substrate's obligation and never occupies this task's delegation slot.

Work orders. The compiler sends the entire chosen assignment, preserving context and instruction roles without an arbitrary host-size cutoff. Direct starts carry the normalized host contract once in instructions and the chosen assignment separately in prompt; coordination context remains complete. Real native/HTTP limits return their actual failure with the original input and any pending invocation retained. Exact-source readers remain optional capabilities. Legacy partial starts still verify their original renderer digest, selector and source intervals, and recover by replaying the recorded request body; removing new partial-start production never certifies an incomplete old run or launches a duplicate after ambiguous dispatch.

Supervision. delegate_wait is model-visible as an event-only sleep, not a caller-sized poll: delegate_supervision.supervised_wait renews bounded transport windows in host code at zero LLM calls, and only a meaningful event (settlement, interaction, fault, addressed message, child signal, control, recovery judgment) becomes a coalesced durable wake, replayed across worker interruption until acknowledged; deadline, ceiling, budget, and cancellation remain outer bounds. Every receipt and wake carries one host-rendered coordination_context (intent note, deadline remaining, root-tree spend, active descendants, remaining paid acceptance capacity) — facts for LLM judgment, never thresholds. A requested future inspection (checkpoint_after_sec) wakes once and is consumed by any earlier real event — no cadence, stall classifier, or hidden polling. An observation the transport could not complete at all (a socket failure, not a received one) is a quiet renewal carrying its typed reason, never a refusal that spends a model round on something the model cannot act on; the reason separates our own read bound expiring against a live daemon (observation_read_timeout, quiet and nothing more) from a socket that carried no answer (daemon_unreachable, the only half worth an owner line); the beat stays three seconds with no backoff, durable counter or outage latch, and deadline, ceiling, budget and cancellation still cut a long unobserved stretch. The outage reaches the owner exactly once per episode, with one recovery line, on the supervising task's own progress surface. The loop holds ONE handshaken gateway for its whole wait and re-establishes it after any tick that returned no observation. On a wake the nanny holds its full tool surface and the parent's captured model/effort. Nanny economics are structurally quiet: the burn baseline resets only on real acts of delegation, and coordination verbs never buy metered silence (nanny_pacing.py).

A run's question is the nanny's to answer. Supervision wakes immediately with typed status="waiting_on_user" on a NEW pendingInteractions entry instead of burning the engine's answer timeout in dead polling; answer keys are echoed verbatim into delegate_answer (custody-gated like cancel, relaying the engine's typed outcomes including subscription_window_exhausted), and delivered interaction ids are acknowledged only after transcript injection, so a question neither re-triggers a round nor disappears across recovery. A question above the nanny's authority escalates to the nearest live ancestor. The codex lane has no mid-run channel: a terminal with outcome_facts.reason=input_required is answered by a plain new delegate_start(subagent_id=..., prompt=...) — never the engine's rerun verb, which would start a run outside this task's custody trail.

Recovery is exact and cause-specific, not generic task resurrection. Only a proven non-signal worker crash (or a planned self-restart's typed handoff) lets a successor adopt the exact run or pending invocation, before any LLM call or new start; anything ambiguous returns typed recovery-required, never a duplicate mutator, and owner restart, panic, signals, deadlines, and cancellation are explicit no-resume causes (delegate_recovery.py). When no physical run exists and none can be started, the model may finish only with a typed zero-run receipt (verify_and_record) after custody proves nothing is open; a session actor's terminal is CLEAN only through its own physical leaf or that receipt — "completed direct child ⇒ clean" does not exist, and physical_leaf_not_started rides the terminal projection as an incomplete/unknown execution axis.

Configured dispatch is exact (subagent_runtime.resolve_configured_actor_dispatch, resolved at dispatch — the last moment route availability is current): a bad choice returns a typed reason, current alternatives, and any reset time with host_fallback: false; the host never ranks alternatives, converts session work to API work, or picks the first healthy row. The WHY is recorded so the next redesign does not undo it: selecting an agent_session row IS the parent's typed LLM decision that this work executes on the harness — the FLOOR the host hardcodes (truth, money, authorship), while topology, decomposition, and supervision judgment remain the model's CEILING (BIBLE P13/P5). At completion actual_substrate derives only from custody rows; a typed startup refusal never authorizes work on a different substrate — another session route or API fallback is an explicit LLM-selected start.

Route health. subagents.route_health is the one health reader for dispatcher, nanny, and review slots — and it does not guess admission: the doctor status describes only the DEFAULT credential store while real accounts live in the engine's credential-profile pool, so admission belongs to the engine, whose start POST answers an empty pool with its own typed refusal (credential_pool_exhausted + earliest reset) at $0. enabled is honored as route_disabled (the owner's switch, not an observation); the other structural refusals are route_not_in_capability_catalog, an access-profile mismatch, engine_rejects_delegated_marker, and positively proven quota exhaustion. A fully-used ratio needs a valid future reset before it can refuse a route; an incomplete reading delegates admission to the engine and never certifies available quota. Explicit active cooldowns still bind. Review slots inherit "the engine decides" — never a silent fallback onto metered API spend; on the auto lane a dead daemon keeps the native fallback with its visible marker, while "daemon alive, pool empty" is discovered at the engine and disclosed. Model-visible selection is a bounded semi-stable catalog (model_visible_subagent_catalog, serialized under ## Available subagents); invalid or disabled configuration projects nothing, and dispatch remains the authoritative live check.

Read-only and mutating session rows share one nanny transport. The only difference is the run shape, whose ONE owner is subagents.delegated_run_shape, asked one question — is this an acting child? — by tools/delegate._derive_authority (live ToolContext) and by resolve_configured_actor_dispatch (configured snapshot); a shape re-derived in two places drifts unsafely in exactly one branch.

task authority access mode isolation execution.delegated
acting subagent (valid write surface) workspace_write agent live true
Ordinary root with a validated external workspace or a selected Project room workspace_write agent live true
top-level task selecting an exact skill payload (root="skill_payload") workspace_write agent live true
anything else, including a fail-closed subagent readonly ask envelope (default) not sent

WHERE a mutating run's changes are destined is the second, separate record — the host-derived mutation authority (tools/delegate_integration._mutation_authority / _payload_mutation_authority, re-exported by tools/delegate), never model-supplied:

source target_root derivation capture_mode
acting_constraint the child's own task_constraint.write_root, required to equal the genuinely ACTIVE workspace root delegated_snapshot
external_workspace_root the root task's validated external workspace or selected Project room delegated_snapshot
skill_payload the exact payload the fresh skill_payload.write binding resolved delegated_snapshot
readonly ordinary active root (nothing to write) none

A payload target gets a standalone private Git snapshot (subagent_worktrees.provision_payload_snapshot): the live payload is never initialized as Git; the loader-visible inventory is committed as a registered synthetic baseline, and capture trusts nothing under the child-writable snapshot's .git. Disposition applies a live, index-free git apply under a whole-payload content-hash CAS (drift = typed conflict; identical content = idempotent applied). The post-apply outcome set is complete: a live loader hash equal to the recorded RESULT hash is the success; a hash equal to the recorded BASELINE hash with a non-empty touched set is a provable non-mutation that RESOLVES the apply intent (apply_no_op: no success, no disposition, no reconcile queued, retry lane open) independently of whether a result hash was ever recorded; anything else is the ambiguous mismatch, whose intent stays PENDING and whose reconcile marker IS queued because the payload did mutate. A successful apply queues the extension reconcile and the skill's review goes STALE pending fresh skill_preflight/skill_review. The apply_no_op arm carries no content residual: a run whose ONLY change is a mode flip is already refused at CAPTURE time as unreviewable_metadata_change, and any patch that did land content would move the whole-payload hash, so a live hash still equal to the baseline means nothing was written.

A mutating run executes in a PRIVATE EXECUTION SNAPSHOT, never in the shared tree (metered children keep sharing the tree). At delegate_start the host snapshots the target's real state — tracked + staged + eligible untracked, with the sensitive/credential veto DECIDED BEFORE ANYTHING IS HASHED: a blanket git add -A would write vetoed secrets, .env included, into the shared object database. The baseline commit is pinned by a refs/ouroboros/delegated/ ref and checked out as a detached worktree (provision_execution_snapshot); scope.root stays the stable authority target while execution.workspaceRoot names the snapshot. The typed snapshot binding is recorded durably on the custody rows BEFORE the POST; an explicit retry reproduces it exactly (a GC-collected snapshot is a typed execution_snapshot_missing refusal, never a re-mint). The run still runs live from the engine's view, which is why the scoped-HOME/delegated marker below applies unchanged.

At terminal, delegate_wait captures the run's diff against the baseline durably into the task's artifact store; NOTHING reaches the target automatically — the nanny explicitly applies or rejects through integrate_delegated_patch. Only the staging substrate differs between lanes: a GIT workspace target applies under the repo git lock — first PROVING no touched path drifted from baseline_sha (a plain git apply relocates hunks by offset; the touched-path set is read from git apply --numstat in BOTH directions because each direction names only the paths it writes) — then applying and STAGING, never committing; a skill-payload target applies live under the whole-payload CAS. A SKILL-PAYLOAD target captures through the payload adapter over a parent-owned trusted index (_write_payload_patch_artifacts: nothing under the child-writable snapshot's .git is trusted, the baseline comes from the host-owned snapshot registry, and the final loader inventory is staged from raw bytes) and applies LIVE into the non-Git payload — nothing is staged into any active root, no .git or index is created in the payload, and a successful apply invalidates the skill's review and queues the extension reconcile before the verdict artifact. The protected-path gate applies only when the target IS the Ouroboros body; a conflict (proven drift) is owned by the still-running nanny, with snapshot and patch persisting until explicit resolution or discard. Mutation rides an apply-intent protocol: a durable delegate_run_patch_apply_started row lands before any tree mutation, so on replay a pending intent without a disposition answers typed INTEGRATE_DELEGATED_APPLY_AMBIGUOUS, resolved only by explicit acknowledge_ambiguous=true. The provably non-mutating outcomes that instead RESOLVE the intent are a lock error, proven baseline drift, a failed apply, a verified revert, and a live payload hash equal to the baseline after a payload apply. artifacts.delegated_capture_read_target narrowly rebinds artifact_store READS for the owning task's own delegated_runs/ prefix, so a child-drive nanny can inspect the patch it must dispose without widened write authority. Its sibling delegate_shared.orphan_capture_read_target extends the same READ, still read-only and still confined to one capture directory, to the terminal-owner ORPHAN the disposition rule already authorizes: the actor that MAY apply a foreign capture may also read it, resolved by path shape first (so an ordinary refusal never replays the custody log) and confirmed by orphan_disposition_status before any path is returned. A read-only child stays in Claudexor's default envelope — one transport with one derived difference, not a second pipeline.

Terminal reconciliation captures only over PROVEN terminality, and patch_captured means "a usable artifact exists". When the owner task is gone, a settled mutating run's diff is captured through delegate_integration.capture_terminal_patch_for_drive — capture only, never an apply: the decision stays with an owner, and the obligation stays visible (delegate_custody.undisposed_patches) until the durable PATCH_DISPOSED row clears it. An obligation is held by the THING, not by the task that created it: the run's custody rows and the target itself carry it, so a terminal owner's run keeps its debt visible and disposable rather than locking the target forever. While the owner task is LIVE, only that identity may dispose its captured patch. Once that task is terminal, any live TOP-LEVEL task may dispose it. Apply requires the caller's active Git root or fresh payload binding to equal the run's recorded target, or — for an ORPHAN only — to CONTAIN it while both lie under the host-minted subagent-projects root (the swarm aggregator shape: a fan-out into <project>/contributions/<track> clones of the very tree the parent works in, which the host already checkpoint-commits), with every apply-path guard unchanged (recorded-target match, protected paths when the target IS the Ouroboros body, proven drift, the whole-payload CAS, staged-never-committed). Reject requires only the owner's proven terminality, because it releases a dead task's locks and snapshot without needing fresh target authority. The PATCH_DISPOSED row then records disposed_by_task_id, and the delegated_runs_unreconciled projection stays the evidence trail the sweep heals. An owner whose terminality cannot be proven keeps the lock. A run closed absent or left unreadable captures NOTHING — across the provisioning boundary the child may still be alive and writing, and an eager capture would freeze an incomplete patch and serve it forever; the snapshot stays preserved for capture-on-demand, and patch_captured is minted only over a ready manifest so a failed manifest leaves every retry point open.

The stored delegated_runs_unreconciled projection is healed only from the write side, at four seams. delegated_custody_unreconciled is the disclosure a task carries when it wrote its terminal result while one of its OWN delegated runs was neither applied nor rejected. The overlay is ADDITIVE (outcomes.custody_debt_axes): it merges an objective WARNING and sets the top-level reason code, and it never rewrites the derived execution, review, objective or artifacts axes — so the paid verdicts the task actually earned survive. A rail truncation code from BEST_EFFORT_REASON_CODES keeps the single Reason line, since a round-limit or budget-exhausted victim is more usefully labelled by its truncation. Such a task reads Done with warnings with Reason: delegated_custody_unreconciled rather than Failed; the debt list itself lives on delegated_runs_unreconciled plus the delegate_terminal_reconciliation envelope, and sweep refresh/backfill heals it once anyone disposes, while the Reason line stays historical. Two disclosed residuals: a provider-death terminal carrying custody debt still shows the custody code on its one Reason line while remaining Failed (the devtools-only infra codes are not in the runtime best-effort set), and a task with a CLEAN derived outcome plus custody debt no longer increments evolution_consecutive_failures, because a warning is not a failure. project_dialogue derives failed/degraded from axis STATUSES only and never reads objective.warning, so a custody-debt child reads as a plain success in project dialogue even though its own card and the Main/Project rows say Done with warnings. Readers serve the stored projection (projection-over-replay, no live custody join), so a run settled after its task's terminal write stays stale until: the periodic sweep refreshes tasks named in its own reconcile outcomes; the boot backfill (delegate_terminal.backfill_terminal_reconciliations, once per generation) re-audits stored terminal rows still carrying a disclosure — the generation-crossing residual; the cursor pass (delegate_terminal.refresh_recently_settled_terminals, a durable byte-offset cursor over the append-only custody log, bounded per tick) catches terminal-boundary settlements no outcome-driven refresh can reach; or a kill path clears a stale list. Every refresh is audit-only — it never rewrites reason_code or recomputes the frozen delegated_runs_* counters, so a healed row may honestly read unreconciled: [] beside settled: 0; current liveness lives in the delegate_terminal_reconciliation envelope, and patch debt survives every refresh as patch:<run_id>, never a blind clear.

Disclosed delegated-isolation residuals (deliberate): a run whose owner's terminality cannot be proven (its task result is missing, unreadable, or still an unreaped running row) keeps its target locked and cannot be disposed at all — there is no time-based release; a live top-level task with a different active root may reject and release another dead task's snapshot, and the disposition row records who did it; an orphan APPLY widens the same way only by containment — a recorded target strictly inside the disposer's active root is accepted when both are under the subagent-projects root, so an aggregator may adopt its dead children's clones, and the disposer's blast radius grows to exactly the host-minted descendants the checkpoint-commit already writes (every drift, protected-path and staging guard unchanged); an orphan disposed by a non-owner writes its verdict artifact and delegate_run_patch_verdict row under the DISPOSER's task while the capture directory and the PATCH_DISPOSED row keep the OWNER's task id — readers key verdicts by run_<rid> so nothing breaks, but the owner's acceptance packet will not list that verdict; a GC-lost snapshot or permanently failing capture discloses in a typed refusal that its obligation can never be satisfied; the baseline is worktree-primary; the git lock is task-drive scoped, so two nannies integrating into the SAME external tree can interleave apply+stage (the drift check makes the loser's apply a typed conflict); and a credential-shaped file the CHILD creates fails the whole patch rather than shipping a partial diff.

A mutating run asks for containment, reads back what it got, and DISCLOSES the gap instead of refusing the work. In place because Claudexor otherwise hands the harness the operator's real $HOME — which holds the daemon token, so a compromised child could start its own runs at any access level. Four facts, one mechanism. (1) The execution.delegated: true marker rides in the same record as isolation: live, built from delegated_run_shape in one place, so one cannot be sent without the other. (2) The version floor is about the SCHEMA: config.CLAUDEXOR_DELEGATED_MARKER_MIN_VERSION (3.3.0) is the oldest engine whose strict RunExecution accepts delegated; below it the start is a 400 and no run exists, and route_health gives dispatcher and nanny the identical typed blocker before a token is spent. Read-only delegation sends no marker and keeps the 3.2.0 transport floor; an engine between the two floors serves read-only and refuses mutating. The floor is a schema fact, never a containment fact — the OS boundary is platform-dependent while a build declares one version everywhere; threat model and measured bands: docs/DELEGATED_ADMISSION.md. (3) What was APPLIED is asked of the attempt, never of the OS: delegate_wait reads harness_home_isolated, confinement_mechanism, and the proven denied path from the attempt record (gateways.claudexor.attempt_containment); sys.platform appears nowhere in the decision. (4) A missing boundary is disclosed (durable delegate_run_unconfined event, the child's instructions, the parent's terminal payload), not refused — the child already holds a shell in this worktree, and cutting the lane on every boundary-less host costs more than the marginal step it prevents. A recorded FALSE is still a fault: harness_home_isolated: false, or a scoped home that IS the operator's own, cancels as a typed containment fault, exactly like a widened access profile — those two exact facts are the WHOLE breach rule; a MISSING home fact is neither breach nor proof, so unproven is REPORTED.

The model cannot widen its explicitly bounded delegated authority. delegate_start exposes prompt, exact session subagent_id, max_seconds, recovery-only retry_of, and the exact-resource selector; there is no access, mode, isolation, or scope argument, a retry cannot change route, root, access, or permissions, and every run carries host-authored instructions the model can neither widen nor forge. Because Claudexor DERIVES effective access rather than echoing the request, delegate_wait verifies rather than assumes: every fetched run detail goes through _containment_breach, the one reader for both halves of containment (access profile and harness HOME) because they fail identically — a verification written for one half leaves the other trusting an echo. A run enforced WIDER than the task is entitled to is cancelled as a typed access_profile_widened refusal; narrower is fine.

Read provenance on the accounts surface. GET /api/claudexor/status carries a reads block (ClaudexorStatusReads: catalog/accounts/quota, each ok|not_read|failed) because the owned daemon starts lazily: an idle daemon serves empty collections under a 200, and "no account connected" must not be inferred from a collection that was never read — ok makes the matching collection authoritative (empty means empty); one parity-tested client reader (facetReadState) applies the same rule, and the aggregate daemon.state is never the negative answer. Login jobs: the daemon stays the sole process/fence authority and reconcile is an explicit POST, never passive polling (gateway/claudexor_accounts.py).

Git and commit review

review_evidence.capture_commit_review_evidence freezes selected browser/vision calls and same-round automatic image-attachment observations, including explicit unavailable-image gaps, after cheap/free admission and before preflight/triad/scope. The borrowed loop trace and original call refs retain exact redacted arguments/results and the immediately following visible response in the same execution, excluding provider thinking. The original context binding is restored through the attempt's existing _LoopExitContext cleanup even if tools._ctx changes while the loop runs. Adjacency never attests visual inspection. One canonical task source handle holds the selected UTF-8 view; its initial exhibit stays within _ACCEPT_NOTES_CAP with counts, completeness and source identity. Native reviewers read artifact-store ranges under the real canonical root; sessions receive byte-identical bytes in ignored .review-drive/<task>/<sha>.txt; packet-only reviewers receive a bounded, explicitly partial view when needed. Existing request evidence/refs and preflight execution evidence_source_ref bind it. Pending rejoin restores the same source, including a recorded empty selection, and carries it from preflight into triad/scope without selecting new trace. Commit cleanup waits for physical custody; standalone or uncertain views remain retained without a new cleanup registry.

Commit preparation verifies the exact local working-branch ref before an unambiguous checkout. A missing ref refuses without changing the current branch, index or files; remote guessing and implicit branch creation are disabled. Detached work retains the existing checkout -B <branch> HEAD recovery, and managed assisted merges retain transaction-owned precommit verification.

The hermetic runner alone mints ctx._preflight_test_proof after actual green lanes and containment. It binds the assembled checkout tree and installed index, effective pass specs, scrubbed environment, and invoked/resolved interpreter and active Node identity (a Python symlink's invocation path selects its venv). Equivalent ordinary and managed checks reuse that process-held proof only after their distinct baseline checks; a phase label alone cannot establish equivalence. Every workload binds HEAD because even an unmarked default-lane test may inspect committed files/history. Repeated preflights can reuse an unchanged subject; creating a commit changes HEAD and requires a new run. A skip or no-suite None is not a green proof. A restart loses the proof and reruns the suite. Creation and reuse emit preflight_test_proof observations through the existing event log, naming the task, HEAD, tree, index and workload fingerprint. An unavailable log falls back to an explicit diagnostic, never a test failure or a new authority source. Managed durable tests_evidence is forensics of the tested tree, never reuse authority; neither event logs nor the transaction can mint a reusable proof.

tools/git.py owns repository writes, staging, reviewed commit, rollback or restore, tags, push, and CI follow-up. File-edit tools validate their own atomic write shape; mutation_attribution.py captures the root-task baseline and projects only the clean-at-baseline system-repository delta — a changed pre-existing dirty path, stale or missing baseline, or failed scan blocks automatic staging. commit_reviewed(paths=None) stages only that attributed candidate, explicit paths must be a subset, and an empty candidate returns GIT_NO_ATTRIBUTED_CHANGES; managed update transactions keep their separate typed whole-tree authority.

A reviewed commit is bound to one staged fingerprint. A cheap LLM-first advisory pass may run before the expensive gates; it is advisory, and skipping it never skips independently applicable tests, triad, applicable scope review, aggregation, or exact-SHA binding. The hermetic preflight runs the candidate in a disposable worktree and data root; triad and scope inspect the same staged snapshot, aggregation preserves actor evidence and obligations, and any mutation stales the binding. Managed exception: a managed-update resolution commit reviews the declared M0→S subject (tools/review_subject.py) and the commit gate binds S to the exact index write-tree the fingerprint pins. External review wrappers report readiness but do not grant commit authority.

The exact binding includes the git write-tree SHA, ordered HEAD and MERGE_HEAD parents, indexed VERSION, expected v{VERSION} tag, any existing tag target, and the binary staged-diff hash; after commit, tree, parents, VERSION, and tag target are re-read before success or push is recorded, and an existing release tag is never silently accepted or retargeted. Durable review state keeps attempts, obligations, readiness debt, raw actor evidence, and the final commit or tag binding. Raw advisory output lives in state/advisory_review.json (selected by snapshot_hash/ts), not in review_status(include_raw=true), which exposes raw triad and scope attempt evidence. BIBLE supplies review authority, CHECKLISTS supplies criteria, and Development supplies the procedure; snapshot identity, advisory coverage or audited-skip evidence, deterministic results, actor evidence, and final Git identity must all describe the same material.

Review stack

review_cycles.py owns the shared paid-cycle ceiling OUROBOROS_REVIEW_MAX_CYCLES (a positive integer or unlimited; anything else fails closed to the shipped default). One number, four gate meanings: paid reviewer-panel cycles per task for plan review; paid panel runs per task for acceptance, where improvement passes are cycles − 1; paid triad-plus-scope cycles per ROOT task for the commit gate, counted from the attempt ledger at DISPATCH so the ceiling counts money and only undispatched attempts stay outside it; paid panel dispatches for skill review. Byte-identical material replays each gate's recorded verdict for free under that gate's identity.

Review waiting has six independent axes:

axis bound
transport dead-socket read bound
operation typed active-operation lease
logical slot/task deadline
policy budget / cancellation
absolute task ceiling
aftermath late-result custody

review_custody.py is the small worker-lifecycle seam used by review_substrate.py; it does not schedule tasks or create a second timing ledger. Parallel slots hold independent operation ids; the active-operation map only prevents a live physical call from being mistaken for idle — deadline, budget, cancellation, or ceiling still wins. Delegated-session expiry uses the verified cancel path; API/thread calls disclose in_flight and reconcile a late answer before the same retry identity can dispatch again, and a settled terminal API error is retained in the cycle's replayable actor roster so a sibling cannot erase its terminal fact or buy a duplicate physical call. Physical custody is proved by the capture/operation state, never by the synthetic operation id: usage_accounting owns the state vocabulary; only reserved/released proves pre-dispatch, a positive capture outranks a contradictory synthetic label, and dispatched/unresolved without a typed terminal HTTP status stays custody-lost/no-resend. Custody does not infer pre-dispatch provenance from Python's implicit __context__ — a fallback raised inside a prior provider handler inherits that earlier attempt — so only an explicit __cause__ or typed transport metadata can release a paid row. A retry token without a durable invocation is custody-lost for every delegated review surface, and a durable token is valid only for its recorded surface, slot, and operation. Retry custody uses an explicit material/cycle identity when supplied — mutable prompt or history is deliberately not part of it, while a changed snapshot, owner intent, reviewer route, or admitted cycle mints a new one.

Timeout composition: send-time VLM captioning keeps its direct 90-second provider cap; Anthropic's direct route keeps its 120-second provider default; neither is a generic review-reasoning cutoff, and the whole hierarchy is narrowed inside the owner deadline and finalization reserve before dispatch. A returned provider response or typed terminal error is settled even when its body is empty, so bounded repair/retry may apply; a dead socket after dispatch is provider_outcome_unknown and cannot trigger another paid route. A spent owner window yields a typed $0 not_dispatched row before fan-out; under blocking enforcement an in-flight triad row remains pending rather than becoming a final quorum verdict, and a primary call at the deadline boundary enters the local finalization rail (deadline_local), not a provider-outage relabel. A reviewed commit has no independent outer tool cutoff: the foreground caller retains custody until settlement.

Before either parallel surface starts, one locked write records paid=True plus both complete slot rosters and operation ids in the commit-attempt row; a delegated slot patches its reserved row with pending_invocation_id before the provider POST. Exact resume preserves rows and tokens; a missing or mismatched operation is custody_lost under every enforcement mode. The paid stamp records the process-custody server session and pid owning the reviewer threads: owner-loss is proven by pid death, never elapsed time — a TTL would convert waiting into resend authority, and starting another Agent is not evidence this owner died; tokenless rows settle as typed infrastructure failure only after a death seam confirms that exact pid gone, and a legacy row without owner identity stays fail-closed. Plan review re-enters a recorded in-flight cycle only through exact live custody or complete operation-addressed CAS for its recorded dispatched rows; a delegated poll that loses transport after a run id exists preserves the exact durable invocation for retry custody rather than cancelling a healthy unknown run.

The review surfaces:

  • Advisory pre-review (claude_advisory_review.py) — a cheap, staleness-aware error-finding pass; the audited skip covers only advisory admission, never authoritative review, test policy, or snapshot binding.
  • Triad diff review (tools/review.py) — configured reviewer slots cover the Repo Commit Checklist with JSON findings under config.adaptive_quorum (the same SSOT as scope/plan/skill/acceptance review).
  • Scope review (tools/scope_review.py) — touched context plus a Generated Scope Atlas, checking intent/scope/coupling; fail-closed on unreadable touched files, enforcement per OUROBOROS_REVIEW_ENFORCEMENT.
  • Parallel orchestration (tools/parallel_review.py) — two-phase admission: BOTH gate packets are assembled and fit-checked before any reviewer is dispatched, so a deterministic assembly block anywhere dispatches nothing and spends $0 everywhere — paying one lane while its sibling was doomed wastes real money and produces an unusable half-verdict. Money admission (tools/review_admission.py: commit_gate_paid_seats / admit_commit_gate_wave, called by the orchestrator between assembly and dispatch) is all-or-nothing and scope-first: every PAID seat of the wave — the blocking scope seats FIRST, then the triad — is priced with the reservation math reserve_attempt will apply (a packet row's exact cached-block message pair, a configured-subagent api row's exact first native send — instructions, work-order and tool schemas, its later rounds reserve themselves — and each seat's own output reservation, each priced under the usage scope its substrate sends under (review_usage_category + slot, so a warm split of the caller's transcript never stands in for a reviewer seat's cold prefix) against every fence reserve_attempt enforces — the global TOTAL_BUDGET remainder, resolved like the ledger's own global fence and unbounded when the owner set no finite global budget, and the task's CURRENT root fence — the tighter remainder binding) and admitted as ONE wave through review_wave_budget_gate(surface="commit_gate") before any seat spends; a wave that does not fit is a typed $0 not_dispatched record on every seat plus one review_wave_budget_insufficient event carrying binding_axis (global or root) and both remainders, and the block names the binding fence (global budget vs per-task fence), accounted on that axis, the part of it reserved by other in-flight attempts, the wave's bound, the shortfall, what the other axis alone would leave and the knob that moves the binding fence — not a half-dispatched panel whose non-blocking seats hold the money the blocking reviewer needs. Admission is a read-only pre-check without a wave-level hold: the per-seat reservation stays the enforcement, so money a concurrent task takes on the global axis between admission and a seat's reservation can still refuse that seat (a disclosed residual, not closed by a wave-level ledger hold). A paid seat is any api row (packet or native episode); an agent-session row rides the owner's subscription (its ledger row is written at settlement, never reserved) and is neither priced nor waited for. Admission and reservation share the same fences: every executor transition on the way to a seat (the orchestrator's two-worker pool, the per-row scope pool, the triad's loop thread and its run_in_executor hop) runs under contextvars.copy_context(), so the usage scope the wave was admitted with — its bound root fence included — is the scope the substrate's reviewer rows reserve under; a settings reload that changes the per-task limit mid-turn changes neither side, never an admitted-against-$50 wave whose seats reserve against $8. An admission that raises is fail-open like the gate's own unknowns, but loud and typed (review_wave_admission_unavailable carrying the error; the wave dispatches unadmitted and without the hold). An admitted wave submits the scope seat first and starts the triad only once the scope seat's OWN reservation is on the ledger (its category + slot + task identity among the reservations reserve_attempt appended, read from the ledger's root telemetry — a refresh, a settlement, a refused reservation or a sibling task's same-named scope slot never releases the hold), bounded by NESTED_SETTLEMENT_MARGIN_SEC and by the scope seat itself; a hold that ends without observing it emits a typed review_scope_lead_unobserved event. Transport and logical timeout contracts are untouched by this admission.
  • Shared helpers (review_helpers.py, triad_review.py) — pack building, checklist loading, JSON extraction, usage events, obligations scaffolding, reviewer actor records.

A free refusal never wears the form of a verdict: a refusal that spent nothing is recorded as a typed not_dispatched fact plus its reason, never as a DEGRADED panel, a synthetic actor, or a verdict. That is the one shape across every $0 exit — a plan-review locator the evidence policy cannot attach is a named omission row the panel is dispatched with; an acceptance packet that overflows takes the ladder rather than reporting a verdict; a truncated or self-pageable row names its cut; and a request that was never sent is one _error_actor(..., operation_state='not_dispatched') seat that stays in the denominator.

Rationale: diff reviewers catch line-level mistakes; the scope reviewer catches cross-module contracts; running both on the same staged snapshot prevents one result from hiding the other. The managed-update exception reviews the same declared M0→S subject in both lanes. Task acceptance is a separate root-owned post-delivery system specified once under Task lifecycle; it shares adaptive_quorum and these transports, has no acceptance scope actor, and stores its verdict separately from the terminal result.

Structural smoke gates are a deterministic BIBLE P3 codebase-size component. ouroboros/review.py::iter_gated_modules is the one source inventory for smoke, codebase_health, census, and the UTF-8 byte gate (Python everywhere plus first-party web/**/*.js, vendored/minified excluded). ouroboros/size_ratchet_manifest.py is a generated, data-only debt register consumed through AST literals: exact module debt above 1600 lines, exact function debt above 300 lines, the 1001–1500 band with rationale authority, and exact byte debt above 200,000 UTF-8 bytes. validate_size_ratchet proves the live and staged manifests exact against their trees and shrink-only against the merge-aware committed authority — no first-parent history replay, so a fork whose local line predates the manifest is never condemned by inherited topology. Official-repository CI runs the blocking size_ratchet pytest lane while every local surface reports the same findings as warnings; two disclosed residuals — pairwise validation covers only the base→HEAD interval, and the official block presupposes branch protection. Within a validated pair, debt can shrink but cannot be swapped, re-entered, grow on the byte axis, or survive stale. MAX_TOTAL_FUNCTIONS (ouroboros/review.py) remains the coarse runtime ceiling. A deterministic hot-store growth invariant sits beside these gates: agent_startup_checks.py::hot_store_growth_notes (surfaced by context_health.py::build_health_invariants and once per worker boot) stats nine hot stores plus the archive/chat_*.jsonl aggregate — state/consciousness_observations.jsonl, state/usage_attempts.jsonl, logs/events.jsonl, logs/tools.jsonl, logs/supervisor.jsonl, logs/task_reflections.jsonl, logs/progress.jsonl, state/scheduled_tasks.json, and state/skill_review_root_tasks.jsonl — against justified byte thresholds in ouroboros/context_budget.py and emits a WARNING with a remediation pointer.

Three gen/verify inventories ride the same discipline as the size manifest (generator scripts/regenerate_inventories.py, verify tests/test_generated_inventories.py, staleness = red): the frozen-contracts inventory (docs/v7next/FROZEN_CONTRACTS_INVENTORY.md, a machine extraction of §11.1 with every owner/anchor path resolved against the tree and the ouroboros/contracts/ package-coverage gap pinned), the data-layout inventory (docs/v7next/DATA_LAYOUT_INVENTORY.md, every entry of the §1 "Data layout" tree probed as a tracked repo path or a runtime-source literal, so a durable file renamed in code while its tree row survives turns red), and the facade inventory (docs/v7next/FACADE_INVENTORY.md, the AST-derived noqa: F401 re-export surface with per-leaf domains from ouroboros/domains.toml).

The shared hard prompt-size SSOT is REVIEW_PROMPT_TOKEN_BUDGET = 920_000 in review_helpers.py. review_context_atlas.py targets 850K estimated prompt tokens for scope review and deep self-review, and the final 920K gate stays in each caller as the hard stop (plan review builds no Atlas: its packet is sized per slot by plan_review_runtime.plan_slot_fit). Scope review additionally reserves output headroom inside the reviewer's window (_SCOPE_MAX_TOKENS 100K plus a tokenizer margin), because provider accounting can exceed the local chars/4 estimator on atlas-heavy prompts and a physically rejected oversize prompt has no authoritative verdict; scope_review.py gates assembled input on _SCOPE_INPUT_TOKEN_LIMIT = min(920K, window − _SCOPE_MAX_TOKENS − margin).

The cap is DENSITY-CALIBRATED from measured evidence, not model names: usage_accounting.execute_physical_attempt records (prompt_chars, real prompt_tokens, route_fp) witnesses after settlement in the token_density namespace of the canonical capability_evidence.json. The Main reducer uses the newest fresh exact-route witness, then newest exact-model witness, then neutral 1.0. The review reducer uses the densest still-fresh exact-model witness as authoritative — a fresh exact-model witness measures THIS model's real tokenizer and may undercut the conservative 1.65 cold floor, which otherwise shrinks a 1M-window reviewer so far that the managed scope atlas can never assemble — while the floor keeps governing stale, absent, or cross-model evidence (a witness stays fresh for 90 days — _TOKEN_DENSITY_TTL_SEC — because a tokenizer does not drift week to week, and an idle install must not fall back to the floor it could assemble above). Retention keeps the densest witnessed pair plus the newest bounded remainder, so lighter traffic cannot evict its support. review_helpers.calibrated_input_token_limit returns the strictest of the 920K cap, the density form, and the historical absolute-margin form; scope_review._effective_scope_input_limit computes it per call (an import-time constant would freeze the pre-measurement value), and the triad, plan review (through review_synthesis.per_slot_input_token_limits), the task-acceptance packet ceiling, and deep_self_review._run_packed_review (the packed delivery) consume the same helper. The floor has one structural loop: a pack refused BEFORE any send can never record the exact-model witness that would have admitted it. ONE rung breaks it, shared by the packed deep self-review and the commit gate (the gate runs the SAME rung, not a copy) — capability_evidence.cold_start_density_probe, beside the witness store it feeds: when a pack would be refused or degraded for size and the route has no fresh exact-model witness, one bounded send on the exact model (a slice of the real pack, DENSITY_PROBE_MAX_TOKENS output tokens, through chat_observed under physical_attempt_limit(1) so the transport ladder cannot redial it) records the witness, the cap is recomputed and the pack rebuilt ONCE; never on a warm store, never retried, never on a commit whose pack fits. At the gate (review_admission.density_probe_before_size_refusal: the scope ladder's size terminals and degradation rungs, and every triad api slot the irreducible prompt overflows) each attempted probe is a progress line, one review_density_probe review event and, for scope, a density_probe ladder step; a probe the paid ledger refuses (BudgetExceeded) is the typed budget_refused disclosure and the existing size refusal proceeds unchanged.

The cap is WINDOW-AWARE. reviewer_window.resolve_reviewer_window resolves each slot's REAL window from Capability Evidence (no static table); the triad, plan review, and task acceptance size their packs against it — a 200K reviewer treated as 1M-capable loses its whole review to a deterministic prompt-too-long 400 — with a sub-1M window scaling its reserves rather than zeroing the slot; deep self-review's PACKED delivery reads the same resolver but REFUSES a confirmed sub-1M route typed instead of sizing down (its pack is the ≥1M guarantee; the native/session rows serve such routes). An UNKNOWN route keeps the full-window assumption on those four surfaces (triad, plan review, task acceptance, and deep self-review) — for deep self-review, its packed delivery; only scope review fails CLOSED on absent evidence, because its BLOCKING authority is what a wrong assumption would forge — guess-small is not the safe direction, since governance packs run ~169K tokens and a sub-floor guess would decline the whole review on every cold-evidence install. blocking_authority_allowed is a COMPUTED property of fresh evidence (capability_evidence.confirms_at_least(..., require_fresh=True)), never of the configured model name: that closes the two routes to a forged blocking verdict — an expired record read as live, and the designated-default sentinel read as sourced (the sentinel survives as a sizing number only). Concurrent resolutions of one route serialise on a per-route lock; the probe is rate-limited by the evidence TTL and nothing else — a memo that never expires while the record does converts a healthy long-lived process into a permanently blocked one.

Scope applicability follows the owner's context mode. In max, provider oversize and window-authority failures remain failed evidence: they never satisfy the sourced blocking-review floor or become PASS. Under owner-selected advisory enforcement, the commit aggregate may continue with that failure disclosed on an independently bound candidate; blocking enforcement still refuses. Budget/admission, unresolved custody, Stop/deadline, ownership and candidate-binding failures remain independent refusals under both modes. OUROBOROS_SCOPE_REVIEW_FLOOR is retired and cannot change this policy. A structurally oversized repository calls for a smaller reviewed tree, never a weaker blocking reviewer (BIBLE P3). In owner-selected low, run_scope_review reads config.get_owner_context_mode() and returns before assembly with the typed non-blocking status="skipped_low_context_mode" row, distinct from an attempted review that failed. In max, a known sub-1M packet reviewer supplies advisory evidence only; an owner-declared retrieving scope slot can instead establish blocking authority at ≥200K sourced evidence. Non-responded actors retain their provider failure text so the cause stays visible beside the permission decision.

Scope prompt assembly is GUARANTEED-FIT: the assembler walks a deterministic degradation ladder instead of skipping. The five rungs:

  1. Full atlas.
  2. Compact atlas — durable context_manifest keeps full per-file coverage; the visible prompt keeps a compact path/disposition index.
  3. A REQUIRED file the atlas cannot fit is a failure to ASSEMBLE, never a smaller pack (BIBLE P3): the row records budget_omitted, the pack becomes required_artifact_omitted, and the ladder keeps shrinking the fixed part and retries.
  4. Touched files degrade to diff-only, freely degradable first — legal ONLY for merely-touched files, whose complete change evidence IS the staged diff; an artifact owed in full regardless of the change (prompts/, ouroboros/contracts/, protected runtime, canonical docs) declared diff-only is the same typed assembly failure as rung 3. Degraded paths are DECLARED to the atlas so the durable coverage row matches the prompt.
  5. Unchanged hunk context may be removed with -U0, preserving every file/hunk identity and +/- line.

The triad independently applies the same one-pass fit rule before dispatch (touched-path manifest replacing diff-duplicated snapshots, then -U0). On a cold density store both ladders are preceded by the shared cold-start density probe (above): the scope assembler is re-run once under the recalibrated cap after a recorded witness, the triad fit re-sizes its slots once. Before any rung, all three review packs — advisory, triad and scope — apply the same disclosed span-only release-carrier cut (review_file_pack.span_only_release_carriers) over the pair each one reviews; the scope pack declares the cut carriers diff-only to the atlas with that reason and records them as the ladder's first entry. Every step is a disclosed omission (BIBLE P1), never silent. TWO exhausted-ladder terminals remain and both fail CLOSED: the irreducible prompt not fitting, and a required artifact that never assembled. ONE predicate classifies both — review_context_atlas.atlas_assembly_failed over ATLAS_ASSEMBLY_FAILURE_STATUSES — instead of each consumer re-deriving a status test; the terminal STATUS picks the authority branch while the CAUSE travels beside it, a mixed state reports BOTH causes, and scope review and deep self-review are strict consumers that never review the remainder (per-rung casuistry: review_context_atlas.py/scope_review.py docstrings). Plan review is not an Atlas consumer: it reviews a typed SPEC with agent-declared evidence, and a packet that cannot carry the constitutional pack a self-modification plan requires is a typed failure, never a silent reduction.

Planning, deep review, reflection, memory

Plan review, task acceptance, commit review, and deep self-review answer different questions and never inherit one another's authority: planning judges a proposed approach before implementation, task acceptance the delivered objective, commit review a staged self-change, deep self-review the whole system. Post-task reflection and memory persistence learn from execution but approve none of those boundaries.

Plan construction and review

plan_task reviews an INTENTION before the work starts — the same organ for code, research, deliverables, and actions (BIBLE P3). The submitted envelope carries the goal, the plan prose, and a typed domain-neutral SPEC (in_scope, non_goals, acceptance_claims, invariants, decisions with rejected alternatives, deferred, affected_resources, evidence); ouroboros/tools/plan_spec.py validates and normalizes its full operative content without shortening strings or dropping excess items, then mints the ids that are the only valid breaks targets (goal, claim_N, invariant_N, decision_N, deferred_N) — host-minted positionally, because a caller-chosen id could shadow another target and corrupt what a blocking finding breaks — and hashes it. Governance documents always come from the system repository; declared targets and evidence resolve against active_repo_dir_for(ctx), and a path escaping the active subject or an unreadable root is a named omission, never a silent gap.

ONE structural fact tiers the governance pack: constitutional is true iff a declared affected_resources/evidence PATH locator resolves under the Ouroboros system repository (an evidence path only if it exists). A constitutional plan carries BIBLE.md and ARCHITECTURE.md in full (an api_chat row inline; a retrieving agent_session row as mandatory full reads) — the constitutional packet is not tiered, so assembling it without either is a typed failure, never a disclosure — and every other plan carries the runtime heading-derived navigation maps (context_layout.generate_doc_nav_map, never a copy) plus resolvable pointers. There is no plan-kind taxonomy, no agent-declared plan_class, no planning scouts, and no plan Atlas.

Declared evidence is resolved by ouroboros/tools/plan_evidence.py against exactly two allowed roots — the active workspace and the system repository — with the shared sensitive-name policy applied to locator and target; every refused, missing, truncated, oversized, binary, or URL locator becomes a typed omission row in the manifest (the host never fetches a URL), and a need_evidence locator a reviewer names is attached by the host on the next cycle through the same policy. What that policy cannot attach is a named [reviewer-requested] omission row the panel is dispatched with and judges with, never a reason to run no reviewer; the evidence continuation uses a fresh full-packet dispatch only when no exact artifact reference exists, disclosed per slot as a capability_delta. An unreadable referenced artifact instead fails closed with plan_review_exact_artifact_unavailable and never mints replacement authority. A locator may carry an exact range — ::lines=A-B, ::bytes=A-B, ::tail=N, or ::symbol=Name for .py sources — and a source above the per-item byte bound is attached head-first with the cut named. The manifest hash joins the spec hash and constitutional in the wave fingerprint, so changing what the reviewers can see changes the identity of the review.

ouroboros/tools/plan_packet.py keeps the current objective, spec with ids and plan prose complete; evidence, exploration and prior-cycle views retain their separate disclosed bounds. Tail-only requirement changes alter canonical plan identity, and actual per-slot fit refuses an unfit packet rather than removing current requirements. Slots come from reviewer_slot_config over the existing review_execution._review_route_executor seam; a retrieving reviewer is recorded host_file_read_attestation: unobserved. Reviewers return ONLY a typed findings array (blocking with a breaks id · note · need_evidence with a locator); the HOST validates membership, demotes invalid findings with disclosure, retains valid repeated evidence requests without growing the seen-locator set, keeps failed slots in the quorum denominator, and computes the aggregate through config.adaptive_quorum. No reviewer emits GREEN as authority. Useful premise criticism and simpler alternatives are encouraged as optional brainstorming notes, with adoption left to Ouroboros; there is no mandatory competing plan or quota, and repetition alone cannot promote advice to a blocker.

plan_review_state v2 inside the root task result is the bounded durable index: each wave records a hash-bound task_source reference to the full operative spec, the hashes, constitutional, validated findings, aggregate, dispositions and paid state. Current authority readers restore the full spec before comparing plans or binding its acceptance_claims through contracts/task_contract.effective_acceptance_claims; historical readers resolve their selected wave. Raw operative text uses the existing source-handle store, while review evidence remains redacted. The evidence manifest remains in the full wave artifact instead of being duplicated in the bounded index. Compaction and child promotion retain both references; an unavailable recorded source is a typed infrastructure failure, never empty claims or a new unpaid cycle. Legacy inline specs remain readable; older waves compact with an explicit omitted count. A paid actor still physically in flight keeps the wave open as DEGRADED with review_late_result_pending even when settled rows meet the arithmetic quorum, so a late blocking result cannot arrive after a false GREEN. A v1 record is read-only (legacy_v1_projection; an OPEN v1 wave projects legacy_open_requires_resubmission, never auto-closed).

Closure follows the finding class through plan_spec.closure_after_disposition at initial synthesis and later dispositions: GREEN and note-only REVIEW_REQUIRED close immediately in either enforcement mode; outstanding need_evidence closes through a disposition-only plan_task call (no model call, no cost), while a below-quorum blocking finding stays open. REVISE_PLAN can never be closed by disposition; a subsequent paid delta review may evaluate a changed spec or justified rejection when another cycle is available. Paid cycles are bounded by the shared OUROBOROS_REVIEW_MAX_CYCLES; an identical envelope replays the recorded wave for free, except that an open wave whose blocking findings all carry valid reject dispositions may dispatch exactly one subsequent paid delta panel when another cycle is available. A wave is paid iff at least one reviewer slot was physically dispatched; only a nothing-dispatched wave of typed $0 skip rows stays unpaid and never replaces a paid predecessor. Before fan-out the engine captures one panel health snapshot (subagents.route_health, route-level evidence): a slot with positive structural evidence of a spent lane becomes a $0 typed skip row that stays in the denominator; unknown health dispatches (fail-open). The wave records the health epoch and reviewer-roster fingerprint, and a recorded DEGRADED wave replays free only under an identical envelope, matching epoch, and unchanged roster — otherwise another paid panel requires remaining cycle capacity (replay/epoch casuistry: plan_review.py docstrings). When the wave's own typed rows prove the quorum structurally unreachable, the wave carries quorum_unreachable plus the earliest recorded reset, and under blocking enforcement the finalization gate RELEASES while the review stays open: the agent may finalize blocked_with_evidence (reason plan_review_quorum_unreachable), wait through a one-shot schedule_followup, or ask the owner — the host adds facts only, never an answer template. Under blocking enforcement an open wave otherwise holds implementation and an exhausted cap escalates with the typed review_cycles_exhausted reason; every banner, closure-note view and next-step description respects that unavailable paid continuation, while free disposition and exact pending custody keep their existing rules; under advisory the agent may proceed with the wave open (one typed owner-visible plan_review_advisory_open event plus the loud disclosure at finalization). Unavailability, invalid state, budget refusal, and deadline rails remain typed non-authoritative attempts, never substitutes for GREEN. There is therefore no pre-dispatch refusal for a reviewer request the host cannot attach: the only $0 exits are the typed attempts named here and the typed not_dispatched slot rows.

Closed note-only waves still accept voluntary review_disposition annotations through the existing writer. The spec/verdict and paid-cycle count stay fixed; the replacement artifact retains its immutable predecessor. This is optional reasoning history, not a new plan gate.

Deep self-review

The retained report is historical evidence, not a current-tree verdict. New reports record their actual generated_at time in the existing provenance header; source_revision=unknown is explicit because these deliveries do not capture one frozen reviewed Git revision. Neither a live HEAD nor file mtime can fill that gap. Context retains the existing 8,000-character disclosed excerpt, labels unknown legacy provenance, and leaves the stored report intact.

A task with type=deep_self_review bypasses the ordinary tool loop and calls deep_self_review.run_deep_self_review once, on the configured deep_review reviewer row (reviewer_slot_config.deep_review_slot; absent, the packed api row synthesized from OUROBOROS_MODEL_DEEP_SELF_REVIEW, so an install that never saved a row keeps the historical delivery). Like every review surface it has three deliveries, chosen by the row's retrieves predicate and executed through the same seam (review_execution._review_route_executor — the advisory precedent):

  • Packed (a direct api_chat row): the caller resolves the model's payable spelling and its observed route window before building the pack; ReviewContextAtlas supplies repository coverage and the memory whitelist remains full (identity, scratchpad, registry, WORLD, full knowledge index, patterns, improvement backlog); the prompt's omission section is bounded and points to the complete persisted coverage manifest; a required Atlas artifact that cannot assemble is an explicit failure — the flow rebuilds only through its final-fit path and never sends a silently incomplete review. One chat_observed call with tools=None, byte-identical on the wire to the pre-row review (golden digest test).
  • Native inspection episode (a configured-subagent api_chat row): NativeToolRoundReviewExecutor over the repository root with policy["native_data_root"] = the real runtime root, so the reviewer reads the repository with the host's read-only tools (the runtime root is its readable data plane) while the memory whitelist reaches it inline byte-exact. The route-owned task carries the role prompt, up to seven memory files inline byte-exact with every whitelisted entry's disposition stated (inlined / missing / empty / oversized / read_error — memory coverage is by inlining and disclosure, never by receipts), BIBLE.md as a MANDATORY full read (with its size, for chunked reads) and ARCHITECTURE/DEVELOPMENT/CHECKLISTS as navigation maps read on demand; the output contract is the report shape (free markdown, most critical first, a one-line model-side coverage header). Bounds are the shared ones: the window-derived transcript bound with the landing notice, the owner deadline together with the slot's logical window (the task's absolute ceiling narrowed by the deadline), and the paid ledger; exhaustion delivers the collected draft marked native_incomplete. After the episode the host derives BIBLE.md coverage from the executed repository-root read_file receipts, matched on the reader's opened_path and opened_root (never on the model's spelling; a data-plane read never counts): read only when the merged extents of measured receipts cover the whole file (a single result is capped, so a full read is multi-chunk by construction), otherwise partial(fraction), missing (nothing of the file delivered at a repository root) or unobserved (receipts capped below the call count, or a matching receipt without an extent — full coverage is proven by measured receipts alone); the edge rules (the raw-spelling fallback for an extent-less receipt, .. shapes, zero-line deliveries) are _native_read_coverage's docstring. Anything but read is disclosed in the header and as a typed capability_delta, never as a refusal (R8).
  • Delegated session (an agent_session row): AgentSessionReviewExecutor with the same task and contract (no output schema — a report is prose); the session reads the repository itself, the memory files reach it inline, its reads are not host-observed and coverage is recorded unobserved.

Every delivered report is prefixed by the host provenance header — <!-- deep-review provenance: delivery=…, model=…, memory=n/7[, memory_missing=…, memory_empty=…, memory_oversized=…, memory_read_error=…], coverage=…, incomplete=…, attestation=…[, rounds/tool_calls/receipts/end_reason/transcript/landing on the native episode only] --> plus one human-readable line, every comment value and every external value on the human line sanitized and bounded through _header_value (one value per memory disposition, each at most the seven whitelisted basenames), the fact set built per delivery (a session carries no round or receipt facts and its attestation is unobserved by construction) — so a reader, and the next task's context (which quotes memory/deep_review.md), can tell a packed report from a retrieved one and a complete one from an incomplete one (R9). incomplete is derived from the facts each delivery actually carries: the native episode's typed native_incomplete; the packed call's provider stop marker — the OpenAI-compatible normalizer's response_finish_reason == "length" or, on the direct-Anthropic lane (the shipped anthropic:: deep default, which sets no usage finish reason), the message's stop_reason == "max_tokens" — as output_reserve (the report hit the 100K output reserve); a session's completeness is unobserved. The packed pack's memory dispositions ride its OMITTED section and state/deep_self_review_context.json's sibling usage fact deep_review_memory. The packed path records normal usage evidence and writes the coverage manifest to state/deep_self_review_context.json. The retrieving deliveries record the row's «Выполняется как» execution on every outcome (responded, empty response, executor exception) and persist their prompt and response through persist_call; the packed delivery persists its request and response through chat_observed and records its execution only once the call has returned (responded, or an empty response as an error row) — a transport exception on the packed call is the typed deep_self_review_error failure without an execution record. agent.py stores the report at memory/deep_review.md ONLY when the review delivered: every failure returns typed usage (execution_status=infra_failed with deep_self_review_unavailable / deep_self_review_error), lands in the task result and a typed task_error event, and leaves the previous report in place. Availability is route-aware (deep_review_route): the packed row keeps the ≥1M floor — a route whose Capability Evidence (reviewer_window.resolve_reviewer_window, the shared resolver) confirms a sub-1M window is refused typed, because the pack IS this delivery's guarantee and is never silently shrunk to a smaller window (a native or session deep_review row serves such a route); an UNKNOWN window keeps the documented full-window assumption every packet surface shares and is disclosed in the header (window=assumed_1000000); the openai:: route is trusted only without OPENAI_BASE_URL, and an owner-pinned OpenRouter -pro slug is rewritten to the direct default, plain Sol being the shipped default on both routes — a native row needs its routed model's credentials, a session row a healthy delegated route (subagents.route_health); the request_deep_self_review tool and the agent read the ROW, not the model key. On every delivery the model has no mutating tools and does not run plan, task-acceptance, or commit reviewers; its report is durable diagnostic memory under BIBLE authority, not implementation or publication authority.

Rationale. The review ran without tools for as long as it did because its guarantee was the PACK: one ≥1M-context call whose Atlas assembles fail-closed — a required artifact that does not fit refuses the whole review instead of shrinking it — so the reviewer provably held BIBLE.md, the protected runtime and the whole memory whitelist at once, and no smaller-window model could impersonate that coverage. That guarantee is also the packed delivery's limit: it exists only on a ≥1M route (sub-1M installs had no deep review at all), the reviewer cannot follow a call chain out of the files the greedy budget admitted, and "the pack contained X" was silently read as "the reviewer considered X". Retrieving deliveries are admitted now because they buy two things the pack never could — a deep review on any install with a payable route or a subscription session, and host-OBSERVED reads on the native episode (the receipts say which files were actually opened) — while the guarantee they lack stays honest instead of hidden: the mandatory BIBLE.md read is checked after the fact and a miss is disclosed, never assumed; a delegated session's reads are unobserved and say so; the packed delivery remains the default and the only one whose coverage is assembled rather than retrieved. Whole-repository retrieval in one episode is not a goal: the four canonical docs alone (≈1.06M chars) exceed any transcript bound, so the retrieving task is a targeted survey through the navigation maps, not a promise of full coverage.

Post-task reflection

The root post-task checkpoint decides whether an error-bearing or non-trivial run warrants Experience Review. reflection.generate_reflection sends the Light route a bounded task goal, trace summary, tool-use profile, errors, review and child evidence, and the same frozen non-final cost snapshot the task summary uses; it runs outside the tool loop, records its own usage, and its failure never erases the delivered result or changes a review verdict.

A reflection lands where it durably belongs: a non-project root appends the full entry to the canonical logs/task_reflections.jsonl; a project-scoped root appends the full entry to its project drive and the canonical log receives only a bounded pointer row — full project text never enters the canonical log, which feeds future global context. A project-bound task's context includes a bounded labeled tail of its own project's reflections; the headless mirror drive of a split root is never the reflection home, and the Pattern Register update stays canonical in both cases. Every entry carries task identity, evidence, lessons, backlog candidates, and validated memory actions. MEMORY_ACTIONS_JSON permits only scratchpad_append, knowledge_write, and identity_update_candidate, at bounded count and size. apply_memory_actions routes accepted actions through provenance-preserving memory and knowledge APIs. An identity_update_candidate is recorded in the scratchpad for review and is never auto-written to identity.md. For a project-scoped task, only project knowledge is written; scratchpad and identity actions are skipped so local facts cannot contaminate the canonical self. Reflection may propose a future campaign or backlog item, but it cannot enqueue, review, commit, or enable one.

Only the root runs full post-task synthesis once (root_phase_checkpoint makes the paid phase at-most-once across restart); children contribute evidence, never a second global synthesis. The owner's final answer does not wait for blocking synthesis: after the durable result is stored, the final send_message is delivered immediately while the buffered-return copy is RETAINED (queue.put is not a delivery receipt); both copies carry one delivery_id and the supervisor suppresses the duplicate through the durable registry in supervisor/terminal_delivery.py. The same file holds a bounded PENDING outbox — ONE seam for the normal, cancel, and reap terminal paths: a terminal answer is recorded as owed BEFORE it is enqueued and cleared in the same write that marks it delivered, so a crash between settle and send replays it instead of losing it. Every non-ephemeral root's answer enters this outbox at durable-result persistence time under the canonical final:<tid>:<digest> id. Replays back off and are bounded; an exhausted or capacity-evicted row is dropped LOUDLY (full text preserved on disk, typed terminal_delivery_exhausted event, chat notice) — external transports stay at-least-once and that residual is disclosed. task_done still goes last through the buffered return — an early task_done would release the queue slot and start child-drive cleanup while post-task still runs — so a worker reaped during a hung synthesis has already delivered the answer. Synthesis receives a sealed final package from the durable result: submitted final text, its artifact manifest, and completion_observations. The terminal writer preserves full redacted action observations in the canonical artifact store (task.budget_drive_root or drive_root) before compact publication. Completion sources use the existing write-once source_handles/context_checkpoints store with verified task_source refs; the published-ref closure includes completion observations before child cleanup. Sources stay outside deliverables and inferred readiness. Their native reader selector uses get_task_result(include_completion_source=true) with the source task id and its canonical drive. It first returns complete character length/hash, then explicit source_start_char/source_end_char ranges; artifacts.text_source_range_projection shares the unchanged work-order range contract. Source bytes, kind, path containment and SHA are checked before any excerpt. Packet-only summary/reflection receive per-send-tool counts, each family's latest recorded return, and task-related skill readiness with coverage; full-source references are for later readers, not evidence the synthesizer has read. Positive observed facts correct error-trace impressions, while tool success does not prove owner receipt, empty material does not prove absence, and skill readiness does not attribute an owner's action to the task. Recovery uses these stored observations; task-summary requests/responses use chat_observed.

Pooled workers retain their slot until root post-task synthesis settles, for API-only and subscription tasks alike; early final-answer delivery keeps the response independent from that queue timing. Ordinary native post-work stays on its registered actor thread without a pooled worker slot. Its existing TaskModelWait owner remains available through POST_TASK_SYNTHESIS_INFLIGHT after ordinary dialogue admission closes; detached server post-work binds its own live owner in the same registry. Temporary role overrides follow that owner, and the existing task mailbox remains available until the terminal post-task checkpoint. Gateway decisions use the phase owner independently of closed dialogue admission, and activity hydration enriches the existing direct row with its finalizing state and model waits. An open phase remains finalizing rather than appearing completed. Typed quota/auth waits resume only the unsettled call, while stop or unknown outcomes degrade the phase without repeating finished stages. Restart recovery still degrades an indeterminate running phase rather than replaying a possibly paid request.

Durable memory and project focus

context.py assembles static governance, semi-stable memory, and dynamic task evidence without treating truncation as forgetting. consolidator.py publishes dialogue summaries only after every part of a logical block succeeds, preserving raw chat generations and the captured generation cursor. Each full Light request is measured with context_fit against fresh role/account capacity evidence and calibrated prompt density. Local output reservation uses llm_local.local_context_limits, the same normalization as the wire, while the requested output budget stays unchanged. Subscription preparation reads the existing metadata-only role/account catalog and carries its exact binding between parts; actual success/refusal receipts replace advertised identity, and a wait re-prepares against the current account. Missing or stale evidence remains unknown. Oversized source is split without clipping, including within one entry. A real context refusal requires strictly fewer input bytes on the same route. A genuine refusal immediately saves its source hash and route bound in dialogue_meta.json (consolidation_retry) under the caller's consolidation lock, before another part can raise a typed interruption, so a later cycle starts smaller; changed source, route, capacity or output reserve invalidates that bound. Era compression remains a single aggregate request: it can still overflow, in which case the old blocks remain intact. Ordinary failures and empty summaries retain last_consolidation_error; unknown spend stays nullable, and control/resource/unknown model errors preserve propagate_model_error semantics. Scratchpad replacement retains its existing explicit summary and source-journal provenance. When the rendered scratchpad exceeds SCRATCHPAD_SECTION_BUDGET_CHARS, context.py keeps the newest whole blocks that fit and drops the oldest behind an in-band gap marker that names memory/scratchpad.md as the live source; no block is retired by a context build. Knowledge topic mutation and its index rebuild share one stable lock, as do scratchpad mutation and markdown regeneration, so concurrent Presence and owner turns cannot publish an older projection over a newer write. Knowledge writes retain source metadata; identity changes remain on their dedicated authority path. The Development context matrix and context_layout.py own which reference form is resident.

Ouroboros remains one identity across Main, project rooms, and Background Consciousness. A project is a focused working room, not an isolated sub-mind: unified dialogue memory remains available to the one agent, while an executing project task preferentially receives its own thread, journal, workpad, and project knowledge. project_facts.py routes project facts to projects/<id>/knowledge; subagents inherit the root's resolved project id and never derive a new one; there is no per-project identity or scratchpad. An id minted from a DISPLAY name collapses dash runs and carries a short deterministic suffix when the name held characters the slug could not keep, so two different non-Latin names can no longer share one project; the normalizer itself is unchanged, so existing ids are never re-slugged and stay reachable by their explicit id.

The projects registry owns immutable project identity, canonical chat id, optional working directory, lifecycle/tombstone state, routing generation, and activity revision. Admission persists the resolved project id in the task itself. project_lease.py serializes assignment of pooled roots by Project while allowing their own subagent trees; it is not a physical-folder lock and does not withhold tools from ordinary conversation. Binding/history files support routing and presentation, not the lease. Delete closes routing, cancels/quiesces the tree, and tombstones only after settlement, preserving everything for recovery. The durable binding is the SINGLE truth about a task's project: the in-task scope guard and project_facts.resolve_project_id read it FIRST, ahead of task["project_id"] and a worker's in-memory ctx.project_id, which are copies a mid-run conversion never reaches (a guard reading only the copy is how a task already bound to one project minted a second, empty one). ensure_project_scope can create or bind the current root to one project mid-execution: it persists the durable registry binding first, then marks the live queue/lease surface under the queue lock so the lease recognizes the running task as a lane occupant; it is idempotent for the same project, and a task BOUND elsewhere is renamed rather than re-scoped - its scope call carries the requested display name to the project it already belongs to and creates nothing. A project-SCOPED but unbound run keeps the older refusal, since there is no durable project to rename, and a child still cannot escape the inherited scope. An unreadable bindings store is disclosed once and read as "no binding" on every path, hot and guarded alike: the refusal is reserved for the measured case, a readable binding to another project. That refusal precedes every side effect on BOTH conversion paths whenever the binding is readable at one of those checkpoints - no project row, lease mark, broadcast or announcement survives it - the UI conversion answers 409 naming the bound project by id and display name in one human sentence. It re-reads that authority after the naming step and before its first side effect, since naming can await a model call while the task binds itself, and a durable bind refused after the mark restores the lane to the value it held, broadcasts nothing and answers the same 409, so no lane keeps a project the binding does not name. One residual survives that last case, a conflicting bind landing in the microseconds between the re-read and the bind: the empty project row the request had already created stays behind, because the registry deliberately has no primitive that removes a row and its delete lifecycle tombstones the id permanently instead, which would cost the owner that id (and, in-task, the display name it derives from) forever. The row holds no task and no binding, and the owner deletes it like any other project. The lease mark is fill-only by default; the single exception is a conversion that already holds the binding it is about to write, which moves the lane onto that binding.

Project journal.jsonl records curated milestones and workpad.md retains active working context; focused context includes the workpad in full and recent journal rows with a visible pointer to older entries. On root completion, only high-signal blockers, questions, and interface contracts are mirrored once from the ephemeral task-tree ledger into the durable journal, and a finished root whose effective working tree is not the registered working_dir writes one typed "work lives at @ " journal row from facts the task record already holds. A project digest gives consciousness a concise completion signal without pretending to be the raw project memory.

promote_chat_to_task, route_to_project, and steer_task become successful only after their token-matched supervisor facts are durable in the existing task result, queue snapshot, annotation, or mailbox authority; with several possible tasks the LLM chooses, code auto-delivers only the unambiguous one-target case, and an unconfirmed or stale receipt fails visibly rather than launching a second root. A routing/promote decision turn receives host-built ground truth (each project's registry working_dir plus bounded typed projections of recent task results — never raw result text), read through the registry row's durable last_task_result_id pointer with a bounded scan fallback; only the absent-pointer case writes the pointer back, because a non-empty pointer that failed to resolve is typically a split-drive result whose canonical copy-back has not landed. Those recent results are the owner's ROOT results only: a swarm child is not an addressable predecessor (it is reachable through its root) and is counted as its own omission rather than evicting a root from the window, while a project room's direct-chat root stamps the same pointer, so the room's own work is offerable without acquiring letters home. The turn is also shown any routing receipt that already exists for the SAME owner message, a fact on the contract it already receives and never a ban: whether one message becomes a second root stays the model's judgment. Canonical owner routing and project UI are projections over those same authorities: a routing receipt proves admission, not completion; an unread indicator proves a visible revision, not memory isolation.

Skills and extensions

Skill capability grows through distinct gates: discovery and manifest parsing (skill_loader.py), content-hash-bound review (skill_review.py / skill_review_runner.py), owner grants, dependency reconciliation (marketplace/isolated_deps.py), enablement, readiness (skill_readiness.py), and execution (tools/skill_exec.py). Discovery establishes identity, source, provenance, conflicts, and hash; it does not confer trust. Review status, grants, enablement, and dependency health remain independent durable facts under data/state/skills/<name>/. The model-facing catalogue (list_skills and the per-turn Installed Skills section) projects the same live extension facts as /api/extensions; each enablement change best-effort appends a typed skill_enabled_changed disclosure row to logs/events.jsonl naming the actor, empty for writers not yet labelled, and an append failure is logged and never blocks the enablement change. A visible or enabled skill is not executable until skill_readiness_for_execution() says the current payload satisfies every required gate.

A skill may declare a reviewed presence: behavior profile: instructions, knowledge topics, runtime defaults, and portable capability requests that installation-local selections resolve to exact targets. Admission requires the bound behavior skill installed, enabled, freshly reviewed, and complete for every required request, then compiles one immutable positive capability ceiling copied through task_contract — so mutable skill/Settings state cannot broaden a live turn. Bundled native skills and editable marketplace/user payloads occupy separate payload-plane buckets with owner and review state outside the payload; declared conflicts are symmetric and never delete a payload. extension_loader.py and the isolated-dependency layer load only a ready, hash-matching extension; a review PASS alone does not prove its dependencies installed or its widget loadable.

skill_lifecycle_queue.py serializes install, update, review, grant, enable, and removal work and exposes queued/running/succeeded/failed plus stale metadata; stale is recovery evidence, not a fake unlock of a still-running thread. Scheduled work is reconciled by resync_skill_schedules() and can run only after skill_readiness_for_execution(). Schedule evaluation is a DST-aware system using the shared cron/timezone contract. Evolution remains hard-blocked in light runtime mode. These separations let skills expand capability without turning discovery, a UI toggle, or old review state into execution authority.

Skill publication

The passive installed-skill projection never launches Betterleaks and never claims the current bytes are publication-ready. Selecting Publish calls POST /api/skills/{skill}/publish-preflight, which resolves one current payload, captures and scans its bytes, recomputes review staleness, and returns exactly one backend-authored state (ready, warnings, needs_attention, repairable, hard_block); the browser only renders those facts, and only hard_block prevents task creation. The authoritative flow is:

passive index (no scan) → selected preflight → explicit confirmation → ordinary managed task → fresh immutable capture bound to the stored review hash → payload scan → GitHub read-only planning → derived-output scans → first GitHub mutation → validated same-skill pull-request receipt → ordinary acceptance

Every outbound byte derives from the capture; the mutable live payload is neither reread nor rehashed to authorize the transaction, which is what prevents time-of-check/time-of-use drift. Literal Betterleaks high findings block the current outbound call; lower or unknown confidence remains a redacted warning. Packaged installs resolve the bundled betterleaks-standalone; source checkouts resolve the exact managed runtime installed explicitly with python -m ouroboros.betterleaks_runtime install — Publish never downloads it. A top-level skill_publish task can be accepted only when pre-truncation metadata contains a validated pull-request receipt for the requested skill and configured Hub repository; the receipt is a narrow veto prerequisite that never manufactures PASS. A definite failure before any submission-branch request makes an explicit Publish objective fail even under degraded review, with PR not created displayed separately from review/preparation facts. Last-confirmed stage alone never proves that the next request was unsent; a previous unknown branch/PR attempt or valid same-target receipt is preserved.

A failed publication envelope names its cause beside the stage: reason_code (the stage that failed — fork_sync_failed and its siblings), a repair_hint chosen by producer evidence (missing CLI, no configured credential, an observed 401/403, a 409 conflict, a timeout), and the transport's sanitized error_detail with github_status/github_operation when observed — the status is read only from gh's own error shapes (the line-terminal (HTTP NNN) of gh api, its message-less gh: HTTP NNN, and the HTTP NNN: error line of the other commands), never inferred from prose. Mandatory fork synchronization still stops before any branch, commit or PR mutation, completed stages stay readable, and ambiguous PR settlement remains a read-only exact lookup. Publication metadata reads the producer JSON before the shared ToolResult host-note separator, so appended route or safety notes cannot hide attempt diagnostics or a validated receipt.

Owner lifecycle actions share skill_lifecycle_actions.run_skill_action: grant/toggle execute their existing effects in the lifecycle lane, local delete uses skill_uninstall_state, and attestation keeps its deterministic floor in skill_owner_attestation. The host supplies actor identity and checks exact resource/revision plus an existing member chat message, answered quiz or owner-mailbox record when owner intent is needed; the model interprets that source. A generated edit-and-review request confers no grant, attestation, deletion or enable authority. Selected-skill enable calls resolve the actual original owner source when no explicit chat/quiz/mailbox source is supplied; a client allow_enable flag cannot authorize them. The model interprets Repair and run as enable-and-test intent. Presence ceilings precede this source path. A known owner disable in enabled.json newer than that source refuses the old request; load-error reverts and legacy unlabelled snapshots confer no owner attribution. Selected Repair reviews never auto-enable: the model chooses the ordinary toggle_skill call. No permission ledger or HTTP impersonation is added. The configured auto-grant policy remains separate from enablement; explicit owner disable survives review and free replay. skill_exec returns the revision captured before its physical script launch, and extension tool receipts carry the descriptor's content_hash from the same publication as extension_generation, only after physical dispatch; neither is review PASS or a semantic test verdict.

Marketplace update retains the owner's selected version through repeated retries. install.PayloadRollbackSnapshot captures payload/environment and the existing affected lifecycle-state quintet; update and adopt use that same owner. A single restore path verifies required reload before claiming rolled_back, preserves independent enablement/history, and never deletes an already restored payload when a state write failed.

Catalog updates disclose the exact current and proposed version strings in the tool result and host-authored PR body; they do not infer semantic ordering. The existing atomic publication-record owner also handles explicit local clearing: it compares the displayed published object before setting that section to null, preserving unknown siblings. Clearing changes no GitHub PR or installed skill bytes. Both skill views refresh through their existing selection/refresh paths, without polling PR state.

MCP and browser-facing external tools

mcp_client.py owns configured HTTP/SSE and local stdio MCP discovery and invocation. HTTP/SSE entries validate URLs and auth headers; secret_masking.py owns the shared exact MCP token placeholder shapes; load-time placeholder repair runs before environment precedence, so a real environment credential is not mistaken for a wire mask, and it is limited to recognized top-level Settings secrets — nested MCP values are never silently migrated. Stdio entries pass one executable command and an exact string args list directly to the MCP SDK without a shell; optional cwd, literal env, and env_from_settings (environment-name → existing setting-key references) resolve from the same saved configuration for discovery and calls. References override matching literal names; omission preserves SDK defaults. Settings classifies referenced values as ordinary or secret; only secret values receive exact diagnostic masking, including JSON-escaped echoes and short secrets. Unknown fields are retained with a visible not-applied warning while a valid server remains usable; invalid known fields or references produce MCP_CONFIG_ERROR. The Settings response-only auth_configured flag never becomes configuration. The UI preserves both environment forms and user extras. The SDK context owns process shutdown. Discovered tools join the selected initial capability envelope including explicit ephemeral decision turns and enabled, granted extension tools, behind the same network resource gate as on a managed task, and schemas(), get_schema_by_name, and execute agree on that; discovery failure produces an explicit capability omission through list_available_tools, never a silent removal. Descriptions and results remain untrusted data, and every call still crosses registry, resource, safety, timeout, and result-handling policy. Web-tool prohibition and network prohibition are distinct: web=false alone does not disable configured MCP or extensions, while an explicit network=false and tool disables remain enforced at discovery and dispatch.

Browser tools are stateful and thread-sticky because Playwright sessions and greenlets have thread affinity; they cannot be scheduled as ordinary parallel stateless calls. A stateful-tool timeout therefore RETIRES the browser generation: the shared browser_state slot is replaced, the abandoned worker keeps writing only into its retired one, and the close is queued on the retiring executor so it runs on the owning worker thread when the hung call settles. Retired sessions are bounded in-process at _RETIRED_GENERATIONS_MAX per task, after which opening another session is a typed BROWSER_BACKLOG_RETIRED_SESSIONS refusal; generation isolation is best-effort under concurrent replacement (a fully closed class needs a process-isolated browser worker — disclosed future design). Model-driven in-page evaluation goes through _evaluate_bounded: Playwright's evaluate accepts no timeout, so the expression is raced against an in-page rejection, bounding the ASYNC class honestly and no further — a synchronous event-loop block cannot be interrupted from inside the page, and the outer tool timeout remains its backstop. (The one direct page.evaluate is _wait_for_page_paint's immediate paint-flag setter, paired with a 500 ms bounded wait_for_function.) The caller's timeout also becomes the session default (page.set_default_timeout), floored on the action path so the five-second action default cannot strangle a capture. Chromium is the default; WebKit and device descriptors are targeted tools for a real Safari/iOS risk, not a universal acceptance matrix and not a claim that a narrow Chromium viewport is Safari-equivalent. First-party PR helpers are normal built-ins whose mutating operations remain subject to selected-root policy, runtime mode, delegated-child constraints, credentials, and reviewed-publication authority.

Browser target policy is browser_policy.py; its control-service identity is three-valued and never a port-number or pathname guess. A live binding in state/server_port.bindings.json proves an endpoint ours; a service's missing snapshot can mean an older installation or failed publication, so its recorded facts still name what is EXPECTED — the integer state/server_port (the launcher's state/server_process.json can prove main), the Host Service configuration beside an expected main, and the local model's own custody row with its live argv — and a matching expected endpoint whose process cannot be verified is refused as unknown rather than treated as foreign (server_process.runtime_service_identity). Only a live snapshot of that same service supersedes its legacy endpoint expectation: main cannot vouch for Host Service, and retiring Host metadata alone does not erase its still-configured expectation. Legacy local-model lookup still reads the whole process ledger for matching loopback targets without a verified model binding; proven non-loopback hosts skip that read, while unresolved hosts retain the port-specific legacy check. A permitted restricted request reuses its completed target classification, and ordinary unrestricted requests need no owner-operation classification. Positive POSIX identity uses the custody owner’s measured fingerprint contract; unavailable measurements never count as proof. Every other port stays an ordinary target, including an unrelated application that reuses an /api/owner/... pathname; the owner-operation request shapes apply only at a proven or unknown Ouroboros endpoint. A restricted target that DNS cannot classify is a typed BROWSER_POLICY_UNAVAILABLE refusal before navigation. Chromium and WebKit follow HTTP redirects natively and Playwright route callbacks see only the first URL of a chain, so tools/browser.py keeps every observed navigation response and re-checks each document's redirected_from chain with the same target predicate before any page result: an allowed→blocked→allowed redirect withholds the content it produced and keeps refusing actions on that document, while the forbidden hop's request has already been dispatched — an owner-accepted, disclosed residual (pinned by a strict xfail), not a pre-request or DNS-rebinding guarantee. Cookies, POST semantics, service workers and native redirects are untouched: no proxy, CDP fork, pre-probe or refetch.

Budget tracking

swarm_efficiency.fanout_count counts observed fan-out emissions; fanout_interval_sec_total sums the wall-clock gaps between them, including intermediate parent work. New events emit fanout_interval_sec. The result reader tolerates the two retired aggregate names (wave_count, inter_wave_latency_sec_total) without rewriting stored bytes; event readers tolerate inter_wave_latency_sec. These observations never infer semantic work waves or child-wait time.

The existing time/cost/intrinsic pacing checkpoints carry resource_facts: incremental per-tool call/error counts and producer-reported duration intervals, plus own-task, tree and delegated-tree ledger buckets and the shared unreserved global remainder. Tool intervals may overlap and unmeasured calls stay unknown; no argv, stdout, sleep/poll classification or new stop rule selects behavior. Money is read only when the existing note fires through the cached accounting readers; own and delegated buckets explicitly overlap the tree total. The loop usage carrier retains the tool aggregate across native owner-wait continuation, and the same facts appear in the prompt and checkpoint.

usage_accounting.py is the single monetary policy authority for core-mediated model work over the physical-attempt ledger. Every provider send has a unique attempt id and durable lifecycle reserved → dispatched → settled | unresolved, or reserved → released before dispatch; a dispatched row may enter released only through the typed pre-dispatch transport seam (connection/pool failures proving no request bytes were sent), while ordinary timeouts and unknown errors remain unresolved. Each retry is a new attempt. For every inspectable candidate the attempt id binds exact post-transform raw/context identities and the existing-CAS physical manifest before dispatch; specialized SDK/stream boundaries keep the lifecycle receipt but are labelled opaque. The wrapper covers main and direct calls, children, scouts, all review surfaces, safety, synthesis, reflection, consciousness, retries, and opaque SDK calls; root scopes include their task tree and post-task/review work exactly once. A reviewed external script or extension with model credentials is represented as unknown/unmetered at each host-observed opaque execution boundary unless authoritative settlement exists.

Before summary, reflection, or consolidation starts, the root freezes one shared ledger snapshot (settled subtree cost, live reservations, unresolved upper bound, unknown/unmetered count, integrity, explicit non-final state); all post-task consumers receive that same snapshot, and the final terminal checkpoint remains the only final cost authority — a read failure is unavailable/null, never $0, and there is no reconciliation LLM or parallel cost ledger.

The in-task pacing stop is an explicitly unreserved planning threshold in the shared global pool, resolved once per task as a typed CostCeiling (disabled, active, exhausted_soft_land, unknown), published in the start-of-task runtime budget block as in_task_cost_ceiling and consumed by the loop as the SAME object, so the number the mind is shown and the number that stops the task cannot differ. The root of a task tree resolves the minimum of the configured percentage of global remaining budget and the root-tree cap minus a small planning margin; enabled descendants retain that original number instead of taking another percentage of a later wallet. The scheduler forwards the original root carrier unchanged at every non-root hop; a legacy missing carrier keeps a disclosed local resolution, which is never relabelled as the original root fact. Explicitly disabled profiles keep the early axis disabled. A root without a task cap resolves its early threshold from the starting wallet, while actual global/root monetary admission stays independent. Every host cost surface prints the bound that binds first and names which one it is. The loop decides against subtree-accounted spend including in-flight holds, disclosing an own-cost fallback as a lower bound when tree accounting is unavailable; graceful finalization runs before the ledger fence and never weakens it — unknown is not zero. The planning margin pulls the stop earlier, and a post-round affordability crossing then borrows the fence's OWN per-attempt reservation — cache-aware from the task's last settled split for the same provider, normalized route and review surface (plan, acceptance and skill reviewer sends settle under the task id but never pose as the transcript's own split), bounded by that split's TTL horizon and a full write otherwise — so the task soft-lands while one wrap-up call would still be admitted instead of dying answerless at the fence. The price is compared with all known applicable remainders, including a fresh global-wallet projection for global-only tasks; candidate probes share that observation. This does not reserve money: a competing task can consume it before the final atomic reservation. The bounded context proxy only pre-screens: a proxy stop and any prompt the proxy can understate (native image parts) are decided by exact pricing — every proxy stop and every image prompt is first priced non-destructively on a copy of the transcript, and service finalization plus the prepared candidate happen only once that exact probe confirms a stop; a task that cannot reserve even one wrap-up says so in its own words — and a tree already over its ceiling still reaches the ceiling stop when a wrap-up is affordable; the exhausted-ceiling soft landing (a per-task cap at or below the planning margin) prices the same prepared candidate and ends as budget_wrapup_unaffordable rather than a fence pause when it cannot fit. Tree spend is read for every task that has a root, with or without a per-task cap, so a global-only ceiling decides on the subtree, not the parent alone; a reply that still asks for a tool is incomplete on every rail even beside a replace delivery control. An unknown price fails open, a disabled ceiling never arms it, and the cost axis is checked only after tool-call rounds. After a cache expiry between rounds the cache-aware reservation may under-reserve by one write; the ledger fence at the full cap still binds.

One resolver answers what the configured global budget is for every agent-side reader (the supervisor's own startup and settings-reload carriers still parse the raw setting and map absence to no limit — pre-existing, tracked as an issue): an absent setting is silence rather than a choice and resolves to the shipped default, while an explicitly non-positive value means no finite global budget and silences the loop-side global axis. Pre-dispatch pricing is an exact-route, bounded, best-effort lookup from the provider's current catalog: only the normalized exact model id and provider-supplied fields count — no manual price table, prefix inheritance, numeric fallback, or admission allowlist disguised as pricing, which would silently invent authority and grow stale. Unknown price is nullable and fail-open for admission while known spend remains below its limits; it reserves None and settles from provider-reported cost or a later exact price, else cost stays None with cost_final=false. A rejection settles at confirmed zero only when structural provider evidence proves it happened before upstream generation with zero usage. review_wave_admission applies the same per-attempt math against the tighter of the global remainder and the root remainder — the two fences reserve_attempt enforces, the binding axis named — before skill, plan, task-acceptance, or P3 commit-gate reviewers launch (per-slot pack sizes and output reservations when the wave mixes packs; the caller's bound root fence when the root has no ledger row yet), and the managed-update assisted-apply admission floor reuses the same estimator against the global remainder before any destructive merge step; an unpayable reviewer row is bypassed, never swapped to a different model, so the audit stays honest about which model reviewed, and an unpriced slot is disclosed and contributes no invented price while priced siblings still bind — one unknown route cannot disable admission control for a paid wave.

Each reserved attempt records the exact applied global_limit_usd (including resolver fallback), explicit unbounded state and global_limit_source/global_limit_revision; transitions preserve them. Provenance follows the chosen numeric request or bound scope, so an explicit override never inherits another limit's revision. Task-start and fallback sources name the existing budget resolver, not an inferred settings author. No settings revision is manufactured when that owner supplied none; legacy records remain unknown, and compaction preserves exact original rows in its existing archive. Validation, reservation, transition, append, and fsync share one short cross-process lock; network work remains outside it, preserving atomic accounting without holding the lock across provider latency. A torn tail is quarantined loudly, the validated prefix remains readable, and affected projections stay integrity-degraded and non-final because paid work may be missing; failed settlement persistence leaves the attempt dispatched/unresolved, and a root budget refusal is durable, cleared on resume only after proving no paid dispatch occurred. Read paths use incremental validated replay over the append-only log (resume fingerprints, fallback to the authoritative full locked read, bounded caches that change read cost, never accounting meaning), and the appender repairs a crash-torn newline-less tail so it costs at most itself (usage_ledger.py/_usage_rows.py).

For a root task, GET /api/tasks/{id} derives cost_breakdown at read time from the same ledger (own, child, unattributed, disclosed delegated, subscription sessions, unknown/unmetered, finality, authority); it is never persisted and is not a third sum, and an unreadable or unattributable ledger omits the whole object rather than returning a confident zero. One physical reviewer send produces exactly one llm_usage row, emitted by the review substrate and carrying that reviewer's wave and slot attribution, and a delegated session row reports its own route provider and resolved model instead of an inferred one. state.json, task results, llm_usage, /api/state, and /api/cost-breakdown are compatibility projections only; startup's resumable importer records source hashes, imports only attributable usage, and represents ambiguous history explicitly without rewriting source logs or fabricating attempts.

7. Configuration (ouroboros/config.py)

ouroboros/config.py is the SSOT for paths (HOME, APP_ROOT, REPO_DIR, DATA_DIR, SETTINGS_PATH, PID_FILE, PORT_FILE), process constants (RESTART_EXIT_CODE 42, AGENT_SERVER_PORT 8765), every settings default below, and the load/save/env machinery: load_settings(), save_settings(), apply_settings_to_env() (copies hot-reloadable runtime keys into os.environ), normalize_runtime_mode() (one clamp shared by the save path, the read coercion, and onboarding validation), get_runtime_mode()/get_skills_repo_path(), and acquire_pid_lock()/release_pid_lock(); ouroboros/update_channels.py owns get_update_channel()/get_update_branch().

Settings file: data/settings.json under the data root — ~/Ouroboros/data/settings.json by default, with APP_ROOT, DATA_DIR, and SETTINGS_PATH independently env-overridable. Access is file-locked. secret_masking.py is the wire-placeholder authority for known and owner-defined top-level secrets: load_settings() repairs only recognized disk placeholders BEFORE environment precedence is resolved, so a real environment credential is never classified as a mask, and prepare_settings_for_persist() applies the same top-level repair at the common writer boundary; nested MCP values are never silently migrated.

ouroboros/openrouter_attribution.py is the application-identity SSOT for every first-party paid OpenRouter request (canonical URL + X-OpenRouter-Title); a fork must use its own URL rather than competing to rename one app record.

Reading and writing the settings document

agent.handle_task binds the admitted task's settings_integrity.TaskSettingsSnapshot for the complete task entry. Its normalized document and exact environment projection are separate private in-memory views: absent and empty values remain distinct, concurrent tasks keep their own models/keys/Supervisor/Review settings, and explicit task route overrides still win. A short SETTINGS_ENV_LOCK serializes capture/publication only, never task execution. runtime_setting, runtime_settings and runtime_environ reuse that view; settings writers keep reading the current document, immediate effects stay live, and Access retains its boot pin. model_wait.copy_wait_context and task-owned helper threads carry only the existing settings binding alongside their established owners. Out-of-process extensions receive only their permitted typed next-task values through the existing private per-call payload; the child still validates current grants, reads immediate values live and receives no whole settings snapshot or extra credential environment. Failed reload loudly retains the prior environment while disclosing unavailable document-only values.

A settings document on disk was written by whatever release the owner last used, so reading one starts by translating it into today's vocabulary. normalize_settings_raw() is that translation and the only copy of it: type coercion against the declared defaults, the deprecated per-subsystem retention keys folded into the unified one, the retired acceptance-pass count consumed into the shared review-cycle cap, the keys a release retired dropped, the renamed model slots promoted, and secret placeholders repaired. Every step preserves an owner customization written under a former key, and the ORDER is load-bearing (the pass count is consumed before the retired purge would drop it; the purge runs before the slot rename, so a retired spelling is never promoted), so every reader applies it BEFORE the shipped defaults merge — load_settings(), the owner endpoints' _owner_read_settings_raw(), and the Colab re-run's build_colab_settings() alike. "Raw" in that name is about the runtime-mode ratchets it deliberately skips, never about the migrations. It is pure and idempotent — it touches no file and no environment — which is what lets a read stay a read and lets a read-modify-write apply it on every save; both in-process readers share one read primitive (settings_integrity.read_settings_json_verified), so a pinned snapshot that changed refuses the owner reader exactly as it refuses the loader. This seam carries the VOCABULARY normalization only: the provider normalization (server_runtime.apply_runtime_provider_defaults) is a separate, never-persisted derivation every route consumer makes over the effective document, context_fit.resolve_context_fit_route() included.

Five functions persist a settings document, and every one calls serialize_settings() and commits its output through a byte-exact helper (utils.write_text_atomic(), or Path.write_bytes on the config saver's rename-less OSError fallback) — never a text-mode write, which would turn LF into CRLF on Windows. Three persist THIS process's document through prepare_settings_for_persist(), the single point where the disk-authored silence rule and the owner-only context/safety ratchets are enforced against the value ON DISK: config.save_settings(), gateway/owner_settings._owner_update_settings() (which _owner_write_settings() is one caller of), and the packaged bootstrap's packaged_cli._save_settings(). Two are exempt by design and pinned as such: the one-window raw context-pair migration (context_mode_compat.normalize_and_persist_context_mode_compat(), written under the load lock — the raw mapping with only the pair changed, never a defaults-merged document) and colab_bootstrap.write_colab_settings(), which writes a generated document for a foreign data root the prologue's on-disk proofs would answer wrongly for. One scan closes the inventory over ouroboros/**, supervisor/**, server.py and the repo-root launcher.py (tests._shared.settings_writers): a function is a settings writer when it CALLS serialize_settings(), or when it names the settings path or file AND does a write-shaped thing; it counts as ROUTED only when it CALLS the prologue, and naming either in prose is neither. The flagged set must equal the five writers plus the scan's declared non-writer matches, so a sixth writer in those roots fails the tripwire whether or not it is routed.

An owner endpoint changes one decision inside a document it does not otherwise own, so it must write the whole document back. _owner_update_settings(transform, expected_digest) does that read, change and write inside ONE settings lock: the transform receives the document as it is under the lock and returns what to persist, or nothing at all, which is how a no-change decision avoids rewriting the file. An endpoint acting on an earlier read passes the digest that read saw (settings_document_digest(), the same staleness question the onboarding transaction asks), and a mismatch refuses before the transform runs, so a concurrent owner change can never be reverted key by key while the request answers "saved".

LLM output token budgets

Providers name the same output-token budget differently: OpenRouter/Anthropic-compatible calls send max_tokens, while every official direct OpenAI Chat route sends max_completion_tokens — a real provider-wire boundary, not naming style. Direct OpenAI also sends the requested reasoning_effort provider-wide; model-name prefixes are not capability authority, and only exact-route success-confirmed wire evidence may adapt a request. Runtime floors (numeric SSOT: the constants in code, pinned by tests/test_max_tokens_constants.py):

Surface Output-token budget
LLMClient.chat() / chat_async() defaults 65,536
Main task loop (loop_llm_call.MAIN_LOOP_MAX_TOKENS) 65,536
LLMClient.vision_query() and VLM tools (analyze_screenshot, vlm_query) 32,768
Review synthesis dedup 16,384
Chat block consolidation, era compression, scratchpad consolidation 16,384
Execution reflection and pattern-register update 16,384
Post-task summary (agent_task_pipeline) 16,384
Improvement-backlog grooming (improvement_backlog.groom_backlog) 8,192
Post-task evolution promotion decision (post_task_evolution) 8,192
Context compaction round summaries 32,768
Skill publish PR body generation 8,192
Background consciousness loop 65,536
Project naming LIGHT one-shot (project_naming.llm_project_name) 256
Update letter LIGHT one-shot (update_letter.write_letter) 1,024
Provider Test (llm_probe.PROVIDER_TEST_MAX_TOKENS) 16

Default settings

A registry of config.SETTINGS_DEFAULTS (exact defaults stay canonical in config.py; this table is test-mirrored against it). Rows marked env-only are operator environment levers with no settings.json carrier.

Key Default Description
OPENROUTER_API_KEY "" OpenRouter credential
OPENAI_API_KEY "" Official direct-OpenAI credential
OPENAI_BASE_URL "" Legacy OpenAI base-URL override
OPENAI_COMPATIBLE_API_KEY "" OpenAI-compatible endpoint credential
OPENAI_COMPATIBLE_BASE_URL "" OpenAI-compatible endpoint base URL
CLOUDRU_FOUNDATION_MODELS_API_KEY "" Cloud.ru credential
CLOUDRU_FOUNDATION_MODELS_BASE_URL https://foundation-models.api.cloud.ru/v1 Cloud.ru base URL
GIGACHAT_CREDENTIALS "" GigaChat auth key
GIGACHAT_USER "" GigaChat user login
GIGACHAT_PASSWORD "" GigaChat password
GIGACHAT_SCOPE GIGACHAT_API_PERS GigaChat API scope
GIGACHAT_BASE_URL https://api.giga.chat/v1 GigaChat base URL
GIGACHAT_VERIFY_SSL_CERTS true GigaChat TLS verification
GIGACHAT_PROFANITY_CHECK "" GigaChat profanity filter passthrough
ANTHROPIC_API_KEY "" Official direct-Anthropic credential
MINIMAX_API_KEY "" MiniMax credential
MINIMAX_REGION "" MiniMax region (empty resolves global_en)
DEEPSEEK_API_KEY "" Optional. DeepSeek direct provider key (deepseek::... model values, OpenAI-compatible API at the fixed official endpoint)
OUROBOROS_NETWORK_PASSWORD "" Non-localhost HTTP gate password (server_auth.py; unset only warns — see §8 packaging note)
OUROBOROS_SERVER_HOST 127.0.0.1 HTTP bind host (0.0.0.0 for Docker/non-loopback)
OUROBOROS_UPDATE_CHANNEL stable Update channel: stable/qa/development (§8)
OUROBOROS_MANAGED_UPDATE_FETCH_TIMEOUT_SEC 300 Managed-update fetch ceiling
OUROBOROS_RESCUE_GIT_TIMEOUT_SEC 300 Per-process ceiling on rescue Git commands
OUROBOROS_TRUST_NONLOCAL_BIND_WITHOUT_PASSWORD unset Env-only: 1 permits saving a non-loopback bind without a password
OUROBOROS_MODEL google/gemini-3.8-flash Main model
OUROBOROS_MODEL_HEAVY "" Legacy slot: readable for migration/history only, out of active routing
OUROBOROS_MODEL_LIGHT openai/gpt-5.6-luna Light model
OUROBOROS_MODEL_ACCOUNTS "{}" Role-owned managed account pins; empty means Auto, fallback entries retain order
OUROBOROS_MODEL_CONTEXT_WINDOWS "{}" Role-owned context sizing assertions; zero means Auto, not a provider limit or scope acknowledgement
OUROBOROS_MODEL_VISION "" Vision model (empty inherits)
OUROBOROS_IMAGE_INPUT_MODE auto Send-time image routing (vision_routing.py)
OUROBOROS_VISION_CAPTION_TIMEOUT_SEC 90 Caption-generation ceiling
OUROBOROS_MODEL_CONSCIOUSNESS "" Background-consciousness model (empty inherits)
OUROBOROS_MODEL_FALLBACKS openai/gpt-5.6-luna Cross-model fallback chain (fallback_cooldown.py)
OUROBOROS_MODEL_MAX_CONCURRENCY 3 Per-(model,route) concurrent provider-call cap (model_concurrency.py)
OUROBOROS_MODEL_SLOT_MAX_WAIT_SEC 180 Concurrency-slot wait bound
OUROBOROS_PROJECT_NAMING_TIMEOUT_SEC 60 Project-naming call ceiling
OUROBOROS_PROJECT_NAMING_ASYNC_TIMEOUT_SEC 8 Proactive-namer async bound
OUROBOROS_UPDATE_LETTER_TIMEOUT_SEC 120 Update-letter LIGHT one-shot ceiling, slot wait and provider call together (update_letter.py)
OUROBOROS_FALLBACK_COOLDOWN_ENABLED true 429-aware per-process model cooldown
OUROBOROS_FALLBACK_COOLDOWN_SEC 120 Cooldown window
OUROBOROS_FALLBACK_ATTEMPTS_PER_MODEL 1 Attempts per model in the fallback walk
OUROBOROS_REVIEW_NATIVE_MAX_TRANSCRIPT_CHARS 900000 Owner CEILING (chars) on the native review inspection episode transcript; the effective bound is the reviewer window's calibrated capacity, never above this — except that a surface's declared mandatory reading (the advisory's five governance documents) lifts it up to the window, disclosed as native_mandatory_read_exceeds_bound when even the window cannot hold it. No round cap exists (OUROBOROS_REVIEW_NATIVE_MAX_ROUNDS is retired): exhaustion is a typed fail-closed refusal for verdict shapes and a disclosed incomplete product for the report shape, never a silent truncation
OUROBOROS_MODEL_DEEP_SELF_REVIEW openai/gpt-5.6-sol Deep self-review model key — the invisible migration source and fallback for the optional deep_review reviewer row: with no row saved, deep_review_slot() synthesizes the packed api row from it (the historical delivery); a saved row wins and the key is not read — so the provider-default migrations of this key (server_runtime.py) reach only installs that still synthesize from it; a row-configured install keeps its row, by design. Not a Settings UI field any more — the row lives in Agents → Review lanes
OUROBOROS_MAX_WORKERS 10 Active worker dispatch capacity; required owner waits retain additional sleeping processes
OUROBOROS_MAX_ACTIVE_SUBAGENTS_PER_ROOT 6 Live-subagent cap per root (hard cap 500 ids; depth hard cap 10)
OUROBOROS_MAX_SUBAGENT_DEPTH 3 Subagent tree depth
OUROBOROS_DISABLE_MANAGED_UPDATES (unset) Env-only: 1 disables managed updates (git_ops.py)
OUROBOROS_ALLOW_MUTATIVE_SUBAGENTS (empty) Mutative-subagent Auto override
OUROBOROS_SUBAGENT_WORKTREE_ROOT (empty) Acting-worktree root (empty derives ~/Ouroboros/subagent_worktrees)
OUROBOROS_SUBAGENT_PROJECTS_ROOT (empty) Projects root (empty derives ~/Ouroboros/projects)
OUROBOROS_SUBAGENTS (empty) Canonical configured-subagent roster (configured_subagents.py; §6)
OUROBOROS_SUBAGENT_HARNESS (empty) Legacy narrow harness input
OUROBOROS_SUBAGENT_PROFILE (empty) Legacy narrow profile input
OUROBOROS_DELEGATE_WAIT_SEC 120 Default delegate_wait window
OUROBOROS_DELEGATE_WAIT_MAX_SEC 1800 delegate_wait ceiling
OUROBOROS_DELIVERABLES_ROOT (empty) Deliverables root (empty derives the ~/Ouroboros/Deliverables sibling; tool_access.py)
OUROBOROS_GC_RETENTION_DAYS 7 Unified GC retention (retention.py)
OUROBOROS_RESTART_DRAIN_MAX_SEC 120 Restart drain bound
TOTAL_BUDGET 200.0 Global budget (USD)
OUROBOROS_PER_TASK_COST_USD 50.0 Per-task cost cap; also the tree ceiling basis (task_pacing.py): the root resolves min(global share, cap minus margin), enabled descendants retain that original threshold while actual monetary admission remains independent, and the wrap-up affordability rail soft-lands under it
OUROBOROS_RUB_USD_RATE (empty) Manual RUB→USD rate for RUB-priced providers
OUROBOROS_PRICING_TTL_SEC 21600 Provider-catalog pricing cache TTL
OUROBOROS_TOOL_TIMEOUT_SEC 600 Default tool timeout
OUROBOROS_PER_CALL_TIMEOUT_CEILING_SEC 1800 Per-call timeout clamp
OUROBOROS_FINALIZATION_GRACE_SEC 120 Finalization grace window
OUROBOROS_WEBSEARCH_MODEL gpt-5.2 web_search backing model
OUROBOROS_WEBSEARCH_BACKEND auto web_search backend selection
OUROBOROS_MAIN_WEB_SEARCH off Main-loop inline web search
OUROBOROS_MAIN_WEB_SEARCH_ENGINE auto Inline-search engine
OUROBOROS_MAIN_WEB_SEARCH_MAX_TOTAL_RESULTS 10 Inline-search result cap
OUROBOROS_OR_PROVIDER "" OpenRouter provider-routing preference merged into requests
OUROBOROS_SEARCH_CODE_WALL_SEC 45 search_code wall-clock budget
OUROBOROS_PRESENTATION (unset) Env-only: launcher-exported presentation (desktop_window/browser_fallback; absent renders web)
OUROBOROS_USER_FILES_ROOT "" (home) Env-only: user_files jail root (empty = $HOME)
OUROBOROS_OBSERVABILITY_KEEP_RAW unset Env-only: truthy enables raw observability payload persistence
OUROBOROS_GENERATIVE_PROBE 1 (on) Generative-write probe toggle
OUROBOROS_GENERATIVE_PROBE_CHARS 5000000 Generative-probe size companion
OUROBOROS_REVIEWER_SLOTS (empty) (6.1) Structured reviewer-slot SSOT (reviewer_slot_config.py): JSON {triad[], scope[], advisory, deep_review?}; each row is EITHER an inline route {slot_id, route:{kind: api_chat|agent_session, target_id}, effort} OR a roster reference {slot_id, subagent_id, effort} — mutually exclusive (a row naming both refuses typed; the reference materializes route/effort from the Available-subagents roster at load time, an explicit row effort winning over the roster row's) — with a STABLE owner-assigned slot_id (never an array index); an agent_session or managed-model api_chat route may add the optional route.profile_id credential pin (empty = account rotation; direct API-key routes reject account pins); the optional deep_review singleton carries the same row keys minus slot_id (fixed deep_review_slot_1) and, absent, is synthesized as the packed api row from OUROBOROS_MODEL_DEEP_SELF_REVIEW. Empty = read the legacy comma keys + phase-5 route envs as the migration source. Malformed value refuses typed at save AND at review time on every surface, task acceptance included (owner R3); env-apply logs and leaves legacy keys unprojected. The save that FIRST gives the triad a retrieving row (agent session or configured-subagent native inspection) returns the one-time R12 migration disclosure in the save response's warnings — the rows by id and target, and the measured API packet-panel cost it replaces (≈12 s / ≈$0.07 per model row per task, median of the 2026-09-01 OSWorld traces; ≈75 s / ≈$0.82 for a three-row panel on ProgramBench) against minutes of subscription window per task for a session row; a later save that keeps a retrieving triad is silent, and so — by design — is the reverse transition back to a packet-only triad: R12 is a one-time migration notice, not a routing monitor. The onboarding ladder footnote states the same numbers.
OUROBOROS_SUBSCRIPTION_PRESET_VERSION (empty) One-shot install-preset marker; endpoint-authored, DISK-ONLY (ENDPOINT_AUTHORED_SETTINGS) — its absence authorizes nothing, which is why install time is proved by three facts (§2)
OUROBOROS_SUBAGENT_PRESET_RECEIPT (empty) Install-preset receipt; endpoint-authored, disk-only
OUROBOROS_ONBOARDING_COMPLETED_AT (empty) Durable completion fact; endpoint-authored, disk-only
OUROBOROS_TASK_REVIEW_MODE auto Task acceptance-review mode
OUROBOROS_SAFETY_MODE full Safety supervision mode (shipped default full; a fresh desktop wizard may author light); lowering is owner-guarded (the runtime-mode boundary: §6 Safety and runtime mode)
OUROBOROS_SAFETY_MAX_TOKENS 2000 Safety-check output budget
OUROBOROS_SAFETY_CALL_TIMEOUT_SEC 60 Safety-check call ceiling
OUROBOROS_WEBSEARCH_TIMEOUT_SEC 480 web_search ceiling
OUROBOROS_DIRECT_TURN_STOP_WAIT_SEC 2 Seconds the chat lane waits for a stopped direct-chat turn to reach its next round boundary after the typed finalize_now control is armed; past it the outcome is live and the supervisor sweep retries custody rather than publishing a terminal (clamped 0-10)
OUROBOROS_ONBOARDING_SNAPSHOT_TIMEOUT_SEC 45 Bound on the onboarding transaction's settings snapshot read, so a wedged filesystem refuses the step instead of hanging the wizard
OUROBOROS_SETTINGS_DOCUMENT_LOCK_TIMEOUT_SEC 30 Bound on acquiring the settings-document lock for an owner read-modify-write; a timeout REFUSES before the transform runs — the lock is a precondition of the write, never a hint. The same bound also caps the initiating writer (_run_settings_writer: one lock wait plus one held episode; the generic save, the owner endpoints and onboarding completion alike): past it the Save answers 503 settings_save_timeout with saved: null and the body is left to its thread
OUROBOROS_LLM_TRANSPORT_READ_TIMEOUT_SEC 2700 LLM transport read timeout
OUROBOROS_PLAN_TASK_DEADLINE_MIN_SEC 300 plan_task deadline floor
OUROBOROS_ACCEPTANCE_REVIEW_EST_SEC 200 The floor: the minimum spendable seconds above the finalization reserve required to START an acceptance panel; never below 200 s (a smaller value is raised to it, a larger one wins). The improvement window is this floor ×2 under the adaptive improvement policy and ×1 otherwise. A spendable window at or below the applicable line is refused review_skipped_deadline_reserve / improvement_window_inside_reserve. The host predicts no review duration (owner R52): once launched, the review is clamped to the owner deadline and the task ceiling (R23) with the per-send money fence, and panel durations are telemetry only.
OUROBOROS_REVIEW_MAX_CYCLES "2" Shared paid review-cycle cap across the plan/acceptance/commit/skill gates (unlimited = no local count cap; per-gate semantics: review_cycles.py docstring, §10)
OUROBOROS_ACCEPTANCE_MAX_IMPROVEMENT_PASSES (retired) Retired alias: a stored value is MIGRATED into OUROBOROS_REVIEW_MAX_CYCLES (passes + 1) at settings load; a leftover env value is inert
OUROBOROS_ACCEPTANCE_RESERVE_PCT 5 Acceptance budget reserve percentage
OUROBOROS_OBSERVABILITY_RETENTION_DAYS (retired) Retired in 7.0 (RETIRED_SETTING_KEYS): observability rows are preserved indefinitely and the reader never deletes, so a retention knob had no reader; a stored value is stripped at settings load and a leftover env value is inert
OUROBOROS_REVIEW_MODEL_TIMEOUT_SEC (unset) Env-only: logical review timeout (absent = route-owned behavior; late in-flight results stay in custody)
OUROBOROS_REVIEW_MAX_TOKENS 65536 Env-only: reviewer output budget, clamped to the 8192 floor
OUROBOROS_REVIEW_ENFORCEMENT advisory Review enforcement: advisory/blocking (closed enum; anything else coerces to the default)
OUROBOROS_PREFLIGHT_TIMEOUT_SEC 900 Env-only: TOTAL wall-clock budget for the hermetic pre-commit pytest preflight (node lane + both passes; teardown + containment semantics in preflight_runner.py/process_containment.py)
OUROBOROS_PREFLIGHT_SERIAL unset Env-only: 1 selects one serial pytest pass; scrubbed from the candidate environment
OUROBOROS_AUTO_GRANT_REVIEWED_SKILLS true Auto-grant manifest-declared permissions to cleanly reviewed skills (hash-bound; blocking findings never grant)
OUROBOROS_TRUST_NATIVE_SEEDED_SKILLS true Launcher seed/resync writes hash-pinned native_seed verdicts; acts only at seed/resync, no runtime grant endpoint
OUROBOROS_CONTEXT_MODE max Owner context mode (max/low); also decides scope-review applicability — an explicit owner policy coupling, not an inferred model limitation (BIBLE P1/P3); owner routes/CLI only
OUROBOROS_CONTEXT_MODE_AUTO_LOW false Task-local low-mode overflow retry toggle
OUROBOROS_RUNTIME_MODE advanced Runtime mode light/advanced/pro/cyber_pro — a compatibility/self-modification boundary orthogonal to review enforcement; light blocks registry mutation, mutative git/writer argv, and self-elevation; advanced blocks protected core/contract/release paths; pro permits protected rewrites with the normal review notice; cyber_pro extends the same rank-aware execution seam to owner configuration and selected host setup; protected BIBLE/history deletion remains separate; the owner endpoint persists the next-boot value
OUROBOROS_SKILLS_REPO_PATH "" Extra skills checkout path (expanded at read time, never cloned/pulled)
MCP_ENABLED false MCP client toggle (§6 MCP)
MCP_SERVERS [] MCP server list (HTTP/SSE via URL/auth, stdio via command+args and optional cwd/literal/settings-backed env); persisted in settings, never env-exported
MCP_TOOL_TIMEOUT_SEC 60 Per-MCP-tool timeout
OUROBOROS_HUB_CATALOG_URL https://raw.githubusercontent.com/razzant/OuroborosHub/main/catalog.json OuroborosHub catalog URL (automatic fetch limited to catalog JSON; installs verify SHA-256)
OUROBOROS_CLAWHUB_REGISTRY_URL https://clawhub.ai/api/v1 ClawHub registry URL
OUROBOROS_PROMPT_CACHE_TTL 1h Prompt-cache tier (default/5m/1h). The policy acts at the final send-time wire boundary so it can legalize provider ordering without prompt builders creating provider-specific TTL policy; review_helpers.cached_prompt_blocks and usage_accounting._reservation_cost also consult it; usage records the applied tier
OUROBOROS_EFFORT_TASK medium Task reasoning effort (scale none/minimal/low/medium/high/xhigh/max/ultra; Settings exposes all but minimal); provider adaptation is exact-route, success-confirmed, disclosed in request_wire
OUROBOROS_EFFORT_EVOLUTION high Evolution effort
OUROBOROS_EFFORT_REVIEW high Review effort
OUROBOROS_EFFORT_SCOPE_REVIEW high Scope-review effort
OUROBOROS_EFFORT_DEEP_SELF_REVIEW high Deep-self-review effort — the surface default; a saved deep_review row's own effort outranks it
OUROBOROS_EFFORT_CONSCIOUSNESS high Consciousness effort
OUROBOROS_RETURN_REASONING true Ask OpenRouter to return reasoning; direct/local request copies strip OpenRouter-only fields
OUROBOROS_REASONING_SUMMARY auto Readable reasoning-summary rendering; presentation-only, never added to history or returned to providers
OUROBOROS_TASK_IDLE_TIMEOUT_SEC 900 Idle timeout — requires absence of real task/subtree progress; the typed in-flight main-LLM row spares only this rail; a settled child result stamps parent progress, because delivery creates immediate integration work and must not coincide with idle termination
OUROBOROS_TASK_ABS_CEILING_SEC 21600 Absolute task ceiling, activity-independent; deadline and budget stay separate hard axes
OUROBOROS_SUPERVISOR_LIVENESS_DEADLINE_SEC 90 Supervisor/direct-turn liveness watchdog — alerts and recommends /restart but never frees an in-process lock held by a genuinely wedged turn
OUROBOROS_PACING_INTERVAL_SEC 600 Pacing reminder interval
LOCAL_MODEL_SOURCE "" Local-model source (HF repo or path)
LOCAL_MODEL_FILENAME "" Local-model GGUF filename (split first-shard expanded)
LOCAL_MODEL_CONTEXT_LENGTH 16384 Local-model context window
LOCAL_MODEL_N_GPU_LAYERS 0 GPU offload layers
USE_LOCAL_MAIN false Route Main locally
USE_LOCAL_HEAVY false Legacy migration input only; excluded from active routing
USE_LOCAL_LIGHT false Route Light locally
USE_LOCAL_CONSCIOUSNESS false Route Consciousness locally
USE_LOCAL_FALLBACK false Route fallback locally
OUROBOROS_MAX_ROUNDS 200 Max task rounds (hot-reloadable)
OUROBOROS_TRANSIENT_RETRY_MAX 6 Same-model transient retry budget; pre-dispatch transport_unavailable is a separate task-bounded outer wait episode, because no provider attempt was admitted on that route
OUROBOROS_SKILL_LIFECYCLE_TIMEOUT_SEC 1800 Skill lifecycle-lane timeout
OUROBOROS_CLAUDEXOR_HARNESS_INSTALL_TIMEOUT_SEC 300 Harness install ceiling (kills the tracked group, typed refusal)
OUROBOROS_CLAUDEXOR_QUOTA_REFRESH_TIMEOUT_SEC 90 Quota-refresh POST ceiling (clamped 1–90)
OUROBOROS_BUNDLE_DIR (unset) Env-only: launcher-owned bundle root propagated to embedded children for Node/ripgrep discovery
OUROBOROS_BG_MAX_ROUNDS 10 Background-consciousness round cap
OUROBOROS_BG_WAKEUP_MIN 30 Consciousness wakeup floor (s)
OUROBOROS_BG_WAKEUP_MAX 7200 Consciousness wakeup ceiling (s)
OUROBOROS_POST_TASK_EVOLUTION false Post-task evolution promotion toggle; agent self-enablement is blocked at the shell/browser/settings/data-write guards; choosing an objective routes through the Main slot because it is a high-leverage decision, while execution stays behind ordinary review and owner gates
OUROBOROS_POST_TASK_EVOLUTION_CADENCE llm Promotion cadence llm or every_n:k (malformed normalizes to llm)
OUROBOROS_POST_TASK_EVOLUTION_BUDGET_USD 0.0 Remaining-global-budget start floor, not a cycle cap
OUROBOROS_EVOLUTION_PERSISTENT_OBJECTIVE "" Owner-only persistent campaign bias; still passes review gates
LOCAL_MODEL_PORT 8766 Local-model server port
OUROBOROS_HOST_SERVICE_PORT 8767 Host Service port (loopback-only; §12)
OUROBOROS_PRESENCE_MAX_ACTIVE 2 Cross-process Presence turn cap (UI-bounded 1–20)
LOCAL_MODEL_CHAT_FORMAT "" Local-model chat template override
GITHUB_TOKEN "" GitHub token (push/PR/issues)
GITHUB_REPO "" Personal origin repository
OUROBOROS_FILE_BROWSER_DEFAULT "" File Browser default root (explicit root required for Docker/non-localhost)

Direct-provider review fallback (legacy name: OpenAI-only review fallback): when exactly one official direct provider is configured, config.get_review_models() compiles that provider's declarative reviewer-role sequence using provider-prefixed model IDs. Current scope covers official OpenAI, Anthropic, MiniMax, DeepSeek, Cloud.ru, and GigaChat; OpenRouter, legacy-base, OpenAI-compatible, and mixed-provider configurations stay outside it. OpenAI, Anthropic, and DeepSeek run three independent Main-model slots; MiniMax mixes Main/Light; Cloud.ru and GigaChat use their one role model for every slot. _exclusive_direct_remote_provider_env returns empty when OpenRouter, legacy OPENAI_BASE_URL, OpenAI-compatible keys, or multiple official direct providers are present, and the fallback requires provider_models.migrate_model_value to make the main model already start with the exclusive provider prefix — exact prefix checking prevents an arbitrary free-text model from silently entering a single-provider route. This is part of the single-provider independence invariant (docs/DEVELOPMENT.md "Provider Independence").

DeepSeek provider specifics (deepseek::): the official OpenAI-compatible endpoint is a fixed module constant (provider_models.DEEPSEEK_BASE_URL); a proxy or mirror belongs to the generic openai-compatible:: route. The canonical reasoning scale is projected onto the provider's wire dialect at the send boundary (minimal→low, medium/xhigh→high, ultra→max, none→extra_body.thinking.type=disabled; native tiers pass through), a forced tool choice (required or a named tool) is served with thinking disabled because thinking mode accepts only auto/none (probed 2026-09-03), and every tier-changing projection is disclosed on usage; reasoning_content stays on canonical assistant turns for strict v4 tool replay with an explicit empty string for turns produced without provider reasoning, while other lanes strip the field and cross-family switches scrub it. System/assistant/tool content arrays are flattened to strings in the send copy only (the API accepts arrays on user turns alone); canonical block history is untouched. Prompt caching is automatic and cost remains nullable when no exact provider catalog is available. The 1M context claim is admitted only through route-fingerprinted capability evidence or owner acknowledgement. Slash-form deepseek/... remains OpenRouter; only deepseek::... selects the direct route.

GigaChat provider specifics (gigachat::): routed through the native gigachat library, not OpenAI-compatible (llm.py::_chat_gigachat). OpenAI tools map to GigaChat functions; at most ONE function_call returns per turn, so parallel tool_calls collapse to the first; role tool results become role function and must be valid JSON (plain text wrapped as {"result": ...}); the system message must be first, so later system-reminders demote to user. reasoning_effort is deliberately omitted — hidden reasoning can consume the whole output budget and return empty content. GigaChat exposes no automatic live cost source, so cost stays nullable/unknown rather than a hand-maintained tariff. GigaChat models sit below the 1M scope-review floor; when no ≥1M reviewer is configured, the declared alternatives are the owner-selected low context mode (whole-repo scope review declaredly not performed; each commit records a typed skipped_low_context_mode evidence row) or an owner-selected retrieving scope slot at ≥200K sourced evidence (BIBLE P3); the blocking triad still reviews the full staged diff in both modes.


8. Git Branching, CI, and Build

ouroboros is the local working branch; the runtime setting independently selects one official feed: Stable is the newest plain release tag reachable from both main and ouroboros-stable, QA is the ouroboros-stable tip, Development is the ouroboros tip. Promotion and rollback are owner-controlled exact-SHA movements; ordinary restart preserves the local tip, while explicit update owns fetch, target validation, rescue, apply, and rollback. managed is the official read/update remote; origin is optional personal persistence. Desktop and Colab reject a target without a regular non-empty BIBLE.md. External pull requests target ouroboros, do not allocate a release version, and receive the collision-free version when maintainers land and re-review them (procedure: CONTRIBUTING.md, docs/DEVELOPMENT.md).

The local ouroboros-stable ref is also a recovery fallback maintained by explicit promotion; that local role does not select the official QA feed. Launcher metadata is bootstrap provenance only — runtime status, preflight, Colab bootstrap, and apply resolve the selected channel and exact fetched SHA themselves, so an older frozen launcher cannot silently redirect updates. Stable additionally requires the shared plain release tag; QA and Development do not use version comparison as an admission gate.

CI topology

.github/workflows/ci.yml has 16 jobs, grouped by trigger and secret exposure:

Tier Jobs Trust boundary
Fork-safe PR validation quick-test, betterleaks-platform-smoke, benchmark-methodology (also on ouroboros-stable pushes and v* tags) ouroboros pushes, pull requests into ouroboros, manual runs; read-only, no provider secrets, never pull_request_target
Browser consumer proof ui-smoke PRs run only the real Publish admission/card/history scenario on Chromium; manual runs and tags retain full host UI and browser-tool Chromium/WebKit coverage. No provider secrets or added default local pytest lane.
Scheduled lanes system-e2e-mock (keyless; daily 04:37 UTC, manual runs, v* tags — on the release bar), e2e-live (OUROBOROS_E2E_LIVE_OPENROUTER_KEY, $30 cap) e2e-live fires only on its own nightly cron (03:17 UTC, seeding the ouroboros branch tip rather than the default branch the schedule fires on) or a dispatch that opts in through the e2e_live input; without the secret it is one summary line and green, never a pretend run
Stable/tag matrix full-test (no secrets), skill-smoke (OPENROUTER_API_KEY) ouroboros-stable pushes, manual runs, v* tags — never pull requests
Trusted provider run integration-test provider secrets; main/ouroboros/ouroboros-stable pushes, manual runs and v* tags (the release chain needs it) — never pull requests
Tag-only gates and release chain marker-guards, docker-ui-smoke, docker-portable-test (manual runs or tags, no secrets); release-preflight (needs full-test + integration-test) → build (signing secrets) → release (needs build, release-preflight, skill-smoke, ui-smoke and these three smoke gates); vendor-package-smoke needs build and stays informational tag-triggered; a reproducible provider-contract failure blocks tag builds, a typed inconclusive provider outage does not

Quick and full jobs each run a dedicated blocking size_ratchet pytest step — the ONLY enforcing surface for the repository size gates (local runs exclude the marker and warn): manifest exactness on the tip plus the pairwise shrink-only transition against the event base in OURO_SIZE_RATCHET_BASE_REF; an unresolvable base degrades to the tip's parent manifest verified against the parent's own tree — never a skip — while a resolvable base without a manifest fails closed. Both jobs also run the browser-module suite (cd web && node --test tests/*.test.js), the same node lane the hermetic commit gate executes through ouroboros/preflight_node.py. Secret-bearing skill review runs before any step that imports downloaded plugin code — untrusted payload code must never share a process with provider credentials — and a missing required key is red rather than skipped.

Tool-schema compatibility has two layers. Fork-safe PR tests build the complete shipped catalog without MCP or extensions, validate every schema as general JSON Schema plus the cross-provider subset (no empty enum, no root anyOf/oneOf/allOf), and require the OpenRouter/function, direct-Anthropic, GigaChat, and direct-OpenAI projections to preserve the complete tool-name set. The trusted integration-test lane sends that same full registry in one bounded delegate_start canary per physical route without executing the returned call — OpenRouter Gemini/Opus/Fable/GPT/Grok/DeepSeek, the three shipped direct-OpenAI defaults (Main alone among them keeps a second-turn nonce-bearing continuation), direct Anthropic, and optional MiniMax/DeepSeek/Cloud.ru/GigaChat (DeepSeek also keeps the continuation because its tool contract is the reasoning_content replay). Each canary requires positive usage, exact provider/model identity, and a normalized schema-valid tool call; quota/billing, 429, 5xx, and timeout outcomes stay typed inconclusive while contract/auth/model/tool/reasoning 4xx are red; one same-route resend is allowed only for a runtime-classified semantic-empty response (bypassing response caches where supported). response_finish_reason is retained in host usage for bounded diagnostics only; malformed raw arguments are reported by type/position and hash without copying the provider payload.

claudexor-platform-gate.yml proves the managed Claudexor runtime on three OSes: a fixture lane (fake harness, offline, $0) always, and a live lane only on explicit API keys — subscription auth stays deliberately out of CI, because an interactive machine-bound token must not enter CI secrets. dependency-graph.yml reads ouroboros/claudexor_runtime_pin.json and submits that direct runtime relationship to GitHub's dependency graph (runs only on pin/workflow changes on main/ouroboros, plus manual dispatch; contents: write only) without presenting Claudexor as a Python or Node package dependency. The Scorecard workflow (.github/workflows/scorecard.yml) runs on main pushes and weekly, pins every action by full commit SHA, defaults permissions to read-only, and adds only security-events: write + id-token: write for SARIF upload and OpenSSF publication. CODE_OF_CONDUCT.md owns community rules; CITATION.cff owns the software and preferred technical-report citations; site/paper/index.html owns the canonical paper landing page; docs/benchmarks/evidence.json holds the release-bound public benchmark projection; README remains the claim SSOT (guarded by tests/test_trust_metadata.py and tests/test_public_site_metadata.py).

Dev-facing release/review scripts: scripts/run_external_review.py (dual-lane external review; --contributor binding; READY_FOR_INTEGRATION is evidence, never merge authority), scripts/contributor_review_evidence.py (route-neutral contributor-packet binding), scripts/run_plan_review.py (the same engine as plan_task; review-exempt dev tool), scripts/validate_scope_receipt.py, scripts/claudexor_platform_smoke.py, scripts/fetch_claudexor_runtime.py (pin SSOT: claudexor_runtime_pin.json), scripts/cleanup_test_pollution.py (dry-run-first). site/ is the Vite source of the public pages (site/scripts/sync-assets.mjs syncs assets); skills/telegram/ and skills/unix_computer_use/ are the bundled skills; packaging/cli/ holds the CLI wrappers and installers; packaging/appimage/ the AppRun dispatch + desktop metadata; packaging/systemd/ the opt-in user unit (§1 Runtime topology owns its no-restart rationale).

System E2E suite (tests/system_e2e/)

The deep-integration suite drives a REAL server.py per scenario — an isolated repo clone, data root and free loopback port through KeylessIsolatedServer, whose child environment can never carry a provider credential — and asserts DURABLE artifacts through ArtifactOracle readers (task_results/, logs/*.jsonl, state/*), never an HTTP 200 or a harness exit code. Synchronization is durable-event polling (wait_until over those readers), never bare sleeps. A scenario that must act MID-ROUND (the owner stopping a turn while the model is still answering) holds the loopback model on an event-gated ModelGate — the request thread blocks until the scenario releases it — instead of racing a latency sleep. Two loopback models share one HTTP base: ScriptedStubModel, an ordered per-scenario tool script whose review-organ calls are classified by prompt markers and answered with canned parse-clean verdicts before the finalization-turn check, and ReplayModel, deterministic fixtures bound by (lineage, slot, attempt) where a miss or an unconsumed row is red. A scenario may hand the stub a ReviewScript — ordered per-kind verdict queues that override the canned all-clean answers for exactly the scripted red rounds and then fall back — and review calls never consume agent script steps in either direction. The prompt markers are VERBATIM literals from the tree under test, greppedback out of their source files by a default-lane pin, so drift is a named failure rather than a silently mute stub. Reviewer routing is pinned through the structured OUROBOROS_REVIEWER_SLOTS, because the retired comma keys are dropped at load and would fall back to the shipped paid default panel. FakeClaudexorDaemon is a loopback daemon imitation serving the exact gateways/claudexor.py client contract — the authenticated handshake reporting the TREE'S OWN runtime-pin identity, the capability/quota answers route_health reads, idempotency-keyed project registration and run creation with the engine's replay semantics, run detail and cancel — recording every wire request for scenario assertions.

The scenario inventory is DATA (SCENARIOS in harness.py) with a two-direction pin scanning every test_*.py module of the package: a manifest row without a test is red, and a scenario test without a manifest row is red. It covers boot and identity, the review organs (commit triad plus scope in both enforcement classes, plan review, the acceptance loop), egress hardening, typed tool safety, honest cost names, subagent trees, cancellation, its cascade and the owner stop of an in-process direct-chat turn, managed update (fast-forward, carrier transfer, conflicting refusal, crash mid-apply, rollback), delegated transport including the mutating pull-in and its refusal, and the whole skills lifecycle from payload to delete. Scenario tests carry BOTH the integration and serial markers plus the OUROBOROS_E2E_DEEP=mock env gate, so neither the default local run nor either CI pytest pass executes them; the keyless lane is OUROBOROS_E2E_DEEP=mock pytest tests/system_e2e/ -o addopts="", and in CI it is the daily scheduled system-e2e-mock job — schedule or manual dispatch only, secret-free — the one lane those three gates leave open.

Build scripts

build.sh (macOS → .dmg), build_linux.sh (→ .AppImage + .tar.gz), scripts/build_appimage.sh, scripts/build_linux_packages.sh (.deb/.rpm/.red80.rpm), scripts/smoke_linux_packages.sh, build_windows.ps1 (→ .zip), and scripts/build_repo_bundle.py (writes repo.bundle + repo_bundle_manifest.json, the manifest launcher.py bootstrap validates) are the release-invariant owners. Linux PyInstaller runs under the same pinned portable Python shipped in the payload, so its bundled libpython keeps the payload's glibc floor instead of inheriting the release runner's newer ABI. The AppImage builder wraps that payload with digest-pinned tool and runtime bytes; the native Linux builder wraps the same x86_64 payload without replacing its runtime. Native package metadata declares the external Git required by bootstrap while bundled Python/Node/browser stay under /opt/ouroboros; the packages install the opt-in user unit at /usr/lib/systemd/user/ouroboros.service with no activation scriptlet. The release-gating package smoke installs through apt/dnf and proves Git resolution, desktop files, the installed user unit's launcher/cgroup/no-restart contract, the packaged CLI, and a bounded desktop-launcher start on Ubuntu 22.04/Fedora 42; Astra Linux and RED OS vendor-image runs stay informational evidence, because third-party registry availability cannot block publication. The macOS image keeps the app + Applications symlink + optional CLI installer layout. Release tag prerequisite: scripts/build_repo_bundle.py is the release-tag SSOT and verifies the annotated v$(cat VERSION) tag points at HEAD before packaging.

Betterleaks follows the same release-resource discipline: packaged builds stage the pinned binary and license under betterleaks-standalone, source checkouts install the exact managed runtime explicitly under data/state/betterleaks/, and the Publish path never downloads it (§13).

Python dependency resolution has one authority: direct requirements and group membership live in pyproject.toml, uv.lock records the universal solution under the pinned tool.uv.required-version, and source/CI sync with --locked so metadata drift is an error. The two-interpreter packaging boundary rides projections: build scripts export their PyInstaller input directly from uv.lock, the committed requirements-runtime.lock supplies embedded python-standalone and managed updates (pip, no bundled uv), and a one-line requirements.txt pointer serves already-released updaters; CI regenerates the export and requires a clean diff, so neither file becomes a second dependency authority.

Platform builds precompile bundled Python with unchecked-hash bytecode: sealing valid bytecode prevents runtime __pycache__ writes from invalidating a macOS signature, and runtime children route caches outside the bundle. When signing is enabled, hardened runtime, notarization, xattr hygiene, and strict verification remain part of the stable-release path; prerelease artifacts may be unsigned and their evidence must report the actual signing state.

The release proof begins with the final DMG/AppImage/tarball/ZIP, never its staging directory — validating a staging tree does not establish that the published bytes contain or execute the same payload. Each platform shard checks the embedded repository bundle, packaged CLI, and managed Claudexor seed + Node by starting the owned daemon, completing a fixture task, and verifying an identity-bound stop. The AppImage is extracted for metadata/SBOM inspection, then run FUSE-independently to prove version output, CLI dispatch, browser-fallback readiness, payload lifetime, main-executable libraries, and clean shutdown (browser-fallback evidence, not a claim of a native GTK/Qt backend); the nested cleanup proof follows the live runtime → AppRun custodian → launcher chain and requires both the extraction and its private base absent. The proven Linux tarball payload is wrapped into the three native packages, each receiving its own digest-bound package-manager smoke receipt and provenance attestation; a digest-pinned Syft build produces CycloneDX inventories from extracted payload bytes, and the tarball inventory is reused for the identical-byte native wrappers while each wrapper keeps its own digest-bound installation proof. The release job accepts only the seven expected assets, recalculates digests, verifies both predicate types, writes the checksum/evidence capsule, and rechecks the remote annotated tag immediately before publication. Signing credentials stay step-scoped and absent from SBOM/attestation steps.

Public installer naming and links ride the same projection: release_sync.py::RELEASE_ASSET_TEMPLATES is the filename SSOT shared by the proof builder, README, and the install pages. A version bump rewrites only named download references and data-release-download anchors to immutable /releases/download/v{VERSION}/... URLs; the /releases/latest/download/... shape is forbidden because GitHub excludes prereleases from latest. Generated release notes expose direct links only for the seven proof-accepted assets. The default README and GitHub Pages deployment use the stable main boundary (main:/docs for Pages); stable promotion advances main only after the release is published with all seven proof-bound installers, so an unreleased development VERSION never exposes dead installer links and an omitted promotion leaves the previous working release public.

Docker

Docker runs the web and server runtime without PyWebView. The image binds 0.0.0.0 and sets no network password by default: a missing password only warns, and NetworkAuthGate permits requests when no password is configured — publishing the container port without setting OUROBOROS_NETWORK_PASSWORD therefore exposes the owner surface. Set the password (or keep the port unpublished); packaging adds no stronger boundary of its own.


9. Shutdown & Process Cleanup

Ordinary window close preserves this installation's shared Claudexor daemon and work submitted by other clients, on graceful and forced launcher cleanup paths. It stops the server generation's workers/services, waits for its process group or Job, performs owned cleanup and releases the PID lock. The Windows daemon is outside the permitting launcher Job; retained-subtree cleanup also covers worker/server ancestors and the POSIX stray reaper (§1). An old immutable launcher requires a later package update for this guarantee. Same-pin planned restart retains the daemon. A planned restart after a changed engine pin instead invokes server_restart._stop_owned_daemon_for_new_pin in lifespan teardown: compare the landed load_runtime_pin with the serving authenticated engine through attach-only read_owned_gateway, then use the explicit stop below. Unpublished/unreadable pins and unreachable/foreign/stopped daemons do not alter the handoff. Unconfirmed stop retains custody and diagnostics while restart proceeds; the next generation reconciles old runs without inventing spend. Ordinary server shutdown closes its own Host Service listener. Existing reserved-port sweeps remain separate launcher/recovery/Panic behavior; they are not PID-identity proof and can terminate an unrelated listener on a reserved port.

The owner's manual Restart (/restart, the chat Restart button) is a clean stop of everything the current server generation owns, then the re-exec. The bridge handler keeps its checkout-first order: _safe_restart_serialized (the update lock and strict managed-update transaction gate, then safe_restart's checkout, dependency sync, import test and stable fallback) runs before anything is stopped, so every refusal — an assisted-update merge mid-resolution ("restart was deferred"), a failed checkout, an unwritable no-resume flag — leaves the server intact and is answered with one "Restart cancelled" line. Past that point the restart always follows: the durable no-resume intent (state/owner_restart_no_resume.flag plus the panic_stop.flag compatibility pair, consumed by auto_resume_after_restart), then server_restart._stop_owned_work — one durable cancel intent per RUNNING task, live direct/ephemeral activity and in-flight post-task synthesis (cancel_intents.request_cancel, source owner_restart, immediate), kill_workers(force, cancelled) with Panic's reconcile_delegate_custody=False, delegated-run cancellation through the public owner-gone seam reconcile_orphaned_runs(running_task_ids=set()) over read_owned_gateway (attach only: nothing between the cancel intents and the daemon stop may ensure_owned_gateway, which would start a dead daemon), and the attested owned-daemon stop_outcome() exactly as Panic makes it — then the owner's "Stopping active task" notice and the exit. Nothing there is a veto or a deferral: an unconfirmed or raising worker shutdown and an unconfirmed daemon stop each leave a critical diagnostic (the daemon's own process_stop_unconfirmed supervisor row, or the same row when the stop raises), custody stays retained, and the restart proceeds; a stop with nothing to stop is quiet. The remainder is the next generation's existing startup work: owner restart is a no-resume cause (delegate_recovery.NO_RESUME_CAUSES), so no handoff is prepared and nothing is adopted; the startup custody sweep reconciles every open delegated run as owner-gone against the daemon it then reaches (a run the stopped daemon took down answers absent → close_absent_run, no invented spend; a journal-recovered terminal settles; a pending invocation replays under its own key), the boot backfill heals stored disclosures, and an in-flight owned model operation stays a dispatched/unresolved physical-attempt row that already counts as spend upper bound. After an unconfirmed stop the re-exec'd generation attaches to the still-live daemon rather than spawning a second one, and the lifespan's one background warm_owned_daemon() (provisioned homes only) makes the first delegation after a Restart find the daemon serving. Planned self-restart, managed-update handoff and Panic keep their separate contracts; the planned restarts share only the attested daemon stop, and only when the landed checkout pins another engine (above). The explicit stop is installation-wide: manual Restart ends runs served by the owned daemon, including another client's runs. This differs intentionally from ordinary window close.

Panic is a complete explicit owner stop, not a restart: stop consciousness, record evolution owner-stop state, close the campaign and pending promotion, write panic_stop.flag, stop local models and call OwnedClaudexorDaemon.stop_outcome before worker cleanup. It then ends tracked foreground commands, executor processes, services, companions and workers before hard server exit. The shared daemon may belong to a previous generation and is outside the Windows launcher Job; its explicit stop, not closing that Job, owns termination. Failure emits a critical process_stop_unconfirmed row and retains unresolved custody while the rest of Panic continues. The next launch suppresses automatic work under the panic/no-resume flags.

stop_outcome returns stopped, nothing_to_stop, or disclosed unconfirmed. With a valid owned marker and authenticated same-home endpoint it first invokes the already installed managed CLI daemon stop --json, using read-only resolve_cli_command(require_npm=False) and an owned config root with ambient socket/port overrides removed. Operator shutdown needs the existing exact Node, not the npm installation toolchain; the default installer-facing resolver keeps that separate requirement. It never ensures a runtime, wakes a daemon or probes accounts. Exit0 plus the typed terminal CLI receipt is required; an RPC acknowledgement or clean lease release alone does not prove captured OS processes have exited. Known roots/children are observed through the exit window, and a surviving authenticated endpoint is disclosed without chasing a successor. The pre-request custody snapshot limits forced fallback to those original rows; forced signalling still requires both measured birth and command identity or the manager's own Popen. Legacy Windows empty-birth rows cannot authorize signals or be re-attested from a later PID observation, but healthy attached daemons can stop cooperatively through the existing CLI. Missing CLI or failed/unavailable shutdown retains the previous independently proven fallback. Token refusal, invalid discovery, incompatible/malformed responses and a matching error-code string alone never provide that authority; a typed HTTP transport failure or positively absent descriptor for marked startup retains its existing meaning. Only confirmed death permits signal-stop custody removal; concurrent rows survive.

The short manager lock retires the local startup generation before waits and never spans runtime preparation, network or process-exit work. CLI RPC, termination confirmation and startup slack use config.CLAUDEXOR_OPERATOR_STOP_TIMEOUT_SEC; captured-process observation uses config.CLAUDEXOR_STOP_EXIT_WAIT_SEC. These existing constants live in runtime_limits and are re-exported by config; each signalled ledger root retains its exit window. These phases do not promise an absolute total stop duration. Startup joins purpose-filtered custody before and after preparation; engine writer election remains authoritative. A live startup outlives its caller's window (20 seconds by default, independently narrowable): daemon_starting invites joining the same work, whereas daemon_spawn_failed identifies the exited child/build/log interval. The authenticated control handshake and the independent normal-admission wait remain distinct. Native Windows Job/identity/stop tests and frozen/embedded psutil imports require Windows/package evidence; portable fixture success alone does not establish that proof.

Bounded foreground run_command/run_script processes ride the in-memory _active_subprocesses panic registry; long-lived children — services, executor-backed processes, extension companions, delegated runtimes — enter durable exact-identity process custody via spawn_supervised (§1 Runtime topology). Unix process groups and Windows Job Objects provide tree cleanup; durable executor and service records let the host recover after worker death. Normal cleanup may archive logs, while panic skips nonessential finalization — agent-controlled or wedged cleanup must not delay an owner stop. Timeout and signal exits remain distinct in tool results so a killed command never resembles success.

10. Key Invariants

  1. Constitution and identity persist. BIBLE.md is never deleted; identity.md remains a physical file even when its content evolves.
  2. Release metadata has one projection. VERSION is canonical; ouroboros/tools/release_sync.py::version_carrier_desyncs() and sync_release_metadata() keep the PEP 440 form in pyproject.toml and the editable root entry in uv.lock, plus the author-facing version in web/package.json and both root entries of web/package-lock.json, web/modules/api_types.js::GATEWAY_CONTRACT_VERSION, the README badge, and this document's header. Changelog prose remains deliberate. Pull requests into ouroboros leave these carriers byte-identical to their target; integration assigns the release version.
  3. Configuration and messaging have single owners. ouroboros/config.py is the one IMPORT surface for paths and settings, and the vocabularies live in its leaves — settings_defaults (keys, shipped defaults and the retired-key lists), settings_scales (the closed clamps), model_slots (slot resolution and the frozen model target), review_model_routes (the API-pinned reviewer lists), runtime_limits (numeric knobs and their clamps) and settings_integrity (the verified read primitive) — so a new key belongs to a leaf, never to the facade; messages go through supervisor/message_bus.py; concurrent state transitions use the owning file lock.
  4. The attempt ledger is monetary authority. state/usage_attempts.jsonl records every physical model send. State, task, event, and UI totals are projections carrying attempt identity; unknown or unresolved cost never becomes false zero.
  5. Packaged bootstrap is manifest-bound. A packaged install verifies repo.bundle and its manifest once, then runs the managed checkout. Restart preserves its local tip; only explicit update applies an approved exact SHA.
  6. Shutdown is custody-complete (§9). Normal close verifies child death; panic stops all owned work — the installation's daemon by confirmed ledger identity — without allowing agent code to delay it; the unledgered-daemon and Windows residuals are disclosed.
  7. This document is the present-tense map. Structural owners, APIs, durable data, UI surfaces, and the rationale for non-obvious guards update here in the same commit as the code (documentation contract: docs/DEVELOPMENT.md; residue ratchet: tests/test_docs_sync.py); release chronology lives in git and README.
  8. Skill gates do not collapse. Discovery, deterministic preflight, content-hash-bound executable review, owner grants, dependency readiness, enablement, and execution remain separate. A PASS does not install dependencies, and enabled=true does not prove executable readiness.
  9. Startup rescue has one mutation owner. Supervisor recovery writes rescue evidence before reset or blocks while preserving the tree. Worker or agent construction remains warning-only and never stages or commits inherited dirt.
  10. Projection over replay. Interactive status, history, and cost reads are bounded, non-materializing projections; durable owners perform the one authoritative replay or terminal materialization.
  11. UI resources carry a disposer. Every subscription, listener, observer, timer, stream, and live page instance has explicit teardown; navigation does not leave hidden instances mutating visible or durable state.
  12. Frozen contracts extend explicitly. ouroboros/contracts/ is a versioned, backward-compatible ABI — typed shapes together with their parsing/normalization/policy helpers (§11). New capability extends the frozen shape or ships an explicitly versioned successor; existing consumers keep working.
  13. Provider wire adaptation stays exact-route and success-confirmed. Canonical history remains provider-neutral; typed physical projections may change values, fields, or a registered dialect on one provider/endpoint/API/model only. Failed candidates teach nothing durable, task-local cognition degradation never becomes future dispatch authority, and the physical-attempt ledger remains distinct from terminal request-wire history.
  14. Cancellation is intent-then-custody. A cancel intent never rides the canonical task status; no teardown begins before the durable, watchdog-replayable intent row exists, and a failed intent write refuses the cancel typed; the settle owner takes an exclusive claim BEFORE any custody mutation and every mutation is fenced by the claim generation (a probed-ALIVE claimant is never abandoned); a settled RESULT does not prove a dead WORKER, so live ownership gates every settled-target cancel while the worker-side snapshot check fails OPEN toward liveness; natural completion wins a late cancel and the stored terminal row survives byte-identical — the kill is about the process, never the result; the deliverable is durably registered as OWED before the intent settles, and a registration failure leaves the intent open for the watchdog; a cascade acts on lineage ROOTS (one cascade, one summary per tree) and settles only on its no-live postcondition. Owners: task_lifecycle.py, cancel_intents.py, cancel_publication.py, terminal_delivery.py (§5 for the flow).
  15. Cancellation registry corruption stays visible. In state/cancel_intents.json and state/terminal_deliveries.json an absent file reads as empty; enforcement reads disclose a malformed file or row and keep its on-disk bytes, and no mutation rebuilds an existing malformed store from a {} collapse. The shared utils.read_json_dict helper answers None for absence, unreadable or invalid JSON, and non-object JSON alike, and _read_evolution_campaign() collapses all of those to {}.
  16. Verification receipts reconcile by one typed identity key. Sameness is equality of the single (kind, value) key — never a match across kinds — so reconciliation is an equivalence that fails SAFE toward strictly fewer reconciliations; disclosed parts are never the comparison (_outcome_receipts.py).
  17. Review spend has one ceiling. Every paid review gate shares OUROBOROS_REVIEW_MAX_CYCLES; the per-gate meanings live once in §6 Review stack, the SSOT in review_cycles.py.

10.1 Continuity data-flow map

The table below is the canonical map for continuity changes. A bounded view is an interface projection, never a new authority. The actor that makes the decision must be able to resolve the named source through an existing reader; otherwise the view is partial and the consumer remains non-final or abstains.

Surface Canonical source and owner Bounded projection Actor-readable source/ref Decision and retention rule
Owner authority and biography Canonical logs/chat.jsonl, archive generations, and memory/dialogue_blocks.json owned by the canonical drive Main/Project context sections and archive-aware history windows Existing chat_history/archive readers with generation and gap metadata A known gap is disclosed; summaries/blocks never replace exact current owner directives. Raw generations and durable blocks follow their existing retention owner.
Execution evidence Task results, observability call manifests/blobs, service logs, and process-custody records Status cards, terminal rows, bounded tails, and compact child summaries Exact artifact/blob/service-log refs carried by the task result or canonical promotion A projection cannot certify a missing child/source. Referenced canonical artifacts are promoted before child-drive GC; disposable execution scratch follows unified GC. An omitted-to-artifact verification ledger stub carries only its re-projected summary; entries and axes are read from the artifact file it points at.
Terminal task/project memory Root terminal result plus existing task/project summary producers Cognitive Main terminal summaries and the two Project-root UI lifecycle rows (started + terminal completion) Task-result ID, project binding, and summary/source refs Summary is a biography projection, not raw evidence. Terminal outcomes, including failed/cancelled/degraded, remain retained through their canonical result owner.
Background Consciousness observations data/state/consciousness_observations.jsonl, append-only enqueue/ACK rows owned by BackgroundConsciousness Pending count/oldest metadata and a bounded recent observation rendering read_file(root='runtime_data', path='state/consciousness_observations.jsonl') Unacknowledged rows survive restart/overflow/error. Gaps block ACK and the existing direct identity rewrite; only a settled successful cycle appends ACK.
Plan/review authority Exact task-artifact/observability wave bodies, evidence selectors, reviewer route/thread receipts, and the bounded review hot index Review status, latest wave, obligations, and compact findings; a predecessor's inherited plan_review_state is first projected to a compact authority core ordered around the newest wave's identity, acceptance claims, findings, and dispositions, with reviewer transport removed and need_evidence_seen last-priority. Every bounded collection names its total and omitted count; the projection discloses full_chars plus source_ref, and the named include_authority source stays complete Exact artifact/source handle plus SHA/range/thread selectors Missing or partial evidence is DEGRADED/NOT_RUN, never PASS. Exact artifacts remain bound to the reviewed candidate SHA; hot indexes may rotate only after the source is retained.
Task acceptance (three deliveries) The FULL host packet (review_evidence.build_task_acceptance_evidence under the host ladder, with its __provenance__ table), the applied host run retained through canonical task source handles, and the paid-identity wallet ledger The per-delivery work order: the api pack for a packet row; the FULL packet plus absolute pointers and the access disclosure for an agent-session row; the packet without its freely degradable tail plus the real data root for a native inspection row (loop_acceptance_review.acceptance_retrieving_work_order) Exact evidence_refs from the packet's enumerable exhibit vocabulary; absolute pointers to the task's active workspace, task result record, artifact directory, verification receipts and tool-trajectory log; review_projection.panels[].applied_source_ref for the complete redacted applied review Refs resolve against the FULL packet only, never the rendered projection; a session's reads are unobserved by the host (disclosed), a native episode's are host_observed; the immutable-core overflow refuses every delivery, a partial tool-result projection only packet rows; one strict wallet claim per panel whatever the rows' deliveries (owner R11).
Canonical versus execution roots Canonical budget/data root owns identity, authority, biography, results, and promoted observability; execution drives own tools, workspace, transient trajectory, and per-call manifests while a task runs Project/fork/task lenses and status projections Existing canonical-root resolver, task-result pointers, and source handles A fork is an execution lens, not a second mind. Copy-back/promotion precedes GC for anything referenced by a canonical result; before terminal promotion the canonical reader cannot resolve a ref bound to a child drive (tracked as #805), and missing legacy bytes become an explicit gap.

Acceptance source identity is computed before history-dependent packet budgeting. The complete receipt/tool sources and work artifacts still invalidate the binding when their facts change. Applied reviews and completion observations use the existing write-once source_handles/context_checkpoints store and verified task_source refs; source handles stay outside both deliverables and the acceptance artifact manifest. The existing artifact route selects a published ref through its source query and returns digest-verified bytes. Ordinary artifact downloads retain their existing path. api_client.taskSourceDownloadUrl owns the shared browser URL contract. The final packet size includes source references and omission notes. Individual materialized tool records have addresses tool_trajectory:<corpus-sha>[<source-index>], not positions in the moving tail. task_acceptance_review(evidence={tool_trajectory_indices: [...]}) selects earlier records from the retained corpus through the existing source reader. The full trajectory remains partial when its head is omitted; a selected record resolves independently only when its actual arguments and result survive final packet budgeting complete. Selection changes the view, not source revision or authorship: agent-supplied prose never becomes host evidence, and a readable source does not prove the reviewer read or understood it.

Child copy-back uses observability._rewrite_child_ref_tree for typed refs in the owned acceptance-checkpoint and trajectory JSON formats. It preserves the complete source handle and rechecks dependencies even when the outer checkpoint was copied before. Relative source addresses preserve immutable checkpoint bytes; rebased observability refs require a newly addressed checkpoint. A failed copy retains pending custody before cleanup; missing legacy bytes remain unavailable. If rebasing changes trajectory bytes, the existing handle retains the original corpus_sha256 for record citations while sha256 verifies the transported copy. Existing row/reviewer addresses therefore survive cleanup without rewriting claims. Normal copyback selects CURRENT review authority through the existing field reducer, prepares referenced bytes outside the result lock, and publishes only while the selected ref/binding basis still matches. A changed basis repeats preparation outside the lock; unrelated newer fields survive. Pending retry starts from CURRENT, not a stale child body. child_ref_promotion_scope memoizes verified work only for one operation; failures are not cached as success. Same-physical-store copies keep the original manifest bytes/digest and return canonical path spelling without adding promoted_call_manifest; distinct-root copies keep their existing provenance marker and filename. Missing aliases resolve only to the exact canonical CAS/call address through shared verified readers, including the model-send reverse reader. Existing corrupt or wrong-scope bytes never trigger a convenient fallback. No arbitrary JSON-path crawl, new manifest filename format or persistent copyback ledger is introduced.

review_projection.publish_acceptance_checkpoint saves full applied host records before updating the compact task-result field through write_task_result and emitting the existing review_reference invalidation with the terminal task's explicit chat id (including hidden chat 0); loop/plan callers retain their context default and Project binding retains addressing precedence. Actual task attempts and host-only publication revisions order snapshots of each panel; the common merge also covers effective child reads and copy-back. Neither the projection nor its ordering stamp grants review authority. Old records without a retained full source and failed source writes disclose applied_source_status="unavailable".

11. Frozen Contracts v1 (ouroboros/contracts/)

ouroboros/contracts/ is the frozen ABI package for the skill/extension layer: typed protocols and shapes TOGETHER with their parsing, normalization, capability, and policy helpers (task_contract.py, skill_manifest.py, plugin_api.py, skill_payload_policy.py, chat_id_policy.py, schema_versions.py, tool_context.py, tool_abi.py). The frozen property is backward compatibility, not absence of behaviour: shapes and helper semantics stay stable for existing consumers, and the protocols are verified against the real implementations by tests/test_contracts.py. The browser-envelope ABI is additive and lives in ouroboros/gateway/contracts.py with the JSDoc mirror web/modules/api_types.js, pinned by the contract/parity suites — not a second file in this package.

11.1 What is frozen

Contract File Anchored by
Claudexor login/status envelopes — ClaudexorLoginJobResponse (required top-level job) and ClaudexorLoginJobProblem (required error, optional code, bounded required_actions); the daemon envelope passes through verbatim ouroboros/gateway/contracts.py + web/modules/api_types.js tests/test_gateway_parity.py, tests/test_claudexor_owned_daemon.py
ToolContextProtocol — the minimum context surface tools may rely on ouroboros/contracts/tool_context.py tests/test_contracts.py (duck + AST checks)
Tool module ABI — ToolEntryProtocol + GetToolsProtocol; every registry entry satisfies it ouroboros/contracts/tool_abi.py tests/test_contracts.py
Gateway browser envelopes (gateway/contracts.py; ABI 7.0 removed the contracts/api_v1 re-export and the cost_usd*/telegram_chat_id/project_last_viewed/project_hidden compat aliases) — browser WS/HTTP envelope families and TaskCreateRequest (optional caller metadata; executor_ref host-owned; costs nullable — whole-object omission over a confident $0) ouroboros/gateway/contracts.py + web/modules/api_types.js tests/test_contracts.py, tests/test_gateway_parity.py, tests/test_gateway_abi3_removals.py
Provider Test gateway ABI — ProviderTestRequest is exactly {provider_id, overrides?} and ProviderTestResponse exactly {ok, error?} (allowlisted request-local overrides; bounded errors) ouroboros/gateway/contracts.py tests/test_gateway_parity.py, tests/test_provider_key_test.py, web/tests/provider_test.test.js
Presence settings card + CAS update (reviewed defaults, local overrides, state fingerprint; CAS touches only presence-profile state) ouroboros/gateway/presence_settings.py tests/test_extensions_api.py, tests/test_gateway_parity.py
client_surface — optional closed-key normalized/bounded client descriptor, host-stamped, propagated to task metadata and chat history; absence is an explicit honest gap ouroboros/client_surface.py tests/test_contracts.py, tests/test_gateway_parity.py
cancelable + cancel-response cascade — additive fields with host-attested UI gating semantics ouroboros/gateway/contracts.py tests/test_gateway_parity.py, cancel/history tests
Executor-route projection — an opaque dispatch decision distinct from execution evidence; empty means native/no chip; a BLOCKED harness pin projects the route it named onto the live frame from cap_info alone (executor_blocked_route; the durable record keeps its empty route so the completion-seam evidence gate stays closed) and its typed subagent_executor_unavailable terminal renders {harness} · blocked; the sticky renderer is log_events.js executorChip ouroboros/agent.py, ouroboros/subagent_dispatch_notes.py, ouroboros/subagents.py, ouroboros/gateway/history.py, web/modules/log_events.js tests/test_claudexor_owned_daemon.py, web/tests/review_truth.test.js
ChatOutbound.executor_observation — optional event-local task/attempt/run/harness/phase/revision facts; model requires explicit requested/observed provenance, with only requested-model production from the current live timeline. Last activity is neither current liveness nor terminal evidence; foreign/older observations cannot replace their owning facts ouroboros/delegate_progress.py, ouroboros/subagent_messages.py, ouroboros/agent.py, supervisor/events_chat_delivery.py, ouroboros/gateway/history.py, ouroboros/gateway/contracts.py, web/modules/api_types.js, web/modules/log_events.js tests/test_executor_observation.py, web/tests/wire_contract.test.js, web/tests/review_truth.test.js
execution_evidence — started/settled/succeeded/failed counts, delegated_run_failure_states, evidence_read_failed, nanny_nudge_recorded, subscription_cost_usd (None while undisclosed — never 0), subscription_cost_estimated, harness_models, applied_access_profiles; derived from durable custody rows by delegate_evidence.task_execution_evidence, attached in subagents.envelope_from_task at terminal statuses only (never overwriting effective_executor/executor_route), and enriched onto the pushed task_done frame by enrich_task_done_event in supervisor/subagent_task_truth.py; actual_substrate ∈ harness_used/harness_attempted/native_only from custody evidence only; substrate_result_fields = {actual_substrate, delegated_runs_started, delegated_runs_settled, delegated_runs_succeeded, delegated_runs_failed, delegated_runs_source_unresolved, native_contribution}; the wait_tasks compact projection carries dispatch_executor plus that same set; an unreadable custody log omits the substrate claim and counts everywhere — evidence_read_failed means UNKNOWN, never "no run", and absence of evidence on a pre-evidence stored result is never a zero-run receipt; native_contribution is the constant "unknown" (no share/ratio is derivable from custody rows); a verifiable native_only amends capability_delta with delegated_substrate_unused; the log_events.js executor chip renders layered truth with unverified work-order counts spelled out ouroboros/delegate_evidence.py, ouroboros/subagents.py, supervisor/subagent_task_truth.py, web/modules/log_events.js tests/test_execution_evidence.py, tests/test_terminal_delegation_receipt.py, web/tests/review_truth.test.js, tests/test_task_status_flow.py
TaskCostBreakdown (root-only, read-time, never persisted; accounted_upper_bound_usd; authority="physical_attempt_ledger") + cancel_state: "pending" with cancel_reason beside it; the browser's one consumer is log_events.js taskCancelPending ouroboros/gateway/contracts.py tests/test_gateway_parity.py, web/tests/cancel_run.test.js
Task hurry ABI — POST /api/tasks/{task_id}/hurry with exactly {request_id} (extra fields refused), duplicate = idempotent success; OwnerHurryProjection attempt-keyed states; consumers log_events.js taskSoftStopPending/ownerHurryProjection (task-card only) ouroboros/gateway/task_hurry.py, ouroboros/gateway/contracts.py tests/test_owner_hurry_s3.py, tests/test_owner_stop_s3.py, web/tests/task_control_menu.test.js
StateResponse.active_direct_turns/active_chat_activities (phases queued/working/finalizing/budget_paused — one predicate budget_pause_fact decides budget pause); TypingOutbound activity fields; ChatOutbound.task_phase/task_terminal_status ouroboros/gateway/contracts.py, supervisor/active_activity.py, web/modules/chat_activity.js tests/test_gateway_parity.py + the activity test files
project_thread stamp on all seven outbound frame types, stamped at the message-bus broadcast choke; a stamped frame is never adopted by Main (chat_activity.mainThreadAccepts) supervisor/message_bus.py, ouroboros/projects_registry.py tests/test_message_bus.py, web/tests/chat_thread_routing.test.js
Media/link envelopes — media task_id/size_bytes/download_url; LinkAction {label,url} with at most twelve absolute HTTP(S) actions; links in WS_MESSAGE_TYPES; chat.links host topic ouroboros/gateway/contracts.py, ouroboros/tools/core.py, ouroboros/event_bus.py tests/test_contracts.py
Owner quiz ABI — QuizOption {label, detail?}, QuizOutbound (quiz_id, question, options, stake, assumption (required for optional clarification), additive wait_for_answer for a live pooled or ordinary-conversation root that must wait, lifecycle state open/answered/expired_terminal/superseded), separate QuizStateOutbound discriminator, chat.quiz host topic; the producer is the one escalation verb escalate(question, options, stake, assumption, wait_for_answer=False) — a ROOT asks the owner, a SUBAGENT delivers a typed frame to its nearest LIVE ancestor, which answers via forward_to_worker or escalates verbatim, so the owner sees only what no ancestor answered; answers arrive through the ONE ingress POST /api/decisions (family ids quiz:{task_id}:{quiz_id}, routing:{client_message_id}:{routing_token}; interaction: reserved), request-id idempotent, first answer wins, validated against the STORED options; option_index is optional for the quiz family alone — a comment-only answer writes NO answered_index, because a stored 0 would replay as "chose the first option"; injected as the typed KIND_QUIZ_ANSWER mailbox control and broadcast as quiz_state (carrying the recorded comment when the owner answered in their own words, so the live card shows Owner's answer: exactly as replay does); expiry is structural only (the task-done seam flips open quizzes to expired_terminal, and the SAME reconcile closes the paired owner_wait so a terminal task never projects quiz=expired_terminal beside owner_wait=waiting; owner_wait.set_owner_wait's refusal to continue waiting on a terminal result is preserved, not caught) and history replay merges the projection state ouroboros/gateway/contracts.py, ouroboros/gateway/task_decision.py, ouroboros/owner_quiz.py, ouroboros/tools/core.py tests/test_gateway_parity.py, tests/test_quiz_display.py, tests/test_quiz_answer.py, web/tests/chat_decision.test.js
Managed update ABI — preflight, UpdateMergePlan, pinned apply, update_status_ready WS notice ouroboros/gateway/contracts.py tests/test_update_apply_routing.py
ChatOutbound.review_projection — bounded actor findings via utils.truncate_review_artifact, at most MAX_PROJECTED_ACTOR_FINDINGS rows (review_execution_projection.py) ouroboros/gateway/contracts.py tests/test_review_substrate_v2.py, web/tests/review_truth.test.js
Skill preflight statuses — preflight_failed is fresh-only; a stale failure surfaces as preflight_failed_stale; absence means the caller could not know ouroboros/skill_review_status.py tests/test_skill_preflight_repair.py, web/tests/skill_preflight_repair.test.js
chat_id_policy — the SSOT for human-visible vs synthetic chat ids across message bus, history, memory, and consolidation (§12) ouroboros/contracts/chat_id_policy.py tests/test_chat_id_policy.py
task_contract — normalization + effective_acceptance_claims(task, closed_plan_wave), a pure read-time binder where ingress wins over a closed plan wave (acceptance reads fresh claims without mutating the live contract); an open wave binds nothing and is disclosed as a non-binding exhibit instead; pacing interprets the budget profile via typed task_pacing.CostCeiling; Presence promotion and follow-ups copy the ceiling by value (§12) ouroboros/contracts/task_contract.py tests/test_contracts.py
PluginAPI 2.0 — the full 16-method extension surface, negotiated against the manifest (major strict, minor minimum, closed-set capabilities, typed educational refusals) BEFORE plugin import or out-of-process cataloging; an absent manifest field means legacy 1.3 by construction, ExtensionRegistrationError, FORBIDDEN_EXTENSION_SETTINGS, VALID_EXTENSION_PERMISSIONS, VALID_EXTENSION_ROUTE_METHODS (route methods mirrored against server dispatch by test_extension_route_methods_contract_matches_server_dispatch); skill_job_dir(job_id) creates jobs/<sanitized>-<hash>/{assets,output,tmp}; host-mediated permissions (companion_process, supervised_task, subscribe_event, inject_chat, presence) require review/owner grants; the ExecutionMode capability matrix is the SSOT for what a per-call child can proxy ouroboros/contracts/plugin_api.py tests/test_contracts.py, tests/test_extension_loader.py
SkillManifest — unified frontmatter (instruction/script/extension), reviewed scheduled_tasks cron metadata, bounded canonical conflicts, presence: block parsed by presence_profile.py; parse_skill_manifest_text() tolerates missing optional fields and validate() returns warnings without raising ouroboros/contracts/skill_manifest.py tests/test_contracts.py
schema_versions — opt-in _schema_version stamping (with_schema_version/read_schema_version); wired by extension health.json, the projects registry/bindings, and presence bindings ouroboros/contracts/schema_versions.py tests/test_contracts.py

11.2 What is NOT frozen (intentionally)

The full ToolContext dataclass (browser state, review history, model overrides, …) stays mutable implementation detail — the protocol pins only the minimum. Raw WebSocket/HTTP values are unpinned; only the shape keys are. The SKILL.md body is free-form; only the frontmatter schema is pinned. state/state.json and queue_snapshot.json carry no _schema_version key and read as version 0; task_results/*.json is the exception — every write is stamped _schema_version: 1 and admission is task_result_schema.py's, not this section's.

11.3 What to do when extending

Add the field to the active frozen owner — ouroboros/contracts/ for the package ABI, or ouroboros/gateway/contracts.py + web/modules/api_types.js for browser envelopes — keeping existing consumers working, and enforce the new surface in the contract/parity tests (CHECKLISTS item 17, gateway_parity, owns the review-time criteria). Removing anything from 11.1 is a deliberate ABI break: it requires an explicitly versioned successor and a migration note in the release row — the release ledger, not this map, is the SSOT for retirements.

11.4 Recent ABI Retirements

  • ABI 7.0 is one deliberate window, so the breaks land together instead of one per minor:
    • Gateway envelopes. Five compatibility aliases are gone: cost_usd / cost_usd_with_children (stripped at the cost SSOT seams — stored rows are still READ tolerantly), telegram_chat_id from the four outbound frames and the history mapper, and project_last_viewed / project_hidden from UiPreferencesResponse and its endpoint, which now answers 400 on an unknown key. The contracts/api_v1.py re-export went with them, and the contract became executable (gateway/schema.py).
    • Settings keys. RETIRED_SETTING_KEYS in settings_defaults.py is the machine-readable list — read the tuple, not this paragraph, for membership — and load_settings strips its members off disk. This window added OUROBOROS_SCOPE_REVIEW_FLOOR, the flat OUROBOROS_SOFT_TIMEOUT_SEC/OUROBOROS_HARD_TIMEOUT_SEC pair (superseded by the activity model), OUROBOROS_REVIEW_NATIVE_MAX_ROUNDS (the native review episode is bounded by its transcript ceiling, the owner deadline and the wallet, not by a round count), OUROBOROS_OBSERVABILITY_RETENTION_DAYS, and — classified separately as RETIRED_COMMA_LIST_SETTING_KEYS — the reviewer comma lists and route envs (OUROBOROS_REVIEW_MODELS, OUROBOROS_SCOPE_REVIEW_MODELS, OUROBOROS_SCOPE_REVIEW_MODEL, OUROBOROS_REVIEW_ROUTES, OUROBOROS_SCOPE_REVIEW_ROUTES, OUROBOROS_ADVISORY_REVIEW_ROUTE). Their migration note: move the configuration into the structured OUROBOROS_REVIEWER_SLOTS BEFORE upgrading — an install carrying only comma keys comes up on the shipped default panel. The owner is TOLD: the read seam logs the dropped keys once per process, and the first supervisor boot with an owner chat bound posts one system row there (server_maintenance._startup_retired_settings_notice, the same sentence — settings_defaults.retired_setting_keys_notice — naming the keys as NOT honored and, from reviewer_slot_config.authored_reviewer_slots_state, what runs now: the authored panel, the shipped default while the structured key is absent, or NO panel with the parse error while it is malformed and the loader rejects it), deduplicated durably per retired-key set in state.json:retired_settings_notified. A comma-spelled ENV projection survives as the derived runtime plane, never as configuration.
    • Plugin ABI. PLUGIN_API_VERSION is "2.0" with manifest negotiation checked before plugin import or out-of-process cataloging; an absent field means legacy 1.3 by construction, and a hash-bound PASS is grandfathered.
    • Durable task rows. Every task_results/<id>.json write is stamped _schema_version: 1, with deliberately NO legacy converter: an inadmissible row is quarantined with log-only visibility and keeps its id occupied.
    • Typed internals. ResolvedModelTarget is the frozen model-target dataclass at the existing resolution seams, extension registrations publish atomically with a per-publication generation digest, and the dead _call_llm_with_retry, compute_cost_with_children and format_handoff_message surfaces are gone.
    • The migration note is a program. scripts/rc_audit.py is a READ-ONLY pre-upgrade scan of a third-party install that emits a machine-readable scope document naming every incompatibility it finds across the five frozen classes (gateway alias, retired setting, comma list, plugin API, schema stamp), snapping the retired-key lists at execution time instead of hardcoding them.
    • Deliberately NOT in this window: the handler ABI — tool handlers returning ToolResult instead of str — is backlog, so handler signatures are unchanged.
  • 5.25.0-rc.4 retired the native skill upgrade migration banner API (GET /api/migrations, POST /api/migrations/{key}/dismiss, and MigrationsResponse). The migration note is the release row itself: dismissed banner state in data/state/migrations.json is intentionally ignored by current runtimes.

12. Host Service, Companion Processes, and Chat IDs

Chat uploads have one gateway.files storage owner for Host-confined paths and completed multipart spools. It uses the artifact substrate's streaming hash and atomic copy; borrowed spools promise descriptor identity, not an original pathname. The multipart request's worker owns both copying and closing its spool, so HTTP cancellation waits for both and cannot interrupt cleanup. Host uploads use the same settled wait before releasing their existing in-flight slot; the skill keeps ownership of its source. store_chat_upload keeps its Path return. The old 50 MiB chat upload rejection is removed on both ingresses; Host's existing 25-file request count and the separate Files-browser upload policy retain their own contracts.

Telegram document mirroring resolves the captured file_ref and streams an owned file handle through its existing multipart client, with legacy inline bytes still accepted. Its outgoing 50 MiB boundary is separate from the integration's inbound 10 MiB download limit. Oversized files stay saved in the application: a ready, already-running owner-authenticated Mini App can be opened through its existing entry button; otherwise the message explicitly says it cannot mirror the file and directs the owner to the app. Delivery never starts a tunnel or publishes a new public/token-bearing artifact URL, and this notice does not claim the bytes were uploaded to Telegram.

The Host Service is a loopback, authenticated callback boundary for reviewed skills (ouroboros/gateway/host_service.py, 127.0.0.1:${OUROBOROS_HOST_SERVICE_PORT:-8767}). Every request authenticates an opaque x-skill-token bound to the skill's content hash, executable review, enablement, and grants — secrets never enter the token, a payload edit stales it, and the client wrapper refuses stringification (skill_token.py). The frozen route family is exactly: /identity, /tools/schemas, /chat/allocate-internal, /chat/inject, /chat/operations/{operation_ref}, /chat/cancel, /chat/decision, /presence/turn, /presence/work/{work_ref}, /ui/ws-message, and WS /events; permissions still decide which route works for a given skill. Review of a transport skill evaluates identity binding, attribution, polling bounds, panic cleanup, token confinement, and exfiltration — an owner-bound reviewed transport may be a first-class control surface, not a screen-only integration. External slash commands bind a separate positive-identity external owner slot (supervisor/state.py), so an unidentified transport can never bind commands and the local web owner can never lock out a real remote owner. wait_for_response remains restricted to A2A-allocated chats. Named messages wait for their exact operation; only the legacy unnamed path uses chat-wide response subscriptions. Admission is one limiter with two policies (host_service._RateLimiter): the WS relay lane (/ui/ws-message) is a token bucket — a 60-message burst reserve per skill refilling one message per second, so a burst no longer silences a widget for the rest of a minute — while every other lane keeps its 60-per-60-seconds sliding window. A refused relay is visible at the host and aggregated per burst: one warning when the reserve empties, then one warning plus one durable host_service_ws_relay_dropped row in logs/events.jsonl carrying the dropped count when the lane admits again or the idle bucket is swept; the 429 carries retry_after_sec and dropped_in_burst, the child's send_ws_message stays best-effort None, and the in-process broadcast path keeps no bound.

Host attachment copies use one request-local cleanup stack and the shared settled HTTP-worker wait. Partial copy/refusal/cancellation removes only new unaccepted copies. Named ingress retains those copies once its existing canonical write is attempted, through an internal custody callback; a replay never adopts its duplicate copies. Cancellation waits for that admission worker. Accepted bytes survive disconnects and unknown write/queue outcomes. Such an error does not claim acceptance and may retain an unused copy; no new reconciliation or transaction log is introduced.

Operation correlation (#667): a named injected message has operation_ref=<chat_id>:<client_message_id> on 202, 200, 504 and disconnect responses. supervisor.message_bus.accept_local_message serializes check, canonical inbound-row acceptance and enqueue; its existing log_chat writer must succeed before work is queued. The host-minted full origin is carried internally to record_inbound_message, which verifies the source and consumes it without writing a second row. Repeated same-id, same-text, same-skill delivery rejoins even before supervisor dequeue; changed content or source is refused with 409. This preserves the existing in-memory queue: a crash after acceptance can lose delivery and reads honestly as lost, never authorizing a second enqueue. Reads span retained canonical chat generations. Routing annotations and outbound task ids are discovery hints only: task reads and cancellation require the actual queue/task record's complete origin_message_ref to match the authenticated skill's canonical source. DirectActivityRegistry carries that same origin. Named response waits poll this exact operation and its retry-aware effective task result; chat ordering alone never proves a reply. Cancel enters the existing durable intent and cascade-custody owner only for work with that origin and the same installation root. A different or unavailable owner root is disclosed as cancel_unsupported before any intent is written. Unaddressable/foreign work remains cancel_unsupported; unresolved custody remains explicit and never becomes a false cancelled.

Successful extension children also carry a bounded ws_relay_failures count map from the existing PluginAPI transport owner through their result envelope and process facts. Missing transport, network errors and HTTP refusals remain visible without changing send_ws_message -> None or the producer's success. One host warning reports the aggregate; diagnostics contain no message bodies, URLs or credentials. A child that dies before its final envelope may lose this aggregate; the existing measured death/timeout facts remain authoritative. The separate streaming route implementation carries the same aggregate in completion X after body/background work.

Presence flow: POST /presence/turn requires the content-hash-bound presence permission, one binding_id, one exact transport event, and optionally staged files confined to the skill's state root; the binding resolves from state/presence_bindings.json with exact provider/account/conversation/thread origin verification; cross-process locks enforce the installation-wide cap and serialize one conversation_key; a stable event-derived task id makes transport retries idempotent; input and output join ordinary dialogue history with full transport/actor provenance. The one typed outcome is message/silent/tool_delivered/deferred — deferred only with a correlated work_ref, because an unanchored "deferred" would be an unanchored promise — and GET /presence/work/{work_ref} polls the bound late result without exposing the general task API. Promotion of Presence work into a managed task clears requested Project/workspace/source widening: a public conversation may promote long work but cannot choose new authority, and the cost ceiling plus return destination follow the promoted root by value. presence_cancel_work acts only on a work_ref whose stored binding and conversation match the current turn; owner chat and Background Consciousness may initiate_presence on an existing enabled binding.

Companion processes are host-supervised: reviewed manifest-declared descriptors enter durable custody, reconcile after lifecycle changes and restart, and stop on disable/unload/panic; state/extension_generation.json carries the opposite direction — the server's published live set, which a running task worker adopts at a task's start or at a dispatch miss, so an enable after boot is not invisible until the pool respawns. Worker-side changes write durable reconcile requests (state/extension_reconcile/) rather than spawning server-owned children; every reconcile state names the process that answered and whether that marker request was written, and the tool and review receipts pass both facts through. Health observations are process-qualified: aggregate and Skills UI health use the server observation as authority and expose the worker observation with its handoff outcome only as a qualifier, so failed handoff success cannot advance last_known_good; restart-budget exhaustion persists a terminal reason in that health state, cleared only by a later successful start. A companion's cwd is the reviewed payload directory, so a payload edit stales review before reload instead of silently mutating a live process. The live projection is state/extension_companions.json.

Chat IDs: a chat id is a VALUE and absence is None. HIDDEN_CHAT_ID (0) is the hidden partition — the Skill Review panel plus every headless task admitted without a registered project — a REAL destination that no browser surface reads: delivery goes through membership routing (message_bus.notification_chat_route; every producer that tested if chat_id: dropped its notices), a chat_id=0 history query coerces to Main and the Main filter drops chat-0 rows, so explicit panel rows never become ordinary conversation history, and a chat-0 row reaches a project thread only via a durable lineage binding; negative ids are synthetic A2A traffic and never enter a human stream (the id policy SSOT is the §11.1 chat_id_policy row).


13. External Skills Layer

Native and external payloads live in separate data-plane buckets (data/skills/{native,clawhub,ouroboroshub,external}), with review, grants, enablement, dependencies, tokens, and health under data/state/skills/<name>/ (§1 Data layout). Discovery and manifest parsing establish identity, source, hash, provenance, and conflicts — never trust; a conflict declared by either enabled peer is enforced symmetrically without deleting either payload.

After a real skill review, explicit Advisory author finish may accept a revised payload after deterministic preflight without another panel. SkillReviewState.gate_for is the shared readiness/extension/Host Service admission: is_stale_for still reports the original critic hash, author_disposition.subject_hash names the author's current bytes, and Blocking requires fresh critic authority. Repeated finishes preserve the original hash; desired enablement, grants and dependencies remain independent. Skill publication retains its own reviewed-capture requirement.

The executable sequence is install → deterministic preflight → hash-bound multi-model review → grants → dependency readiness → enablement → execution, and the gates stay independent: a review PASS installs no dependencies, enabled=true does not prove readiness, and an extension additionally needs host registration (§10 invariant 8). Mutating lifecycle work flows through one deduplicated queue (skill_lifecycle_queue.py); review jobs retain task/source/hash/attempt/actor/terminal evidence in the private full record, with the compact UI history as a projection. Review ordinals are allocated only after a job starts under the lifecycle lock — retry history stays explainable across hash changes without a UI counter becoming review authority; a started failure/cancel/timeout consumes its number with one idempotent terminal row, while pre-start dedupe consumes none. Skill-review history displays its actual review round, snapshot attempt and revision facts rather than renumbering the retained tail. Accepted rebuttals reduce reviewer thrash; a new payload hash still requires fresh evidence.

Skill review combines the deterministic preflight with the authoritative multi-model checklist review; an optional advisory stays fail-open and cannot replace it. Official catalog payloads get their reduced-noise profile only when the sidecar, catalog file set, local file set, and every SHA-256 match exactly. A deterministic preflight failure persists as PENDING, not BLOCKERS — BLOCKERS could be overridden under advisory enforcement, while PENDING is non-executable in every mode.

skills/telegram/ is the bundled transport: an in-process extension owns binding/polling/injection/settings while a supervised companion owns the sidecar, tunnel, menu rollback, heartbeat, and singleton; the text bridge works when the Mini App/tunnel is unavailable. The payload is seeded with hash-bound native provenance and stays disabled until token/permission grants; it is never a marketplace install. Its state classes stay separate under data/state/skills/telegram/ (settings/binding, companion config, menu rollback snapshot, delivery cursors, verified tunnel cache). Colab bootstrap waits for native discovery plus a fresh executable seed projection, grants only missing grantable items under the owner auto-grant policy, then enables. Inbound Telegram files (documents, video, audio, voice; this integration's 10 MiB download cap) ride /chat/inject attachments — paths confined to the skill's own state root that the host copies into the browser paperclip's data/uploads store and stages like any chat attachment, so a caption-only text or a text-less file is one ordinary owner message; owner answers to quiz cards (a tapped option or a reply to the card) ride /chat/decision, the same task_decision.answer_decision ingress as the web card; the bridge status reports degraded (telegram_startup_deferred) while a dead network defers token validation, and the task-done push takes its word and icon from the stamped outcome_phase.

The optional model_experience manifest prose — what the skill adds to the model's context and what it costs in tokens — travels to the model-visible surfaces (the list_skills JSON and the installed-skills context section), and a manifest without the section keeps the exact prior rendering; manifest parse refusals teach the repair through the fix_hint every SkillManifestError carries.

Marketplace installs are bounded archives staged privately and landed atomically with per-file hash checks; install metadata drives isolated dependencies (marketplace/install_specs.py); manual instructions remain guidance, not execution. Exact download descriptors and per-spec source/postinstall opt-ins share that dependency owner: fresh review and hash-covered declarations authorize literal build/check argv, verified resources and package caches survive under state/skills/<name>/dependency_cache, and existing deps.json records resolved package metadata, bytes, outputs and diagnostics. Delivery is distinct from a declared executable check; absent checks remain unknown, and declared-output drift invalidates installed readiness. Large resources stay outside the payload/Git patch. Delegated payload capture uses the existing review byte classifier and preserves binary resource descriptors such as PNG/Wasm. Go compiles to a private temporary executable before running literal argv and preserving its exit status; Deno translates current reviewed effects and actual forwarded env names without adding an OS sandbox. Extensions import through staged trees (_stage_extension_import_tree under __extension_imports/), so concurrent workers cannot remove a peer's live import and stale trees stay reclaimable. Per-call child processes may proxy tools/routes/WS/UI/settings/companion descriptors, while persistent subscriptions and supervised tasks require in-process or companion lifecycle — reported through the generic capability matrix, never inferred from a platform name. Isolated children run the same staged loader in a private base with a scrubbed env, so native crashes cannot kill server.py. In-process extensions are more powerful, which is why namespacing, declared permissions, per-skill tracking, and atomic unload are an executable contract rather than convention. Transport metadata records source and session generically; skill repair enqueues an ordinary managed task with an exact selected resource and revision admission (skill_repair_admission.py); a legacy skill_repair selector remains readable but selects no reduced profile. Shell, browser and delegation remain normal task capabilities; an installed payload needs no mandatory Git copy, and review never forces unload merely because the caller is repairing it.