ouroboros/docs/architecture/06-agent-core.md
Ouroboros 6a2ee44ac2 docs/tests: narrow the transport-death repeat wording to Host-executed effects; state the 409 condition exactly
Reviewer P2: a lost response proves no Host-executed tool or correspondent action from the failed attempt, not that nothing ran anywhere — priced read-only provider-owned retrieval (server-side web search) may rerun on the repeat. The 409 and owner notice apply when the round ends unresolved without a further permitted repeat (deadline or a finalize control can refuse the first repeat). Test comment on the owner-notice writer corrected; chapter budgets untouched.
2026-09-26 02:00:57 +03:00

658 lines
303 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# 6. Agent Core
This chapter maps execution and cognition inside a worker: task lifecycle, tools, context fitting, safety and runtime mode, reviews, delegated custody, planning, deep self-review, reflection, memory and project focus, skills, external tools and budgets. Pooled tasks and direct turns, including consciousness wakes, enter `OuroborosAgent`'s tool loop (`loop.py`); reviews and post-task operations use separate executors. Shared contracts do not make these paths one loop.
### Task lifecycle
A queued task enters through a reviewed transport, is admitted by the supervisor queue (§5) and runs in `OuroborosAgent`; each direct-chat turn runs on its own in-process agent, tracked by the process-local `DirectActivityRegistry`, and creates no `PENDING`/`RUNNING` queue record (§3). The root pipeline (`agent_task_pipeline.py`) captures the task contract and immutable context core, runs the loop, preserves a delivery candidate, stores the result (`task_results/<id>.json`, every write stamped `_schema_version: 1`; an unstamped, future, malformed or retired-key row is quarantined by `task_result_schema.py` with its id kept occupied) and artifacts, emits lifecycle and usage evidence, performs the root-only post-task work (`post_task_synthesis.py`, checkpointed by `post_task_checkpoint.py`) and publishes the typed outcome. Queue admission proves only that asynchronous work was durably accepted; completion, objective satisfaction, artifact finality, verification and review acceptance remain separate facts. The loop's tally rides `loop_outcome.usage` while the top-level `total_rounds`/`prompt_tokens`/`completion_tokens` are the ledger's answer (`reconstruct_task_cost`); an internal exception is `failure.kind = "runtime"`, never a fabricated provider failure, and a capture the crash lost is reported unknown, never as zero counters.
The separate review and post-task executors (`review_execution.py`, the post-task pipeline) use the same evidence and context contracts. `supervisor/cognitive_operations.py` tracks in-flight LLM, review, VLM and tool work as typed `cognitive_operation` leases for the idle rail; deadlines, budgets, cancellation and absolute ceilings remain independent.
Between the sends of one loop execution the transcript is append-only — each send a prefix extension of the previous — because OpenAI-family caches reuse a request only when it is a byte-prefix of the next, so a replaced tail or a rewritten earlier message discards the whole conversation cache. `transcript_prefix.py` records and never blocks: `unsent_in_previous_send` allows `_append_or_merge_user_content` to merge only a tail absent from the last observed send; an absent slot or missing observation means unknown, so the producer appends a new row, compaction seams stamp `sanction_rewrite`, and a break is the `prompt_prefix_break` checkpoint fact (`kind` = `system_rewritten` | `tail_replaced` | `rewritten` | `shrunk`, plus `sanctioned_by`; a context-fit reprojection after a real overflow is an ordinary break). The observation follows a usable ordinary model response; it is not a ledger of every physical send, and existing image eviction is unchanged. A `FINAL ANSWER:` line is latched every round as the latest typed candidate, so review, nudge and forced-finalization paths never erase a structured answer; marker prompting is gated on `task_contract.answer_protocol="final_answer_line"` (`answer_protocol_active`) while the latch and extractor stay unconditional, and `outcomes.extract_final_answer` refuses the outcome-tier ledger identifiers as answers — internal enum vocabulary is never a deliverable.
`model_execution` is a compact projection of the last ordinary solve response the loop accepted (text or tool calls), selected from marked `llm_call_refs`; it keeps the initially requested model/local route, the route used and the provider's own label separate, and empty responses, forced finalization, review and post-task calls cannot replace it. Result, terminal frame, history and card share it without changing dispatch or claiming final-answer authorship; the terminal `task_summary` keeps it when the full result ages out, and a current result or live card outranks that fallback. After tool work, host-driven fallback or model-wait model changes record an authoring handover. The successor's first tool-less response gets one recovery round with delivery-control JSON held; a second yields `execution=degraded` and `authoring_handover_incomplete`. Ordinary no-tool turns and explicit model changes stay unchanged. Verification and skill checks still run; later tools clear only this warning, preserving incident and recovery history. Each handover gets one recovery prompt under ordinary stop, deadline and budget limits.
`DeliveryCandidate` is retained before verification or review so a later notice, reviewer failure, deadline or provider outage cannot erase a useful answer. `outcomes.py` combines the execution, objective, review, artifact and child-absorption axes without converting one into another — the terminal custody overlay (`outcomes.custody_debt_axes`) is an instance of that rule, not an exception; verify-before-done receipts and exact artifact references are host-attested evidence, and declarations or answer prose are no substitute. Individual failed tool calls alone do not degrade a delivered answer: unresolved calls retain their typed signal/exit evidence in `execution.unresolved_tool_errors`, cosmetic errors keep their separate bucket, and either kind adds `residual_tool_errors_without_review` when the canonical objective is `not_evaluated`, including after delivery/child-state normalization. Host acceptance or qualified Advisory author completion establishes objective success; unresolved FAIL/DEGRADED verdicts, empty or failed execution, provider failure, deadlines, incomplete delivery, deferred children and failed artifact verification keep their own outcome rules. A forced exit may publish the best current candidate only with its typed rail and evidence-freshness disclosure, and lifecycle may remain `completed` while the objective or review axis records a best-effort or unaccepted result. Custody debt heals from the write side while the stored `reason_code` may not be rewritten, so the owner-facing Reason line resolves the code at RENDER time — the custody warning only while the row's own `delegated_runs_unreconciled` list is non-empty, else the execution reason, both when both are real; a healed debt is never restored, and the debt is a warning beside the rail cause, never a replacement for it — as is every standing limitation of the same answer (deferred child results; a plan review still open at delivery, worded by the outcome class recorded on `execution.plan_review` when there is one), which the normal rail types on the result exactly as the forced rails do, so one clause per fact reaches the card, the durable row and Ouroboros's own memory.
Host plan/orphan disclosures ride beside the model answer as `terminal_host_notice`; delivery sends the model answer alone and the disclosure stays a field of the result (an owed pre-upgrade notice row still replays as the untyped System row it was), and the custody audit is its own typed row (`terminal_custody_notice` on the send event → `system_type="custody_notice"`, `card_row="timeline"`, its own owed delivery id), so the open-delegation fact is a row of the task's card and the unreconciled-runs note never enters the assistant text. CLI gets one host-labelled status (`terminal_host_notice_text`); external Presence speech excludes host notices (chapter 12). Stored answer bytes do not change: the answer keeps its hash and its real PASS or FAIL when only the notice changes, and equal text cannot revive a superseded verdict. Every reader that hands a result on — synthesis, parent handoff, `get_task_result`/`wait_task`/`wait_tasks`, the `task:<id>` plan-evidence reader — carries the notice as a separate field, and the child-result and plan-evidence hashes include it, so a changed child limitation invalidates an old parent disposition or an old review. A host-salvaged terminal is labelled, not hidden: the durable row carries `Preserved intermediate output (not a final answer):` plus the bounded excerpt every terminal row uses, the untruncated copy stays with `get_task_result` and the stop receipt, and only a row written to the chat the stop receipt itself reached (`cancel_receipt.delivered_chat_id`, recorded after a successful send) reduces to its label beside the pointer; the two durable writes are not atomic, and no writer trades its only pointer for the label. One event is disclosed once, at the layer that owns it: the forced orphan note leaves a child to its own terminal row only where that row reached the reader the note addresses, and a rail that ended a routing turn is named by that turn's own row, never inferred from another layer's stamp.
The delivery-control protocol is resolved here and only here (DEVELOPMENT keeps the rule and points here). The candidate carries sticky loop-local provenance that its lineage has seen a host-issued delivery-control episode; without one, exact JSON is ordinary text. In a marked lineage both the ordinary and the forced resolver intercept recognizable whole-body envelopes and balanced trailing protocol attempts — valid `keep` resolves to the retained candidate, valid `replace` to `full_answer`, anything malformed preserves the retained candidate. Both strip one whole-body fence and treat a balanced protocol object at the very END of prose as a protocol attempt (`utils.extract_trailing_json_object` + `loop_delivery._parse_delivery_control_body`); the trailing-object rule deliberately refuses substring scanning, because quoted protocol literals mid-prose are legitimate text, so a control object quoted mid-prose stays prose and a truncated trailing fragment remains prose. During an ordinary acceptance continuation, `finalization_control="acceptance_feedback"` makes a complete revised answer ordinary prose and keeps `keep`/`replace`/`pending_review` optional. Prose resets the pending-review choice to `wait`; it never means `finish`. An outstanding effect, owner-revision or child-action control retains its stricter rule, and unread owner source still requires acknowledgement. Empty or recognizable malformed control bodies retain the candidate rather than becoming its replacement. The ordinary resolver takes one repair round then degraded-preserve; the forced resolver resolves purely and never re-loops — malformation preserves the retained candidate with the typed `delivery_control_degraded` reason, which is this forced rail's own code, while the ordinary repair path records `invalid_delivery_control_after_repair`. Every degradation carries the cause it computed: `outcomes.derive_loop_outcome` falls back to `delivery_control_degraded` only for a degradation that reports no cause and publishes the `loop_outcome.degraded`/`degraded_reason` pair the benchmark ledgers read. A malformed attempt in a marked lineage never leaks JSON, even after the transient latch clears. Presence forced calls (turn or root, not ceiling-only child) always arm `presence_finish`; original bytes reject duplicate keys (chapter 12).
The child-absorption gate is an action gate: while undispositioned direct children remain, the loop HOLDS the candidate (`child_absorption_or_revision_required`) instead of arming the JSON-only control instruction — the hold-vs-arm split exists so the model never receives two contradictory instructions in one round. A typed keep cannot close the gate; after the one bounded reminder it forces the best-effort `children_unabsorbed` rail with a current `id [status] sha256` listing. The absorption digest's `## child` header carries the child's typed custody debt (`delegated_runs_unreconciled`, bounded, with a `get_task_result` pointer) as visibility only — the parent's authority over that patch is exactly the orphan rule, and a child's debt never relabels the root card (DESIGN §4). Finalizing over an UNDISPOSED OWN delegated patch is deliberately NOT gated: the consequence is disclosed where the decision is made (the `integrate_delegated_patch` schema, the apply receipts) and lands as the additive Done-with-warnings custody overlay rather than a hold; a pre-finalization reminder and propagation of child custody debt into the acceptance-subtree snapshot remain disclosed deferred gaps.
Provider death is the one forced rail that is NOT a best-effort completion: it salvages the best available text but stamps `infra_failed`, so the task terminalizes `failed` with the typed `provider_unavailable` reason and an immediate "provider outage — NOT completed" owner notification; a waited-out transport outage reaches the same rail through its deterministic no-resend branch (`transport_unavailable_no_resend`, or `provider_outcome_unknown_no_resend` when the round still holds a transport-death repeat record). The rail makes its one forced model call only while a call can still land: a transport that spent its same-model retry wall stamps `_llm_retry_wall_exhausted`, a no-call gate beside `context_overflow` and `provider_outcome_unknown`, so the forced rail never re-pays a second retry window over a proven-dead provider (disclosed residual: the marker is a last-invocation bool on the shared usage dict). Terminal delivery preserves producer authorship: `terminal_provider_notice` is host presentation beside the raw result, and receipts and secondary System incidents carry the wait duration, finalization cause and unknown-outcome warning and never recommend a blind rerun. Every forced rail stamps one closed-vocabulary producer word at the single forced-finalization sink: `model_final` for complete model text (host-authored notices stay separate), `host_notice` for host-written terminal text (an owner System row, never automatic external Presence speech), and `host_salvage` only on the provider-death rail, where managed/direct delivery replaces the text with the enriched outage receipt and the full bytes stay in task details. Missing origin stays unknown; `agent.py` resets it on host text replacement. Presence synthesis seals the recorded adapter body and internal result separately (chapter 12), never a delivery receipt.
Finalization controls are typed owner-mailbox entries, not injected owner prose. The supervisor may request one bounded tool-less answer, salvage the last persisted assistant text, and retain a full canonical copy when a preview would truncate it. A grace episode has one durable control and can be revoked atomically when the task itself resumes; descendant activity does not count as the task's own progress. A process that cannot be killed remains visibly running, and custody checks prevent another runtime instance from reaping work it does not own.
Disclosed cancel-lifecycle residuals (deliberate): a cascade over a tree with no resolvable lineage chat whose typed handoff-row append ALSO fails still settles; an empty-intent release can add liveness noise to a foreign claim's forensic trail (bounded by the generation fences); cascade postcondition timing can flake under heavy load (the watchdog re-feeds — one retry, never a lost teardown); and the cost projection of a task whose delegated runs stayed open may read `cost_usd=0`/`cost_final=true` while a run is still live — the disclosure line names the open runs.
#### Task acceptance
Host-enforced task acceptance is a root-owned completion coach, not the P3 commit gate. `off` disables it; `auto` and `required` review observable effects, typed deliverables/criteria, and an explicit root `task_acceptance_review` nomination, read-only research included. Queue membership alone does not qualify; ordinary conversation, exploration or cognitive-memory updates alone do not qualify in `auto`, and no prose or tool-count classifier decides their meaning. Child reviews remain advisory evidence superseded by the root decision.
The explicit call nominates the complete ready result and returns `deferred_to_host_acceptance`, `authoritative=false`; after the whole tool-result block the host advances the same acceptance operation ordinary final delivery uses, and early feedback does not seal the task. Main authors the effective criteria; the paid subject is result bytes, criteria and material effects, so a status question preserves a running review while a new criterion can buy review of unchanged text. A reviewer panel is advice for its author, never a signature on bytes the reviewers did not read: the wave wakes the original Main through the mailbox at its quorum and again when the last slot settles, each wake carrying each reviewer's own verdict (`acceptance_settlement.announce_acceptance_settlement`). The optional answer control is prepared before either ready-feedback shortcut, so a settled panel or queued wake skips parking without losing control provenance. Mailbox readiness is distinct from owner authority: transport and waits wake for all entries, while acceptance source capture, acknowledgement and final sealing use the drain's typed owner/control boundary; system, descendant, independent-task and peer-task messages alone imply no owner revision, while owner/principal messages, quiz answers and typed owner controls keep their handling. Only a still-pending panel may park the turn; `acceptance_settlement.awaited_panel_has_settled` lets the next round, control repair included, run once it settled. Final delivery without re-nomination uses the same task-panel feedback whether ready or pending; delivered feedback stays on its exact `review_runs` row, never in phantom pending state. For a pending panel Main waits (the default; the only option under blocking; Cyber Pro: Main's final response is its decision) or consciously finishes through the delivery control's `pending_review` key; a panel that settled PASS on the earlier revision accepts the task on the reviewers' word (`previous_revision_accepted`, said on the owner row) only when nothing but the answer text changed — same owner source, criteria and material evidence — while a changed subject or any other settled verdict hands delivery to the ordinary path with the collected verdicts in its dialogue history. Before a new panel or any capacity refusal, every recorded still-pending panel of the same root is collected at $0 over its recorded request and roster (`review_dispatch.reconcile_pending_acceptance_runs`), so a re-authored subject cannot discard verdicts the tree already bought; they enter the next panel's dialogue history. Stop, missing custody and unfinished work keep their observed outcomes.
The host-collected `repo_diff` uses one bytes capture (`repo_diff_capture.capture_repo_diff`) for preview (`review_evidence_sections.collect_turn_diff`) and source (`artifacts.materialize_repo_diff_evidence`); a legacy partial preview without its capture stays unavailable, never certified by a later read. git writes to a private spool under a real subprocess timeout; memory holds a bounded prefix per section, the rest stays on the spool for retention. Every gap — nonzero exit, timeout, unreadable repository, memory cut, undecodable run — is stated: an unreadable diff is a gap, never an empty (clean) tree. A root PROVEN plain (`.git` and `HEAD` both absent by lstat of the root alone; any entry, even a broken gitfile, or an unreadable or missing root stays Git-required) is `applicable=false`, complete, without a Git call — no baseline, never a clean tree, no `source_unavailable` partial — and its preparation identity is a bounded content hash (files by content, symlinks by target, no traversal, nested `.git` skipped; special, unreadable or over-bound material is unknown); Git retargeting variables are scrubbed and the root's parent is the discovery ceiling, so no foreign repository is captured. Two never-interchangeable identities leave it: the EXACT bytes, streamed into the private observability CAS (`retain_private_capture` → `repo_diff_capture.write_blob_stream`, 0700/0600, host-forensic, never published; a failed retention is disclosed, not implied), and the REDACTED text projection, decoded and redacted WHOLE per section before any presentation cut so no credential straddles a cut (marked diff lines redact by BODY, file headers kept; a PEM block in a hunk is masked WHOLE by the existing `secret_masking` pattern), carried by the packet, the actor-readable `repo_diff_source_ref` and every owner surface. The packet's `repo_diff_capture` publishes applicability, digest, size, gaps and whether raw bytes were retained, never bytes or a handle; readable hunks of binary, non-UTF-8 and mixed patches stay beside located gaps, an undecodable run being a located mark, not content. Git stderr, exception diagnostics, failed or timed-out stdout and bound-cut sections are withheld from public projections (a failed process can end inside a secret) though privately retainable; complete sections and gap metadata are redacted before cuts. Small complete captures are retained privately by `collect_turn_diff`, larger ones by the exact-source consumer (or locally without one); spools are released on projection or later-section failure, and a half-created spool pair closes what it opened.
Local packet preparation failures have their own material identity (`acceptance_preparation.preparation_source_identity`), fixed before packet budgeting, subject hashing or assembly: semantic criteria, canonical receipt content, artifact bytes, child results, service-finalization facts and the repository or plain-folder content inventory — never the paid delivery fingerprint, tool history or indices, repeated writes, staging of unchanged bytes, candidate prose, owner-message counts or exception text. A proven unborn symbolic HEAD uses the empty-tree base; other Git failures stay unavailable. Incomplete or unreadable material is explicitly unknown; lost availability keeps the last known identity, failed attempts and spent retries (a transient unknown never reopens preparation), and size/mtime never certifies equality. Changed material or criteria reopen it; otherwise one more attempt needs an explicit source-bound `acceptance_retry` (`material_change`, `repair_evidence`, `owner_retry`) after the current failed attempt was exposed: the host resolves the current `owner_source_sha256` or an existing verification receipt's exact content (index = locator only), Main judges substance, rationale and basis are no grant identities, a spent source grants nothing again, and a retry beside a terminal author stance is refused. The paid identity is unchanged.
Feedback exposure binds the responding Main request, incident and attempt. Informed Advisory finish or Blocking unfinished stop retains the answer without rebuilding evidence, bound to that attempt, answer and owner acknowledgement with material rechecked; only root auto/required acceptance qualifies, never child/off lanes or outstanding action gates. Reconcile, dispatch and application keep separate stages and paid/unknown custody; preparation failure authorizes no resend. `loop_delivery._delivery_evidence_state` and its `observed_delivery_evidence` reader serve retention, nomination, post-tool, controls, publication, forced exits and stance merge. Failed reads yield UNKNOWN (empty), preserving the last known fingerprint and revision. An answer retained over unknown evidence remains one candidate across unchanged repeats, with keep allowed but no verified subject: `evidence_current=false`, `evidence_status=unavailable_local_preparation`, empty subject hash, unaccepted binding. Unknown subject rebinding revokes current approval; authoritative UNKNOWN cannot qualify for keep. Forced exit preserves text unaccepted with an unreadable-evidence notice. Publication hash failure has the same projection. The host acceptance pass accounts the incident; retention never opens one.
Local decisions carry `origin=local_acceptance_preparation` without rewriting reviewer decisions. Publication follows §10: settlements remain visible; delayed snapshots cannot erase newer warnings or resurrect resolved ones. Binding and admission alone are not dispatch; a processing failure after handoff or around existing custody carries `origin=host_acceptance_processing` beside the real runs, never as fabricated critic findings. The incident rides the existing acceptance decision and `review_projection.acceptance_incident`, panel or none; `derive_loop_outcome` projects an unresolved preparation terminal as unfinished/blocked in Blocking (forced rails included) while Advisory keeps best effort with stronger rails, unfinished stops and earlier critic failures visible. Cards and history show the concise warning and expanded detail (completion stays the only notification); resolution clears the warning, keeps history, and a first-attempt success on repaired material resolves the previous warning under its original incident id and attempt count. The final cause states both the local failure and the terminal rail without inventing reviewer rework.
A completed reviewer from an older plan wave is attached as a historical supplement through the locked task-result writer and exact producer CAS: it never rewrites the original verdict, aggregate, closure, author dispositions or current-wave pointer, settles its historical cost without another cycle, and reaches a terminal parent without a new model turn. Task acceptance has the same twin: a panel that settles after its task is terminal is collected at $0 over the recorded operation, republished on the task's review projection with a host-composed `late_settlement` note (the verdict, the revision it covered, settled after the terminal; honest about a reviewer whose physical outcome is still unknown) and announced once in the task's room as a System row stamped `card_row="reviews"` (`acceptance_settlement.attach_late_acceptance_settlement`, deduped by delivery id) — a timeline item of the card, read by the next turn from chat history; no model turn starts.
Before an eligible panel runs, `supervisor/queue_transitions.py` closes subtask admission under `_queue_lock` (transport: §10 invariant 25) and `task_status.find_child_tasks` proves the subtree quiescent; a fence that refused or never answered buys no model round — the panel runs on the disclosed rail `admission_fence_available=false`. Early settlement releases the fence; final delivery seals ingress by the queue's typed answer (`loop_delivery._seal_admission_before_delivery`: the worker's own earlier seal, a real owner follow-up, or no answer — delivered with `admission_released=false` and, where blocking reviewers approved, noted `admission_close_unconfirmed`); a changed subject reopens review, keeping the earlier one. Reads use the canonical `budget_drive_root`. The reviewer packet carries verbatim owner directives, the full contract and criteria, canonical deliverable identity, terminal child state, verification receipts, artifact references, touched-skill lifecycle facts (visibility, never a gate) and an explicit omissions manifest; a required component that cannot be assembled makes the affected actor `DEGRADED`, never a silently smaller prompt. The packet is SIZED against the quorum's real windows (the `reviewer_window`/`review_synthesis.quorum_input_token_limit` seam the triad and plan review use), resolved once per task so the bytes cannot drift between the binding build and the staleness rebuild. Non-core sections shed through a DISCLOSED ladder — predecessor authority envelope, trajectory tail and its results, artifact previews, agent-supplied evidence, last a diff preview that keeps the durable `repo_diff_source_ref` — each shed a row in `omissions_manifest`. A slot whose window cannot hold the rendered prompt is a typed `$0 not_dispatched` row while the rest of the panel reviews; a packet still overflowing after every shed stamps `__immutable_core_overflow__` naming the oversized sections and refuses the panel without spending. `__unresolved_partial_artifacts__` withholds packet rows only for a tool result whose exact source is genuinely `source_unavailable` (a retrieving row reads the source itself); a budget shed with a durable, actor-resolvable source ref is an omission, never an unresolved partial.
The packet is delivery-conditional; the FULL packet is not. Every triad row reaches the panel as configured (`reviewer_slot_config.triad_delivery_slots`): a packet (`api_chat`) row receives the assembled packet; a retrieving row receives the route-owned work order `acceptance_retrieving.acceptance_retrieving_work_order` writes onto `ReviewRequest.slot_session_tasks` — the same task-stable contract and output contract (`review_execution.review_output_contract` as `policy["output_contract"]`), absolute pointers to the task's ACTIVE workspace (`review_repo_dirs_for`'s subject root, never the governance repo, as `session_root`) and its result, artifacts and receipts, and the packet in the form its delivery can use. An agent-session row gets the FULL packet (its run is unobserved by the host, so the packet is its only attested view) with the disclosure that access outside the workspace is not guaranteed and a refused read is absence of evidence, not of the artifact; a native inspection row gets the packet WITHOUT its freely degradable tail (tool-trajectory rows and artifact previews, manifested as `retrieving_delivery` omissions) plus the real data root (`policy["native_data_root"]`), as its episode reads those sources itself (`host_file_read_attestation: host_observed`). Reviewer `evidence_refs` resolve against the FULL packet on every delivery, never the rendered projection, so a retrieving row citing a real receipt is clean like a packet row. The wave budget gate prices API money only and DECIDES on one work-order send per paid row (a session row rides the owner's subscription, unpriced); that admission is the whole money rule (no rounds multiplier, no second read-only pricing pass), so a panel's cost is bounded at dispatch by the per-send wallet binding, not predicted; only a floor that does not fit is refused `review_wave_budget_insufficient`, and a packet row's second physical send for format repair never buys a retrieving row a second episode or session. Child-task and `off`-mode acceptance stay advisory and run packet rows only.
Paid identity binds that semantic subject together with substantive nonempty obligation dispositions (`acceptance_paid_identity`); forensic source hashes and ingress counters alone do not buy a panel. A resubmit with the same paid identity reuses its recorded verdict for free: a clean replay can authorize acceptance, a non-clean replay keeps its verdict and the `identical_acceptance_refused` outcome — no repeated payment, and no cosmetic edit needed when real criteria or evidence change.
The configured slots are independent actors with adaptive quorum (`config.adaptive_quorum`: 2-of-N for N≥3, both for N=2, a single reviewer as loud `single_reviewer_no_diversity`; a fewer-responded shortfall stays a loud infra quorum failure); each receives one substantive interaction on its bound route — at most two physical sends for a packet row, one bounded episode for a native row, one delegated session for a session row — and a retrieving verdict is equally authoritative. Transport status, parse status, semantic verdict, criterion support, route, quorum contribution and binding hashes stay distinct, so an unavailable or malformed response cannot masquerade as a negative judgment; a panel that refuses before any transport projects `not_dispatched` on every row and on the panel, and a slot released at the dispatch barrier projects `awaiting` for transport and parse until it settles — distinct from `success`, `timeout` and `provider_transport_error`, never a failure or a verdict. `PASS`/`FAIL`/`DEGRADED` are reviewer verdicts; the host-owned completion decision is separately `accepted`, `revision_requested` or `finalized_unaccepted`, written only by `loop_acceptance._set_acceptance_decision`. A clean quorum supplies critic approval; an informed Advisory author finish supplies separate current-author authority. Neither changes the original verdict. Material FAIL and typed unavailable outcomes reach Main before it chooses how to respond, including after the last paid panel; a no-quorum outcome with real minority findings retains that partial feedback. A terminal technical failure keeps `reason=review_degraded` when no permitted author completion follows.
A clean criterion is evidence-resolved, not only well argued: reviewer `evidence_refs` must be exact members of the packet's enumerable reference vocabulary, and a claim id resolves only through `acceptance_support_refs` linked to a passing host receipt for that claim. Agent-supplied, declared-intent, unattested and non-resolving sections never certify success; an OPEN plan wave binds nothing — its claims are disclosed as `acceptance_claims_source='none_open_plan_wave'` beside a non-binding `plan_claims_exhibit` inside `DECLARED_INTENT_SECTIONS`, so citing it never resolves and the task is distinguishable from one that never had claims. An unresolved reference keeps the actor's record for audit but removes its clean contribution (`criteria_refs_unresolved`). This total, fail-closed resolver is why the task cannot certify itself by echoing its expected outcome.
Actionable findings enter the durable obligation dialogue with stable identity: fix, rebut with an evidence-bearing disposition, or ask the reviewers to declare the issue unreachable or a stable disagreement. A re-raise must name an existing obligation id or is disclosed as new; a valid rebuttal retires the row, an invalid one reopens it with both positions preserved, and the reviewer's `reviewer_rebuttal_response` rides into the next panel's catalog so it can tell "already answered" from "never answered". Each panel receives the bounded `acceptance_dialogue_history` OUTSIDE the hashed evidence material — so reading the history cannot mint a fresh paid binding.
Every received review outcome permits ordinary author inspection, correction or rebuttal, including the last finite panel. `loop_acceptance.merge_agent_acceptance_stance` binds `author_action=finish|stop` to the actual outcome exposed in Main's request and its current tool/owner-directive state; only queueing feedback or predeclaring a decision is insufficient. `loop_acceptance_review._finish_advisory_author` permits informed Advisory completion on unchanged or revised work after criticism or disclosed review unavailability, even when paid capacity remains. The current author hash stays separate from the critic hash/verdict. Under Blocking, corrections may be retained and the attempt stopped, but disputed advancement still requires fresh reviewer approval; a later independently initiated task can resume under its own admission, never an automatic root or budget reset. `stop` means unfinished work and grants no approval. The earlier-revision PASS rule for final task-message delivery above remains distinct; it grants no new commit, plan or skill authority. Typed `dialogue_status` votes retain the reviewer's position: invalid votes abstain, material-free continue is `continue_without_findings`, and zero valid votes is `inconclusive`. A terminal opinion may close another critique of that case, but cannot suppress feedback or choose the author's stop. Auto and Required share these enforcement rules once eligible; explicit task-local response limits and real task rails remain. Cyber retains its own BIBLE P0/P3 authority without rewriting verdict, cost or custody.
Pacing predicts no review duration: a task has three host-owned rails — a deadline, a paid-cycle cap and a wallet — and a panel starts iff the cycle cap has room, the wallet can buy one work-order send per paid row, and MORE than the configured floor (`OUROBOROS_ACCEPTANCE_REVIEW_EST_SEC`, never below 200 s) remains above the finalization reserve; that floor applies only to a new critic. Author reaction uses ordinary remaining task time above the existing finalization reserve, budget, cancellation and round limits, without a reviewer-sized floor or adaptive multiplier. The two refusals remain `review_skipped_deadline_reserve` and `improvement_window_inside_reserve` for their respective owners. `task_pacing` owns both predicates (`review_launch_allowed`, `improvement_pass_allowed`); the launch rule is evaluated ONCE per panel, at loop admission before a panel is built — the paid dispatch claim (`review_dispatch.task_acceptance_paid_dispatch_stamp`) checks cancellation and the paid-cycle wallet only, and no other surface evaluates time. That claim is minted in the locked `task_results.task_acceptance_review_accounting` at first physical reviewer dispatch, and a claim without a recoverable terminal host run is UNKNOWN, never permission to re-dispatch — the double-spend fence. Once launched, a review is clamped to the owner deadline and the task ceiling with the per-send money fence: a panel the deadline cuts is a typed DEGRADED outcome, never a free skip. Disclosed residual: a panel whose evidence build consumed the margin after admission still dispatches and may be cut, and packet repair/retry sends and native rounds mean the total is NOT bounded to one floor wave. Panel durations are `task_acceptance_review_timing` telemetry (`delivery`, `deliveries`, `native_rounds`, `native_rows`) that no gate reads back; the structured review axis is mirrored as top-level `review_status`. A revision row names the pass it starts and the causes the wave recorded, never the aggregate word as its own explanation — a DEGRADED wave that still fed an improvement capsule is not the no-quorum terminal that shares the word.
A forced turn sends the round's exact tool envelope — same schemas, same server-web flag — so the provider prefix stays a cache hit; "tool-less" describes the instruction and the host, which executes no call, so a reply that still asks for a tool, with or without prose, is incomplete on every rail and degrades through the host fallback path instead of publishing as the model's final. Every forced rail uses the common terminal recorder; an eligible task with no panel run records an eligible bypass with zero runs and the rail's trigger. No forced rail can take another model round, so a dangling `revision_requested` terminalizes as `finalized_unaccepted` (`revision_unavailable_on_forced_rail`) on EVERY forced rail through one shared helper, naming the prior reason; `accepted` and `finalized_unaccepted` decisions are never overwritten, no bypass reason is stamped over a panel that ran, and the pair stays outside the blocked-terminal set so the objective remains best_effort. The forced `children_unabsorbed` rail still runs the panel for an acceptance-eligible root with a quiescent subtree, with the undispositioned-children debt in its evidence, keeping a forced delivery distinct from both clean acceptance and no-panel-warranted. A superseded panel remains an audit row with its `run_count` intact; `pending_delivery_acceptance` is a transient eligibility, never a terminal state. The root's post-task phases use the minimal `root_phase_checkpoint`: startup replays only a durable `pending_once` phase, and an indeterminate `running` phase is disclosed as degraded rather than replaying paid work. A deadline or late unread message retains unaccepted work rather than claiming a new subject was reviewed.
#### Headless finalization and workspace patch capture
A workspace task's completion compares against the captured preflight base — task-local commits stay in the delta, not `git diff HEAD` — and the patch is bound to `task_constraint.base_sha`; a moved HEAD fails closed only for `self_worktree` (a shared tree relies on reverse-patch verification), and an unborn repo diffs against the canonical empty tree. `workspace_patch_capture.py` streams the tracked binary diff plus admitted untracked files under the pure rules of `workspace_patch_rules.py` (5 MiB per untracked file; `untracked_capture_veto_reason` is the one composite the patch and the delegated-run execution snapshot both ask), excluding each vetoed entry with a per-file reason; otherwise eligible oversized or binary untracked outputs ride complete manifest+zip file artifacts, and tracked files whose old or current size exceeds 50 MiB stay in the same file-reference manifest instead of a giant Git patch. Generated output (`dist/`, `build/`) is governed by the project's own `.gitignore`, honoured through `--exclude-standard`, not by a host name rule — git-ignored files are outside the capture universe and are not listed as exclusions, because a project whose deliverable IS its build output must not have it silently dropped; a sensitive-looking untracked credential is excluded per-file and disclosed as `sensitive_blocked`. `workspace_patch.json` is written for EVERY workspace finalization, no-change and failed included, and is the truth source for CLI strict-patch (it distinguishes omitted, no-op and failed); `workspace.patch` exists only for `ready_with_changes`. A forked or empty child drive under `data/state/headless_tasks/<task_id>/data` is execution state: the result copies back to the canonical root, declared artifacts rebase to `data/task_results/artifacts/<task_id>/`, and once the canonical result is terminal a late copy-back cannot overwrite the parent-owned terminal marker or the cost/round/token fields. The startup prune (`headless.prune_headless_task_drives`, after prior-process custody) removes a child drive only when the canonical parent is terminal, artifact finalization is terminal, retention has elapsed, the recorded child path matches the expected directory and no child-ref promotion is pending — everything needed after deletion must cross the canonical handoff before a task is presented as settled. The capture manifest explains its acting/admission/empty-tree/capture base and current branch/upstream observations; these do not establish task authorship or change application-patch bytes. An auxiliary comparison uses explicit `vcs_diff(base=..., head=...)` inputs rather than guessing a target. The CLI contract stays in §1 CLI / Headless Boundary.
#### Owner routing verbs
`promote_chat_to_task`, `route_to_project`, `steer_task` and `ensure_project_scope` ride one receipt rail: an act succeeds only once its token-matched supervisor facts are durable in the task result, queue snapshot, annotation or mailbox authority; among several possible tasks the LLM chooses, code auto-delivers only the unambiguous one-target case, and an unconfirmed or stale receipt fails visibly instead of launching a second root. Receipts are retained per `(owner message, routing token)`: an earlier act's receipt stays readable by its token (`chat_annotation_receipt`) while the message's latest row is the UI projection and the picker's liveness test. A KNOWN rejection returns `rejected` with its reason, never a timeout, so `UNCONFIRMED` keeps its one meaning: no matching receipt exists. A receipt proves admission, not completion; an unread indicator proves a visible revision, not memory isolation.
WHO is speaking is ONE host-minted fact on the event (`control_routing._routing_issuer`; the model has no argument): an OWNER TURN (a direct turn the owner door stamped: `is_direct_chat` and `run_origin.owner_ingress`) or a TASK speaking for itself (a promoted root inherits the stamp as ancestry, not as issuer; a root relaying an owner message it just drained; a consciousness wake-up, a Presence event or the auto-resume template on the direct lane, which nobody typed — a wake's promotes mint consciousness roots inheriting its origin, ledger category and autonomy level). An owner turn's steer travels as owner text: `[Message from my human]`, the owner corpus, the generation bump that supersedes a reviewed answer, the room veto from the registry lane of the issuing chat (a Project room reaches its own roots, Main every host-listed root) and the owner acknowledgement. A task's own words NEVER travel as owner text: they go through the one task-message writer `forward_to_worker` also uses, as `independent_task` provenance, to any host-listed active independent root (hidden roots included, no room veto; a Presence sender only to its own binding's work, chapter 12), render as `[Message from independent task <id>]`, enter no owner corpus (`owner_source_sha256` and the acceptance premises stay the owner's), carry no attachments, and are confirmed WRITTEN or refused with the host's reason. A relay keys and publishes its acknowledgement on the owner message it drained; other task acts use their own synthetic receipt id without a chat acknowledgement. Neither receipt grants owner authority; author and target ride the receipt and one `task_message_routed` Logs row, and the receiver's `task_message_injected` row names the sender. Independent roots learn the roster from a `[INDEPENDENT_ROOTS]` TAIL note (`peer_roster.py`, `ROSTER_NOTE_CAP` = 40 rows shown, the cut disclosed), appended only when the roster changed and never merged into a sent row. Rows identify live direct conversations without claiming an owner initiator; messaging one is the model's call. A root may publish one bounded `update_focus(text, source_ref)` record, exposed in grouped notes and paginated `live_roots` and retained by `recent_tasks` on dormant results. A `source_ref` names a reader; `update_focus` answers it through that reader under the caller's registry admission (`disabled_tools` included) and retains the exact answer (≤256 KiB) write-once on the canonical root as `focus.source_handle`, a native `task_source` ref peers read from any drive via `get_task_result(include_focus_source=True, focus_source_sha256=…)` against the PHYSICAL author's record: the digest selects the immutable historical file, so no later focus or retry can substitute its evidence (the roster's `retained_source`); a refusing reader or changed snapshot refuses the focus (`FOCUS_SOURCE_UNRESOLVED`), and the roster drops a settled root's focus. Focus has authored time separate from host observation and carries no TTL or owner authority. Explicit authorized project journal/workpad reads return exactly the requested source; foreign scoped writes refuse; children/Presence keep their capability ceiling.
`steer_task` relays the owner's exact ingress bytes only on the turn's FIRST routing act, while it still acts on the message that started it; the window ends with the latest owner message the turn DRAINED (`ToolContext.last_owner_delivery`, stamped at the loop's mailbox drain; a message only written to the mailbox ends nothing) or with a landed promote/route/steer receipt already on the origin message (a refused or unconfirmed act carried nothing, so the next act still relays). Past either, the turn RELAYS its own words and its receipt is keyed on the message relayed (the drained delivery's own client id, else the synthetic `agent-steer:<routing token>` id), never again on an origin this turn already routed. An agent-authored steer belonging to no owner message earns its receipt under that synthetic id, confirmable through the same `routing_wait` poll; no chat row carries the id and the owner message's own receipt (what a later decision turn reads) stands. Each steer's mailbox entry is keyed by its routing token, so several instructions under one origin are several deliveries while a retried emit of one steer stays one.
Every promote/route refusal names its typed reason; producers holding a fact beyond the code (the workspace family, source and attachment staging, persistence failures, the project fence) add `detail`, the cause with its repair (self-explanatory codes such as `duplicate_task_id` or `worker_pool_unavailable` get no prose), and `_persist_promote_rejection` is the single durable writer. `workspace_admission.workspace_repair_hint` composes the workspace repair from the typed source of the refused folder; an empty Project whose genesis provisioning fails is typed `workspace_provisioning_failed` on the promotion path, never a fallback onto the system repo. The exact Ouroboros repository root named as `workspace_root` is the documented default, mapped to the `workspace="none"` sentinel and disclosed; a subfolder of the repository, the data drive and every `/api/tasks`, Presence and `source` caller stay refused, typed. Who is told follows who can narrate: a tool-issued act gets its receipt, error row and `detail`, and the model narrates; a host-issued act (`host_initiated`) gets ONE typed System row (`task_not_started`, or `task_start_unconfirmed` for an unconfirmed admission) in the chat the owner wrote in, from the one publication boundary `_handle_promote_chat_to_task` wraps around every outcome, never a bubble in Ouroboros's voice. A steering refusal whose owner act carries no owner-message receipt uses `steer_not_delivered` through `supervisor/steering.py::_refusal_row`, in the issuing chat. The Project start row is announced only once the task is enqueued. An admitted promote reports the destination the admission receipt returns, not the requested `project_id`/`project_name` (an implicit scope is re-resolved under the origin claim lock, §6 Project binding by task and by origin); a promote landing in a project other than the one the request's work already has says so: promote stays a free choice, disclosed. A promote by a root whose Swarm planning obligation is still UNMET (`force_plan` with no plan wave recorded, `owner_hurry.unmet_force_plan_obligation`) carries the obligation onto the new root and releases the promoter in the admission transaction (`supervisor/plan_obligation.py`), disclosed in the result; a met obligation moves nothing, and `ensure_project_scope` keeps it.
A routing/promote decision turn receives host-built ground truth (each project's registry `working_dir` plus bounded typed projections of the lane's recent ROOT results and live roots, never raw result text) through the registry row's durable `last_task_result_id` pointer, stamped by ROOTS only, with a bounded scan fallback; only the absent-pointer case writes back (a failing non-empty pointer is usually a split-drive result mid copy-back). A swarm child is not offered (reachable through its root) and counts as its own omission instead of evicting a root. The list is a HINT sized by `runtime_limits` and scoped to the lane that receives it (a pooled task holds none and names any id it read through `recent_tasks`/`get_task_result`); the door is the predicate in §10 on the result itself, never on the caller's room or the landing project, and a helper predecessor or a landing outside the predecessor's project is disclosed in the receipt (`_predecessor_notes`), a free choice like the second-project note. `list_projects`, `route_to_project` and the promote receipt read the project registry through the canonical data root, because a task on a forked execution drive never carries `state/projects.json`. A routing receipt on the SAME owner message is a fact, never a ban: a second root stays the model's judgment.
### Tool capability and execution
`tool_capabilities.py` is the SSOT for the tool classes (core, meta, cognitive-memory, parallel-safe, stateful-browser, untruncated, capped-result, reviewed-mutative) and the child allowlists, which remain deliberately narrower principals. `tool_policy.py` chooses the initial capability set; `ToolRegistry` remains the execution authority; `loop_tool_execution.py` owns timeouts, concurrency, live evidence, result handling and mutative ceilings. Ordinary top-level presets share one built-in name surface: project focus changes the default target, while root policy, runtime mode, task-contract disables, credentials, resources, selected-resource bindings and delegated-child profiles narrow independently — a registered or discoverable tool is not thereby callable for a particular target. Nothing disappears silently: lazy discovery returns an explicit capability omission or `CAPABILITY_UNAVAILABLE`; a policy-filtered REGISTERED tool (a contract-disabled extension/MCP name included) is "hidden by policy: <reason>" (`ToolRegistry.policy_hidden_reason`), never "Not found". A name dispatch cannot run gets one answer (`ToolRegistry._name_miss_result`): its typed reason plus the callable names of the ONE namespace it addresses (`tool_policy.tool_namespace`), or above 12 their count and `list_available_tools(namespace=…)`, never another namespace's or a hidden name. `list_available_tools` (`tools/tool_discovery.py`, bound by the loop) reads `schemas()` in every mode; `loaded` is residency, not availability.
Outcome classification keeps policy refusal separate from execution failure, so an expected authority boundary never becomes the task's headline failure: `user_files_path_blocked`, `cwd_blocked` and `artifact_output_undeclared` are typed non-failure surfaces, while a declared output that cannot be registered is the genuine `artifact_output_error`. The identifier register (`tool_result._EXACT_IDENTIFIER_CODES`) applies the same rule: a typed refusal is recorded as a refusal, never as `ok`. It covers the whole routing family (`STEER_REJECTED`/`STEER_UNCONFIRMED`, `PROMOTE_REJECTED`/`PROMOTE_UNCONFIRMED`, `ROUTE_REJECTED`/`ROUTE_UNCONFIRMED`, `ROUTING_UNCONFIRMED`, `NEEDS_MANUAL_TARGET`), so a promote or route that scheduled nothing is never read as a SUCCESSFUL call; a `_REJECTED` result carries its `(reason: detail)`, and the refusal stays visible to the agent (`is_error`) without degrading the execution axis — the host refused, the agent did not fail. File discovery distinguishes a confinement refusal (an error), a failed listing (a failure) and a path holding nothing — a MISS (`LIST_FILES_NOT_FOUND`, a warning beside `DATA_NOT_YET_CREATED`) that leaves the rest of the outcome uncoloured, because the recovery scan credits only a later success on the same target. The delegate tools (`delegate_start`, `delegate_wait`, `delegate_cancel`, `delegate_answer`) reach the same register from the producer side: `delegate_shared._fail` writes the `ok: false` + `host_code` envelope BESIDE the domain payload (whose `reason` keeps its own vocabulary), so every substrate refusal (daemon, ownership, containment, cancel, subscription window) is RECORDED as a refusal, never counted as a successful call — mapped to `TOOL_REPORTED_FAILURE` (never degrading), a malformed call to `TOOL_ARG_ERROR` (degrades, feeds reflection), never to a timeout or a generic tool error — and a terminal run in state `failed` or `cancelled` is a SUCCESSFUL observation of an unsuccessful leaf.
Tool API v2 exposes neutral canonical names directly (`read_file`, `list_files`, `search_code`, `write_file`, `edit_text`, `edit_batch`, `apply_patch`, `run_command`, `run_script`, `verify_and_record`, service tools, `commit_reviewed`, `vcs_*`, `schedule_subagent`, `manage_schedules`, `schedule_followup`, `wait_task`, `wait_tasks`); legacy public names are neither exposed nor translated. `search_code` uses one basename-glob selector, including comma brace alternatives, before ripgrep or the Python fallback reads allowed files; selected-file count and content matches remain separate facts. The file tools share a path-based public ABI; payload-borne paths (`edit_batch` entries, `apply_patch` targets) miss the dispatch seam that rewrites a `path` ARG, so both ends canonicalize through `tool_access.canonical_repo_relative_path` — one normalization contract keeps a guard from judging `repo/BIBLE.md` while the write lands on `BIBLE.md` — and `_ROOT_ARG_REPO_WRITE_TOOLS` is the single set every repo-write fence keys on. `verify_and_record` (`tools/verify.py`) runs the declared check through the SAME pre-execution guards as the process tools — a post-execution check cannot gate a receipt already written — and appends a host-attested receipt to `<drive_root>/task_results/artifacts/<task_id>/verification_receipts.jsonl` (`expected_match`: substring by default, `exact`, `exact_line`, `json_equals`, `bytes_equal`); it reads only public task info, reports an owner-settings change as a typed note and never auto-reverts it (a revert cannot prove causation), and `delegation_zero_run` writes only `incomplete`/`unknown`, after the custody scan proves no open run.
#### Resource roots and physical file identity
Filesystem tool output is self-locating: results use canonical `root:path` labels and `run_command`/`run_script` echo the resolved `cwd`. A direct Project room selects one active physical folder for reads, writes, edits, process cwd, VCS and delegation; governance stays at `system_repo`, and a missing folder keeps its selected address with an availability note, never a fallback to Ouroboros source. Plain folders support ordinary file/process work and directory delegation (`delegate_directory.py`); Git-only snapshot and integration paths fail typed when their target lacks Git. Physical file resolution preserves the requested address: an absolute path inside any selected base normalizes to that base, an absolute path with no root binds to the permitted root holding it, an outside absolute path is refused before `safe_relpath` can turn it into a similarly named file, and the repo basename-prefix, lineage and delegated-artifact read redirects share the resolver guards and handlers use (`ToolContext.repo_path`/`drive_path`). A Docker workspace's mapped backend absolute address (`workspace_executor.map_backend_path`) is accepted only inside `active_workspace`; other roots and unmapped absolute paths stay confined.
`user_files` is the first-class root for user-visible files under the owner's home (the Ouroboros repo and runtime control plane are refused); `task_drive` is task-scoped scratch; `artifact_store` is task-scoped under `data/task_results/artifacts/<task_id>/`, created lazily on a write or output registration, where deliverables written through `user_files` or declared process `outputs` are copied for audit, and a rewritten user-visible file keeps its prior copy under `task_results/artifact_versions/<task_id>/` (§1 tree). `subagent_projects` and `deliverables` grant `read`/`list`/`search` only; a read-only child reads `deliverables` and its lineage's (parent/root, each also on its own non-symlinked headless drive) `task_drive`/`artifact_store`, never a sibling's; a continuation also reads the `task_drive`/`artifact_store` of the one predecessor its contract names (`predecessor_authority.source.task_id`, one hop, read-only; a child carrying the envelope reads it too) (issue #1232); a top-level task writes the Deliverables container through `user_files`. Process admission keeps the original argv and the prepared physical target, and that one binding serves every other authorized destination; command words, script examples and unknown interpreter effects establish no write intent — the selected Safety Supervisor receives the full task source (§6 Safety and runtime mode) — and the post-execution shell audit is observational: no replay, rollback, interpreter or attribution proof. Invalid path-like prose yields no finding. A failed audit adds a bounded diagnostic (typed results record its exception class in `output_audit_unavailable`) while keeping stdout/stderr and completed exit, signal, timeout and runtime facts; it cannot invent a spawn failure or undeclared output. Computed targets stay a disclosed parser limit, not grounds for a new semantic scanner.
#### Credential fence and byte masking
Ordinary-mode protection of credentials, Bible and identity keys on PHYSICAL owner locations, source operands and actual VCS metadata, never on names: ordinary project names such as `tokens.json`, `credentials.json`, `.nojekyll` or `.well-known` do not identify host credential stores, and an owner's input or output is never refused on a SUFFIX or a WORD in a file name — on attachment ingest, on export/Deliverables, on a `user_files` mutation or in the git lanes — beside the exact-name and retained dotenv-tail rules. The enumerated physical owner stores (`credential_shapes.owner_credential_locations`) are the mutation-fence authority; path-selected attachment ingest checks exact credential leaves and credential/control directory components; ordinary `.config`, `Library`, `settings.json` and exact `~/.ssh/config` mutations need no special authority; ordinary git capture adds content evidence (`workspace_patch_capture.pem_private_key_reason`, shared by the patch, the coop checkpoint and the attached-folder snapshot). This is not blanket protection for all credential stores: unlisted locations such as `.cargo/credentials.toml`, `.terraform.d/credentials.tfrc.json` and `.kaggle/kaggle.json` keep ordinary root configuration access, and children keep source visibility for ordinary `auth/`/`tokens/` paths and public PEM certificates. `core_secret_paths.restricted_data_roots` anchors owner/control checks on the child, canonical parent and configured admission roots, shared with vision/media; runtime secret/control entries keep their predicate even inside the project. Reads and searches share ONE byte masker in every scope: restricted repository views mask known token formats and complete PEM private-key blocks while ordinary long identifiers, hashes and source bodies pass, so search and read deliver the same bytes; owner-home output does not mask opaque runs either, and known credential bytes stay masked on delivered views.
#### Web access mechanisms (three distinct paths — do not conflate)
The three web paths differ in who chooses the query, which model reasons, which authority performs the fetch, and where evidence is recorded:
| path | reasoning model | fetch authority | evidence |
|---|---|---|---|
| Main-loop native search (provider server tool, attached only when the main-loop setting allows) | the same solve model decides whether to search; no second model enters the scaffold | provider-side | citations and request counts fold into `llm_usage` and the task's host-attested retrieval fact (context for acceptance, never a criterion by itself); absence means only that this path recorded no search; the provider-side query is not available to the host and is not claimed as logged |
| `web_search` function tool | a provider-backed call can introduce a second model and its own cost | the configured web-search route or a keyless retrieval backend, executed by `ToolRegistry` | arguments and bounded results in `tools.jsonl`; never native-search usage on the answering call |
| Browser tools (`browse_page`, `browser_action`) | the main loop | a local stateful Playwright session that can fetch or act on arbitrary pages | arguments and result previews as tool evidence; browser state and local action semantics are not equivalent to either retrieval path |
For ordinary delegated profiles the URL, route, private-range and control-plane guards apply to local-readonly and acting children alike (target policy: `browser_policy.py`, §6 MCP and browser-facing external tools); a Cyber-effective acting child inherits the selected access while explicit task restrictions remain. A concrete private origin is reachable only through host-established `resource_policy.allowed_origins`: ordinary Main/root scheduling names it in `schedule_subagent`, delegated and external Presence callers can only inherit or narrow it (`tools/control_subagent_spec.py`), and the origin normalizer rejects paths, credentials and wildcards. Of the delegated profiles only a valid acting child may run JavaScript `evaluate`, on its current page; acting children keep the pre-existing shell-to-loopback `/ws` route without WebSocket authentication. This separation is methodological authority: an evaluation or acceptance claim must name the path used.
#### Context fitting, retry, and compaction
`config.get_context_mode()` is the effective Main sizing/rendering source and `config.get_owner_context_mode()` the persistent owner-intent/P3 source; they differ only during the auto-Low compatibility window (`context_mode_compat.py`). `prompts/SYSTEM.md` and `BIBLE.md` are tier-0 and full in every projection (`context_layout.TIER0_ALWAYS_FULL`). In Max, `ARCHITECTURE.md` is full-resident for every Main task class because it is Ouroboros's capability/tools/access map (economy comes from dropping the handbook for external work, never from hiding the map), and `DEVELOPMENT.md` follows the active repository binding: full when the work targets Ouroboros's own body, a visible on-demand pointer for a bound external workspace or an API/CLI/scheduled external surface (explicit `context_requires_development`/`context_requires_self_body_docs` flags override the binding). In Low and Nano BOTH books are replaced by their book navigation (`context_layout.book_navigation`: the authored chapter introductions plus each chapter's heading index with line ranges into the PHYSICAL chapter file, because the composed book exists only in the prompt and an offset into it would address the wrong bytes), and a subagent child receives that navigation of both books in every mode, against BIBLE P1's wording (issue #1026). Disclosed residual: Low and Nano do not consult the explicit per-task `context_requires_self_body_docs` flag, although the per-task requirement should win in either direction (issue #1019).
The one path that trims tier-0 is the local-model overflow compactor (`llm_local.py`, re-exported by `llm.py`): when a local window cannot hold the prompt it replaces every `## ` section outside its keep list with an omission notice — SYSTEM.md sections (hence the prompt's load-bearing floor lives in its preamble) and the knowledge index (the one disclosed exception to the always-loaded index) — keeping Identity, Shared understanding and the dynamic memory sections (Scratchpad included), and raises `LocalContextTooLargeError` rather than truncate silently. Disclosed defect, not design: BIBLE.md's own `## Principle` headings are not on that keep list, so the compactor also cuts the principle bodies (issue #1018).
`context_fit.py` renders Max, Low and Nano projections from one immutable core and measures each ordinary Main candidate on one labelled density basis (`context_fit.estimate_context_prompt_tokens`) against the selected route capacity, while budget reservation keeps RAW chars because over-counting money is the safe direction. Owner Nano uses its compact source view and bounded output target; Low adds an elastic 200K total-context economy target (`context_budget.py`); crossing either target is never a synthetic failure. With unknown capacity, owner Max gets one honest Max call while owner Low or Nano may reclaim toward its known target. Predicted pressure retains the selected document projection and may request one deficit-sized mutable-history pass; only an actual provider overflow authorizes task-local Low, and one same-route semantic recovery is permitted only when the final candidate has the same route, round and response reserve and strictly fewer context-bearing bytes. Owner mode and P3 applicability never change. `context_health.build_health_invariants` runs ONCE per task attempt, so the Health Invariants block is a task-start snapshot; its delegated-run lines are visible to every task but ownership-aware, so a non-owner is never handed a call that structurally refuses (§6 Delegated subagents).
`context_compaction.py` is a requested materializer, not a second threshold, timer or retry authority. The helper-summarized path selects a positive-reclaim prefix of completed assistant-call plus contiguous matching-result units; owner turns and malformed, interrupted or visually opaque units stay verbatim, and an unfinished Anthropic native unit is ineligible. A non-empty selection first writes an exact private checkpoint, then summarizes complete, gap-free hashed map/fold input; independent covered units may apply while a failed unit stays raw, and opaque custody never enters summarizer text. No eligible positive reclaim means no checkpoint, no summarizer call and no transcript mutation; the one route+round latch prevents repeating the same automatic pass.
Main also uses this materializer for an actor-authored working view. At the existing successful-response boundary, after retry/reprepare/fallback and before returned tools execute, `compact_context.record_context_view` retains the canonical source and active schemas; the acceptance observer still sees the separate send projection. `compact_context` can retain exact inspected unit IDs, restore checkpoint sources, or pair an authored note with an explicit `keep_last_n` count. IDs take precedence; note-only keeps raw units, and count-only uses the helper path. Original source bytes remain retrievable, newer owner/tool messages stay untouched, and no helper model rewrites the authored note. The completed tool boundary measures the full proposed messages, schemas and receipt through Main's existing fit arithmetic; estimates are disclosed, not new capacity or economy-target admission gates. Actual overflow retains its existing recovery. Publication updates the view and schemas together; no-op observation leaves cache identity and the live turn unchanged.
Ouroboros stores one provider-neutral, function-shaped conversation. OpenRouter omits the neutral `tool_choice="auto"` while building the physical request, before request-wire identity is bound, so an implicit API default cannot become an unsupported parameter under strict routing. Direct OpenAI agent/tool traffic stays on Chat Completions: `openai_chat_custom.py` translates only the physical copy into Chat custom tools and normalizes returned custom calls back into canonical function calls; the dialect ladder is fixed at custom/function/none (ordinals 1/2/3), only an exact dialect rejection advances a rung, the custom→function control-flow fallback is never learned as a durable dialect action, and the compatibility layer never migrates the conversation to Responses or changes the model/provider/API surface. Direct Anthropic tool turns keep a private route-bound receipt with the complete native assistant `content`, replayed byte-for-byte on the immediately matching same-route continuation and scrubbed on any provider, endpoint, API-surface or model change; opaque values never reach summarizer or public projections. Owner-requested `none` is sent as `thinking.type=disabled`; no guessed legacy `budget_tokens` mapping is invented.
`request_wire_recovery.py` is the single provider-neutral adaptation driver for both exception and HTTP-200 body-error shapes, keyed on one credential-free exact route/profile: only bounded `set_value`, `drop_field` and registered `replace_dialect` actions, at most eight per call. A learnable reactive action becomes durable — the 14-day cross-process store `state/request_wire_compatibility.json` — only when a semantically valid normalized response is bound to the exact settled physical-attempt capture; provider prose cannot switch routes. Disclosure needs only the identity half: the candidate a RETURNED response carried is disclosed as `usage.request_wire` once its capture passes the exact-identity check (`request_wire_attempt.validate_wire_attempt_identity`), because a ledger write that fails after the provider answered cannot unmake which request shape was sent; the finalizer's place on the successful-return path is what proves a response arrived, and money keeps its own authority in the attempt ledger, so an unsettled attempt discloses but never teaches. Disclosures aggregate in order as bounded `request_wire_history` — not the physical-attempt ledger; `state/usage_attempts.jsonl` remains the complete monetary authority. Private custom-argument receipts never enter stored history or public observability.
Retry budgets are failure-class specific: empty/incomplete responses and transient 429/5xx failures may retry the same model with deadline-bounded backoff; auth, quota, permanent bad requests and confirmed oversize fail fast; exhausted compatibility returns to the configured model fallback chain rather than becoming a second router. `LLMClient` keeps leading system messages authoritative and demotes later notices to visibly marked user notices.
Transport failures are classified by physical-attempt custody; each class has its own owner and rails. A REMOTE pre-dispatch failure is `transport_unavailable` (`loop_llm_call.classify_llm_exception`: released custody, $0, on a non-local provider). It takes one physical attempt per call; pacing belongs to the round-level wait episode (`ouroboros/loop_transport.py`), which tries the fallback chain once only when `USE_LOCAL_FALLBACK` makes it local (remote candidates never dial over a proven dead egress), waits (durable `network_wait` events, progress that keeps the idle rail alive, owner-interruptible backoff), then redials the SAME round for free. A managed task waits as long as its existing rails allow — owner deadline minus the dispatch-admission reserve, budget, Stop, absolute ceiling; never a new setting, and `OUROBOROS_TRANSIENT_RETRY_MAX` bounds only transient PROVIDER failures — and its exhaustion is its terminal result. The interactive class (every turn stamped direct-chat: owner chat, Presence turns, consciousness wake-ups) waits the same way but has no queue rails and ordinarily no owner deadline, so its episode is bounded by the task idle timeout (`OUROBOROS_TASK_IDLE_TIMEOUT_SEC`) measured from each episode's entry: it limits idle WAITING only and never cancels in-flight work, the shorter of it and an explicit deadline binds, the durable `ended` detail names the expired rail (`interactive_wait_window_exhausted`, or the deadline's own), and exhaustion emits an owner note alongside the provider-outage terminal. Closing `network_wait` rows exist only for cooperative exits — an external kill, Panic, crash or direct-turn hard stop leaves none — so correlate a confirmed `task_done` or durable terminal result; silence never proves completion.
The mid-flight class is `provider_outcome_unknown`: a dispatched request with no terminal provider fact keeps its unresolved monetary bound and is never resent. A NEW logical request afterwards needs a unique host-attested input, granted per caller class. Ordinary managed tasks and native API children use the upstream-observation continuation (§6 Review delivery), whose wait admits a marked new physical attempt on observed connectivity. A configured-session supervisor with exactly one live delegated leaf latches a durable unknown-provider hold (`ouroboros/delegate_hold.py`, closing any transport episode as `network_wait ended / hold_latched`) and parks in the ordinary `supervised_wait`, resuming in a NEW round only on a meaningful leaf wake whose receipt is that input (control wakes and an ineligible hold grant nothing; budget admission stays fail-closed on the unresolved bound); without a hold it uses the managed continuation. Interactive primary rounds get the bounded transport-death repeat below: no-effect by construction (no Host-executed tool or correspondent action can stem from the failed attempt; priced read-only provider-owned retrieval may rerun; §12), not an event retry. Every other surface — forced-final, fallback-chain candidates, review actors, safety, probes, web search, consolidation/summary/reflection, external-harness delegated runs — keeps no-resend (budget 0, or never enters `call_llm_with_retry`); an unknown outcome not typed as a transport death is never resent.
The transport-death repeat rail: a typed death (`transport_custody.is_retryable_transport_death` — httpx `ReadError`/`WriteError`/`RemoteProtocolError` through the explicit `__cause__` chain, or a requests wrapper carrying `ProtocolError`/`RemoteDisconnected`; never a timeout, a provider status/body error, a pre-dispatch failure, a local provider or a loopback route) lets the PRIMARY main-loop round dispatch alone (`_dispatch_round_model`, `transport_death_retries`) repeat the SAME logical request at most twice per round. Each repeat is a NEW physical attempt with its own ledger row, and the earlier rows stay `unresolved` at their upper bound (a fully dead round reserves up to three: honest accounting over a cheaper rail). A round record (`execution_id:round:round_idx`) counts the repeats and names the class the terminal reports; `llm_non_retryable_same_request` marks only the exhaustion. A current typed `finalize_now` control or the deadline can refuse a not-yet-sent repeat (`finalize_control_pending` names the control case); the round then exits through the unknown no-resend terminal. A round holding a repeat record sends nothing but those typed-death repeats: a repeat failing with any other class ends the round on the same terminal — no transient burst, empty-response retry, compaction retry, forced-final dial, fallback chain or local-only pass — because the earlier request may still be live; only a repeat that never left the host (released custody) stays the free wait episode's to redial, and when that window closes the terminal is still `provider_outcome_unknown_no_resend`.
Retry rails nest, each with its own owner and bound; the table exists so the multiplication is visible in one place (a physical send is always its own `execute_physical_attempt` lifecycle, whichever rail asked for it):
| Rail | Owner | Bound | What it repeats |
|---|---|---|---|
| SDK `max_retries` | client construction (`net_transport.py`, `llm.py`) | 0 everywhere | nothing: every physical send is visible to the ledger |
| request-wire recovery ladder | `request_wire_recovery.py` inside one `LLMClient.chat` | at most eight typed compatibility actions per call | a rejected request in a corrected wire shape |
| transport-death repeat | `call_llm_with_retry`, primary round dispatch only | ≤2 per round, 4 s then 8 s deadline-aware backoff | a dispatched request that died with a typed transport death, as a new attempt |
| transient burst | `call_llm_with_retry` | `OUROBOROS_TRANSIENT_RETRY_MAX` attempts per call (the outer ceiling of the same loop) | typed 408/429/5xx and empty/incomplete responses |
| round redial | `loop_transport.py` wait episode | a managed task: owner deadline, budget, Stop, absolute ceiling; an interactive turn: the task idle timeout measured from episode entry (`OUROBOROS_TASK_IDLE_TIMEOUT_SEC`), or an explicit deadline when that is shorter | a released ($0) pre-dispatch failure of the same round |
| served-model redo | `llm_substitution.py` inside one `chat_claudexor` | `OUROBOROS_SERVED_MODEL_REDOS` per call, and none under a pin, an admitted candidate, a spent send budget or a spent owner window | a round the engine says another model answered, as a new operation. It nests inside the rails above, so a repeat granted by one of them starts a call whose redo budget begins again |
| review physical rail | `review_substrate.py`, `review_native_episode.py` | 2 sends per packet or session actor on the P3/acceptance surfaces; a native-retrieval slot carries no send count (its bounds are the episode's own, enumerated below); no rail elsewhere | a released review send, never an unknown one |
The cached OpenAI-compatible clients, the no-proxy per-call clients and the web-search clients share one transport factory (`net_transport.py`) that sets platform-guarded TCP keepalive socket options — on Linux and Darwin the idle threshold, probe interval and probe count from `config.py` (Darwin: `TCP_KEEPALIVE`, not XNU's 75 s × 8 default), on every other platform (Windows included) `SO_KEEPALIVE` alone — so a NAT/VPN mapping silently dropped during a long silent reasoning stretch is detected by kernel probes instead of hanging until the read timeout. When any proxy httpx would honor is configured, the cached and web-search clients skip the explicit transport (httpx env-proxy mounts require it absent). Disclosed residual: proxy-routed installs, the Anthropic-native `requests` lane, every non-Linux/non-Darwin platform and a handful of library clients run without keepalive tuning.
Main-loop model cognition is a separate typed in-flight fact: immediately before each exact provider-call seam the worker sends a direct supervisor `started` event bound to the exact attempt, execution, round, call and retry, and every terminal sends the matching fact, so a stale terminal from an earlier retry or attempt cannot clear the current row. The supervisor keeps only that process-local active row, with no elapsed-time expiry, consulted only by the idle predicate. `OUROBOROS_LLM_TRANSPORT_READ_TIMEOUT_SEC` (default 2700 s) is a configurable dead-socket bound, not a cognition deadline; deadline, budget, cancellation and the absolute ceiling remain independent hard axes.
OpenAI-compatible response choices keep their outer `finish_reason` as the bounded observational usage fact `response_finish_reason`; it never enters canonical history and changes neither the empty-response classifier nor retry policy. The trusted provider canary accepts a schema-valid native tool call even when the assistant also returns text (length/hash recorded as warning telemetry); a malformed native argument, invalid schema or missing call remains a red contract failure — no prose parser, salvage, provider hop or unbounded retry.
Prompt caching is stable-first: governance and task-stable contracts precede mutable evidence, review builders disclose the stable/dynamic boundary, untrusted payloads stay outside the governance cache block, and `context_fit.seal_task_transcript` owns the single message-side breakpoint (the task-contract boundary until a rolling tool-result seal qualifies, migrated in the same call so the four-breakpoint cap can never drop it). OpenRouter derives its sticky session from stable policy/model and a copied first-user content projection without cache or host-only metadata, so moving a seal in either direction does not change the session. Explicit review affinity and provider-reroute rotation retain their precedence. Provider-specific cache hints are sent only where supported and receive one exact retry without the rejected hint; rejection evidence is durable and route-specific. Gateway response-cache recovery is narrower and reactive: only a main-loop `provider_incomplete_response` arms a fresh-response request for later attempts, and only the generic `openai-compatible` route renders LiteLLM's `extra_body.cache.no-cache` control — direct providers and OpenRouter never receive it, and no URL/model heuristic exists.
#### Vision and local image evidence
`analyze_screenshot` and `vlm_query` are bounded secondary-model calls through `LLMClient.vision_query`; `view_image` attaches a local image natively. Send-time image routing (`vision_routing.py`: inline vision, captions or placeholders) works on a per-send copy and never mutates canonical history. Model-wait reprepare starts from that canonical source, preserves original images and prepares the new route's send projection; a typed invalid native continuation retains its existing reset without replacing the source with captions. Image payloads are validated, capped/downscaled and confined to readable roots derived from the Tool API policy matrix plus the protected-artifact rule. `vision.attach_local_image_to_context` is the single attachment seam for explicit `view_image` and the typed `auto_attach_image` opt-in — same durable copy, trust boundary and live-image eviction budget; auto-attachment failure is non-fatal. The loader's own refusals are typed `TOOL_ARG_ERROR` results (a policy owner's refusal keeps that owner's typed marker), so a known failure is recorded as one. Vision/local-media tools are not web tools and may be withheld by the task contract. The uploads and skill-output roots resolve through the task's canonical data owner and count as independent image admission before user-home confinement, per-path secret/owner-state/project-store/protected-artifact checks still applied; an admitted image is copied into the task's `uploads/views` before same-round attachment, so the original and its retained copy both stay available to explicit readers.
#### Background consciousness and Evolution
Background Consciousness is Ouroboros taking a turn in its own Main chat when nobody wrote to it. There is no second mode: a wake-up is an ORDINARY Main direct turn (`_is_direct_chat`) with the same system prompt, context builder, tool loop, chat log, activity block and usage ledger, so consciousness structurally cannot count its context, its money or its tools differently from Main. Only three things are new: an alarm clock, a wake-up MESSAGE that stands where the owner's words would be (`prompts/CONSCIOUSNESS.md` rendered by `consciousness_wake.py` from existing readers, never silently cut), and the three owner settings beside the existing interval bounds. The alarm (`consciousness.py`) owns no thread: the supervisor pass calls `tick(now)`, whose ordered typed outcomes are disabled → a live wake → a live owner direct turn → not yet due → the rolling-24h allowance (an unreadable ledger is the disclosed skip `allowance_unknown`, never a silent block) → an owner chat must be bound → launch through `supervisor.workers.handle_wake_direct`. The next wake is last finish + interval (`set_next_wakeup` or the default, clamped into the owner's [min, max]; a failed or unadmitted wake doubles it up to max), `notify(reason)` pulls it forward to one shared floor, an owner message never wakes it, and a wake's own finish never re-arms it. Version awareness stays ambient: every task's and wake-up's Runtime context carries `official_update` (running version versus the official target at the last fetching check, plus the update letter, `update_letter.py`); no check forces a wake, and mentioning an update to the owner is the mind's judgment.
The rendered wake message puts the fresh cause and settled facts before outstanding cards. Failed/cancelled or host-authored Presence terminals stay visible even on the direct-turn lane, with their stored outcome, task pointer and deferred work reference; ordinary owner direct turns and the wake's own turn stay excluded. This exposes an external failure without starting another wake or prescribing recovery: consciousness inspects the task and actual delivery through existing tools. Cards in `open` and `expired_terminal` retain late-answer semantics, explicit state, age and bounded preview. `consciousness_last_wake_at` preserves the factual boundary across restart; the best-effort in-memory trigger may fall back to heartbeat.
Observe retains a narrow positive path. `consciousness_authority.OBSERVE_ARGUMENT_NARROWED` is the set of mutating-table names Observe still SEES, because each has a genuine read-only or own-work use that withholding the name would destroy along with the mutation: `schedule_subagent` and `delegate_start` for read-only research, `write_file`/`edit_text`/`journal_write`/`workpad_write` for its own notes, and `cancel_task` for the children it started. What such a name may be ASKED to do is decided by one predicate, `observe_argument_refusal`, which the dispatch guard consults directly — no fail-open wrapper, because a policy that cannot be consulted must not read as permission. It refuses a write root outside `task_drive`/`artifact_store`, a mutative `write_surface`/`may_mutate`, and a delegated session asked for `workspace_write` or a payload selector; an OMITTED access is the read-only shape (`_derive_authority` pins it) and stays allowed. `cancel_task` is the one name whose narrowing needs durable lineage, not arguments, so it rides the tool's EXISTING own-child custody check (`join_ledger._cancel_task`), the same seam a delegated task meets: Observe may stop the children it started, never the owner's running work. A read-only child may delegate further read-only descendants under the inherited depth cap but never passes a mutating delegation budget onward (`delegation_may_mutate`). `manage_schedules` (`tools/followup.py`, beside `schedule_followup` — one module for the two tools over one table) is argument-aware, not name-wide: the root turn may list and change, a delegated task may only list, and a Presence conversation gets neither — Presence authority is not expanded here.
Three facts distinguish a wake-up from an owner's turn, and all three ride its metadata. Its ORIGIN is `metadata.initiator == "consciousness"`, inherited by everything the wake starts: it labels every surface; the wake speaks to other tasks as a task (`[Message from independent task <id>]`), never as the owner, because no owner door stamped it. Its AUTHORITY is the owner's autonomy level (`metadata.consciousness_autonomy`, owned by `consciousness_authority.py`): Observe researches, keeps internal memory/project notes, stops the children it started, inspects and governs existing schedules through the audited narrow surface, and may delegate read-only research, while the dispatch guard refuses owner-file, source, skill/settings and publication writes; Act (the default) is everything the runtime mode allows except editing Ouroboros's own code and prompts, evolution, restart and settings; Full adds evolution. A level has two structural consequences — an exception list in `task_contract.disabled_tools` and a per-task `runtime_mode_cap` of `light` below Full — plus, at Observe, the argument-level narrowing of the names it keeps (`observe_argument_refusal`); for a consciousness-origin task that list binds at DISPATCH ONLY, so the wake's tool schemas and cached prompt prefix stay byte-identical to an owner turn's while a withheld call is refused when it is made: the typed dispatch refusal is the mechanism, and the wake message names the level. The stricter of install mode and cap binds at the repo-mutation, protected-path, `start_service` and shell-write gates, at the acting-child write surface and subagent patch integration, and in a started task's Runtime block, so the owner's level holds even on a Cyber Pro install; a capped tree is refused a self_worktree child and a system-repo patch integration in every install mode, the mutative-subagent toggle included. Its MONEY is one rolling 24-hour allowance covering the wake-ups and every root tree they started, read off the usage ledger (`consciousness_allowance.py`) by the alarm before it wakes and by the ONE admission door (`supervisor/queue.py::enqueue_task`) before any work it asks for starts — no separate "consciousness task" type, only the existing tools through the existing door; a Presence cycle a wake starts (Act and above) is an independent root outside this allowance. What is left of the allowance (capped by the owner's per-task cap) is the wake tree's graceful ceiling (`root_cost_ceiling_usd`, inherited by its members): the in-task stop lands a planning margin before it while the ledger fence stays at the owner's per-task cap, so a wake never dies before its first call; a remainder at or below that margin counts as exhausted.
An active campaign owns an explicit objective, campaign id, transaction and task claim (`supervisor/evolution_lifecycle.py`); `evolution_mode_enabled` is only its scheduling projection. Dispatch, review, commit, publication and restart revalidate that exact authority: a restored row without a live uncommitted claim is cancelled, a reviewed commit binds to the claim by exact SHA before publication, and if authority changes after commit the commit moves to a private inspection ref and any attempt-created tag leaves the normal namespace, without resetting concurrent index/worktree edits; exact terminal replay resumes only incomplete effects, never double-counting. The restart-verify claim (`state/pending_restart_verify.json`; one writer helper serves the supervisor's evolution restart and the agent's `restart` tool) is written whether the supervisor restarts automatically — `OUROBOROS_EVOLUTION_AUTO_RESTART` off skips only the restart — so the owner's manual restart verifies the cycle by exact claim rather than by the weaker markerless reconcile.
Campaign cleanup is deterministic and custody-aware: a no-op or abandoned cycle may restore the transaction base while preserving dirty/ahead work in recorded stash or local refs, and cleanup is skipped when another task, a live test or the operator kill-switch makes reset unsafe. A byte-identical diff whose last terminal is a review-verdict block is refused for free from the first block; changing the diff, or a rebuttal whose content hash is new to the streak, creates a new paid reviewable case, and the shared cycle cap bounds paid triad+scope cycles per root task. Checkpoint and outcome rows (`evolution_checkpoints.py`, append-only) preserve git/memory identity, cost, rounds and explicit omissions so the promotion loop learns from failed and absorbed cycles. An agent-requested restart first drains heartbeat-fresh running tasks up to its configured bound, then fails closed rather than cut another task silently.
Post-task evolution is owner-gated and default-off (`post_task_evolution.py`): the worker may recommend one backlog item by writing a durable request — never enqueuing or enabling evolution itself, and never from an evolution, deep-self-review or subagent task — and only the supervisor's idle tick converts it into one normal campaign cycle. Its CLOSED-objective projection distinguishes absorbed/abandoned objectives from this attempt's typed review-cap or unavailable-review stop: a temporary review block does not permanently prohibit reconsideration. Retained task sources remain discoverable, without scheduling a retry or resetting campaign counters. An owner stop closes the campaign, clears queued requests, persists the `evolution_owner_stopped` flag, and cannot be autonomously reversed. Evolution is hard-blocked in `light` runtime mode at the owner and post-task start entry points, again at idle campaign enqueue, and again at assignment (§5); it uses the normal task/review path in `advanced` or `pro`.
Loop self-checkpoints remain plain user-message reminders: a second structured, tool-less reflection protocol in the hot loop produced unusable records and destroyed cache continuity. The reminder's recent-call list carries each call's recorded outcome and, for a failed call, the first line of what the tool answered — host facts only, no threshold, no repetition verdict and no stop. It is built from `llm_trace.tool_calls`, so name, arguments and outcome come from ONE record: joining the visible transcript to the trace by call id cannot be made sound, because providers reuse ids (GigaChat answers `call_0` every round) while compaction removes whole units from anywhere, so a surviving call was printed carrying an evicted namesake's error — by id, and still by (id, tool) when both calls were the same tool. Without a trace the transcript's calls render with no outcome, not a borrowed one. Durable learning belongs to the post-task reflection flow (§6 Post-task reflection).
### Safety and runtime mode
Tool dispatch uses the registered capability, the explicit task/resource contract and one prepared physical target. Explicit task resources and the actual Git target are host facts, distinct from any semantic judgment about the call; this is not an OS sandbox.
Runtime mode is effective Access and a self-modification ladder. Light refuses mutation of the Ouroboros repository and control plane (`LIGHT_MODE_BLOCKED`; skill payloads under their own root stay editable) while user deliverables outside them remain buildable; Advanced evolves the application layer but not the protected set owned by `runtime_mode_policy.py` — `SAFETY_CRITICAL_PATHS`, frozen contracts, release/managed invariants, one policy shared by the registry, git tools and gateway guards, refusing typed `CORE_PROTECTION_BLOCKED`; Pro may write that set, announced by `CORE_PATCH_NOTICE`, and its commit still passes the normal triad + scope gate. `runtime_mode_policy.mode_has_unrestricted_agency` derives Cyber authority from effective Access; saved next-boot settings never rewrite a running task or a prior request. In Cyber Pro, internal review and Safety findings are advice under BIBLE P0/P3: configured enforcement, independent assessments and the decision to proceed remain distinct facts. Real unavailable resources, malformed executable forms and external errors remain failures, not invented success, and an explicit read-only assignment keeps its meaning.
Process admission preserves the original argv and the resolved cwd. Command words, quoted source examples and unknown interpreter effects are not proof of a forbidden action: `shell_parse.py` and `tools/write_shape.py` supply observed targets and syntax facts, never permission judgments. The Safety Supervisor (`safety.py`) judges each call from the complete retained initial and later owner directives, the task contract, the task constraint and the prepared resource target; the bounded newest-first recent-conversation view (`_SAFETY_CONTEXT_CHAR_BUDGET`, omission marker reserved inside the budget) supplements that source and never replaces it. Coverage is `OUROBOROS_SAFETY_MODE` Full/Light/Off over each tool's `TOOL_POLICY` entry (`skip`, `check`, `check_conditional`; an unknown tool falls to `DEFAULT_POLICY` = `check`): policy-skip tools and established safe conditional subjects make no call in any mode, Light additionally skips conditional checks, Off skips every Supervisor call, and every waved-through check leaves a durable `safety_mode_skip` event. `full → light → off` is a decreasing coverage ladder, so outside Cyber Pro a downward step is owner-only (`POST /api/owner/safety-mode`). A selected assessment runs on the light-model route. No second consent store, appeal pass or repeated execution exists.
Foreground commands, scripts and run-kind verification accept `env_from_settings` through the same Settings-selection authority as service startup. `tools/process_facts.py` resolves references once after admission and masks diagnostics before result and receipt persistence. Foreground execution keeps its inherited environment; Docker transports target values through inert host aliases, so target PATH or DOCKER_HOST cannot reconfigure the host client. Verification matches raw output before masking. Python selection uses the effective child environment without a version probe; wrapper and non-local backend resolution remain unknown when not established. The existing Node health check runs after gates with the selected PATH.
#### Safety Supervisor outcomes
The first line of a Supervisor result is the typed outcome. SAFE passes silently; SUSPICIOUS allows the call with `SAFETY_WARNING`; DANGEROUS, an unrecognised status, a failed check and an answer still unparseable after one repair retry (`safety_parse_retry`, `safety_parse_failed`) block with `SAFETY_VIOLATION`. Outside Cyber that warning/refusal stands; Cyber records a negative or unavailable assessment as advice (`SAFETY_ADVICE`, durable `safety_advisory` event) without manufacturing SAFE. A provider 429 is an infrastructure fact about the Supervisor, not a verdict about the call: after one deadline-capped retry the call is blocked with the typed non-verdict `SAFETY_UNAVAILABLE`, which tells the agent to retry the same call rather than reword it, and a confirmed storm arms a process-local latch (`_SAFETY_STORM_COOLDOWN_SEC`) that answers in-window checks without a provider call, because re-probing from the highest-frequency light-model consumer would amplify the storm it reports; each is a durable `safety_check_rate_limited` event, while a structured insufficient-quota refusal stays PERMANENT and blocks as a verdict. The serialized subject has its own 250,000-character budget (`_SAFETY_SUBJECT_CHAR_BUDGET`, rendered `ensure_ascii=False`): an over-budget subject is refused fail-closed with `SAFETY_SUBJECT_TOO_LARGE_BLOCKED` plus a durable `safety_subject_too_large` event — never truncated, because anything past a cut would run unreviewed. Only the no-backend cases fail open with `SAFETY_WARNING` (no provider key and no local lane; a remote key that does not cover the light model's provider, with no local lane; a local runtime chosen as FALLBACK that fails or is rate-limited), so a misconfigured runtime does not hard-block every unknown tool; an explicitly primary local runtime (`USE_LOCAL_LIGHT`) that fails is a real `SAFETY_VIOLATION`.
Registry dispatch passes its prepared binding into `start_service`; a direct handler consults the same Supervisor only when no prepared registry binding was supplied. Each admitted process effect executes once; post-execution settings/repository observations annotate the original result and never repeat or roll back the operation.
Outside Cyber Pro, read-only shell git is allowed everywhere; mutating shell git only when its resolved target lies outside the Ouroboros system repository and runtime data drives (`git_shell_policy.py`; network-disabled tasks still fence network git). Acting `self_worktree` children remain read-only because patch capture requires an unmoved HEAD. `git init`/`git clone` are judged by destination, relative destinations and path-valued retargeting flags included; `resolve_shell_cwd` supplies the same canonical cwd to target checks and execution, and command text is not an independent runtime-file permission classifier. The generic Tool API VCS family (`vcs_status`, `vcs_diff`, `vcs_pull_ff`, `vcs_restore`, `vcs_revert`) defaults to `root=active_workspace` and accepts explicit `root=system_repo`; protected Ouroboros path names constrain generic restore/revert only on the explicit system target, so a project's own `BIBLE.md` or `contracts/` remains ordinary project content, while `preflight_review`, `commit_reviewed`/`vcs_commit_reviewed`, `vcs_rollback` and promotion remain system-repository lifecycles. `vcs_diff` without refs retains unstaged/staged behavior; `base` compares a resolved local tree to the worktree or index, while `base` plus `head` compares two local trees. Results name requested refs, exact tree IDs and comparison kind, without fetching or implying a merge-base comparison.
Outside Cyber Pro, a task contract may declare `resource_policy.protected_artifacts[]` as execute-only black boxes (`protected_artifacts.py`): registry guards allow the declared execution but refuse reads, copies, hashes, static inspection and trace/debug wrappers over those paths. Light-mode cognitive writes are redirected to `update_identity`, `update_scratchpad` or `knowledge_write` instead of raw memory-file edits; a corrected cognitive redirect is advisory, while an ignored user-file root correction remains a blocking deliverable failure.
### Review delivery
`review_execution_projection.review_actor_progress_text` says in one sentence per row who the reviewer is (the frozen route display name), what it was asked for at start, and at settlement whether it answered, how it ran (the observed harness, model and account, or that this was not reported), or why it did not (the engine's reported sentence, quoted; a spent window's reset instant); a `$0` refusal reads as not sent. Requested effort/model is not provider-observed truth; duplicate model slots keep distinct identities in the record. No global last-execution record supplies task progress authority; no code, id or custody state reaches the row.
Review delivery has two closed route kinds in `review_execution.py`: `api_chat` and `agent_session`; vendor and harness names are route targets, not new kinds. A slot is bound to one immutable route before its first send and never falls back to another transport — a route that cannot deliver raises on its own slot (`ReviewRouteUnavailable`). The API executor memoizes the assembled review messages, so its durable prompt record and its at-most-two physical sends share the same bytes; one logical interaction may spend one bounded second send on transport or empty-output recovery, never while a dispatched outcome is unknown (task acceptance may also spend it on malformed-format repair). That two-send rail belongs to the PACKET api row: a native tool-round episode carries no send count — its bounds are the ones below — and neither retrieving delivery is sent a format-repair resend, because their executors canonicalize their own answers. The hosted-agent executor starts one read-only delegated session through the shared Claudexor nanny (route-owned instructions and retrieval pointers, its own tools, no assembled API pack); extraction canonicalizes the collected transcript and never launches a second session, and custody, cancellation, full-artifact recovery and settlement stay on the delegated transport contract (§6 Delegated subagents). Mutating external coding work uses the delegated subagent path. The Claude Agent SDK gateway is retired; two residues stay on purpose: `reviewer_slot_config.py` migrates legacy Claude-SDK advisory targets at parse time (`_migrate_sdk_advisory_target`; an unmapped target is typed `legacy_claude_sdk_target_unmapped`), and the legacy filename `tools/claude_advisory_review.py` hosts the live generic advisory gate.
A session reads with its harness's own tools. `review_session_reads.session_read_facts` derives diagnostic `harness_observed` ranges from completed calls in `attempts/*/events.jsonl`, using the same manifest and range vocabulary as native reads with weaker provenance: an inferred shell range does not attest the exact bytes the model saw. Simple standalone `cat`/`nl`, `sed -n`, `head`/`tail` and bounded file reads have modeled extents; compound commands, pipelines, redirections and searches do not. File reads without enough range information, including ordinary Cursor `read`/`view_file`, stay unobserved, as do missing or unparsable journals. These limits never invalidate the answer, remove a response from quorum or trigger another review. Changed candidate bytes remain a source-gap diagnostic; publication binding is independently checked by the commit gate.
#### Native tool-round episode
`review_native_episode.py` (`NativeToolRoundReviewExecutor`) delivers every retrieving API row: configured-subagent reviewers on any surface, plus every scope, deep-review and advisory API row. Scope declares `ReviewSlot.native_retrieval_override` through `review_substrate.scope_delivery_rows`; deep review and advisory bind the episode directly. One episode is one logical review attempt: `chat(tools=…)` rounds against a fresh instance-local inspection registry (`read_file`, `list_files`, `search_code`, `query_code`, `vcs_status`, `vcs_diff`, `compact_context`; the `local_readonly_subagent` constraint, network off) until the reviewer answers. There is no round cap (P13; the retired round-cap key: §11.4). The bounds are four. The SEND bound `review_native_transcript_bound` is min(owner ceiling `OUROBOROS_REVIEW_NATIVE_MAX_TRANSCRIPT_CHARS`, the reviewer window's density-calibrated capacity in chars) measured as the wire size of the next send (serialized messages plus tool schemas, recomputed after every append) and enforced before every provider call. The owner `deadline_at` (passed by commit triad, scope, advisory and task acceptance) binds together with the slot's independent logical window: each send's transport timeout is clamped to the window remainder and a spent window refuses before dispatch. Each inspection call adds the loop's per-tool timeout, narrowed by the inherited dispatch deadline; past it the call is abandoned as `native_tool_abandoned` on an `error` receipt. The paid ledger prices every physical send as its own attempt. Mandatory reading never raises the send bound: a surface's `policy["native_mandatory_read_chars"]` declares the whole required surface, which the reviewer reads through successive working views (`compact_context` authors a new view; the data plane is the host's empty scratch unless the surface opts into `policy["native_data_root"]`, and only the scratch is removed at episode end); when one view cannot hold the declaration the shortfall is disclosed typed as `native_mandatory_read_disclosure` = `native_multiple_windows_required` beside the declared chars (`native_mandatory_read_bound` is a separate advisory estimate — never a floor, a grant or a refusal). A surface may also declare `policy["native_required_sources"]`; exact host-observed intervals and byte-identical inline delivery fold into `native_read_coverage`. Missing or unobserved reading remains diagnostic evidence beside the verdict, never an episode stop, lost quorum seat, commit block or automatic paid repeat (§6 Scope review by retrieval).
Typed ends: a bound that leaves no room below the first send is a pre-send refusal (`native_bound_below_first_send`), as is a registry without projectable inspection schemas (`native_inspection_unavailable`); both are `not_dispatched` ($0) in review custody and leave the receipt keys (resolved model/provider) empty, so the public execution wire never mints a native run that never sent. The host posts one `[EPISODE_BUDGET]` landing notice per working view at `NATIVE_LANDING_FRACTION` (80%) of the bound, and every tool result is clamped to the room left below it so no single read can jump past the bound; exhaustion is a typed refusal (`native_transcript_cap_exceeded`) for verdict shapes and a disclosed `native_incomplete` product for the report shape. A round carrying tool calls but no well-formed one and no prose is the malformed end `native_round_without_progress` (a progress floor, never a ceiling); an empty answer stays the episode's honest end on the empty-response rail. An episode that ran paid rounds is dispatched (settled) in custody whatever exception ended it — deadline, transport or the paid ledger (`budget_exhausted`). Every end — delivered, refused or errored — leaves one `review_native_episode` custody row and its facts on the actor usage (`failure_custody` for the error actor). The default output contract of a retrieving row follows the surface's output shape (`triad_review.default_output_contract`) unless the surface hands over its own `policy["output_contract"]`.
Native read evidence rides the actor usage as `native_tool_receipts` (`outcome` = executed/refused/error/withheld, `host_file_read_attestation: host_observed`); an executed `read_file` receipt carries `start_line`, `end_line`, `total_lines`, `eof`, `opened_path` and `opened_root` from the reader's own stamp (`ctx.last_read_view` in `tools/core_file_tools.py`; `tests/test_native_tool_round_executor.py` pins the writers), so a bounded read is distinguished from full coverage without trusting the model's path spelling — a missing stamp proves no extent. Every native end publishes `native_rounds`, `native_tool_calls`, `native_transcript_chars` (last physical send), `native_transcript_bound`, `native_transcript_refused_chars` when a next send did not fit, `native_landing_notified` (posted) versus `native_landing_sent` (physically sent), `native_end_reason`, `native_custody_row`, `native_view_changes` and `native_history_source`; a non-delivering or incomplete end after any assistant round also keeps `native_terminal_round`, a bounded, redacted, structurally valid JSON copy of the envelope and tool results that ended the episode, because receipts alone cannot reconstruct it. These facts survive on failed actors through `failure_custody`, deadline, ledger and transport ends included.
#### Late completion and typed refusals
Late review completion is retained in the operation-addressed prompt/response CAS (`review_substrate.py`, `review_custody.py`): the response manifest stamps the complete producer outcome with the original task/root, attempt, slot/route, operation, subject, contract and roster/epoch binding, so exact reconciliation verifies that source and runs the existing surface reducer without a second paid attempt; pending, unreadable, mismatched and unknown outcomes keep custody and their full source references. Plan waves carry their dispatched set, health epoch and operation ids through the same path; free replay spends no new cycle. Skill Review reserves every chunk digest/retry key and slot operation id in `review_job.review_wave` before paid dispatch and keeps that binding in terminal history; the next authorized call may record `review_resume_of` only for the immediately preceding unsuperseded wave with identical task/root/attempt, skill/state roots, group, content, contract, rebuttal and chunks, and then reaggregates complete CAS at the original wave id without another paid stamp, even at the cycle cap. Missing or partial producers stay pending, unstarted chunks acquire no verdict authority from their reserved ids, and a new task/root, changed material/contract/rebuttal, superseding lifecycle or explicit cancellation cannot inherit the old wave. Late completion alone never revives a job or enables a skill: review-state persistence, grants/dependencies and enablement keep their lifecycle guards.
Terminal task delivery and later inspection derive a bounded host notice from the delegated-custody audit after cleanup — confirmed terminal cancellations, live runs, unresolved invocation ids and undisposed patches stay distinct — beside unchanged model text and continuation narrative, so a later cancellation never rewrites historical authorship or the answer hash. A valid acceptance PASS with partial, missing or rejected criteria remains a parseable contributing verdict; only the clean/applied host decision can authorize objective completion.
Malformed structured slot configuration is refused at save and is a typed loud review-time failure on every surface — commit triad, scope, advisory, plan, skill review, deep self-review (`deep_review_slot()` raises; `run_deep_self_review` returns the typed `deep_self_review_unavailable` result instead of a report) and task acceptance (a typed DEGRADED panel, `reviewer_slot_config_invalid`); no surface silently chooses the opposite route or a default panel. Technical failure and commit permission are separate facts: producers keep the failure origin (context, delivery, format or window authority) beside the original status, source and findings, the shared physical exception projection keeps budget/admission and authority refusals distinct, and what each enforcement mode permits after such a failure is the protocol's rule (DEVELOPMENT "Review & Commit Protocol"); neither path changes quorum, and custody, Stop/deadline, ownership and candidate binding stay independent. Contributor slot evidence retains configured-subagent identity while its stored wire uses the mutually exclusive reference-or-route form, preserving native retrieval and its API budget admission.
### Usage ledger substrate vs. accounting policy
The startup seal audit selects its bounded manifest set before reading live attempt rows; missing live ids share one lazily validated archive-id set for the whole audit (an UNKNOWN result included) while forward checks continue independently, and the next audit refreshes that observation through the unchanged archive reader and its integrity/cache checks — no per-manifest archive traversal, no change to monetary authority or archive retention.
`usage_ledger.py` owns the durable append-only physical-attempt ledger (cross-process locking, sequence and transition validation, append+fsync, replay, loud tail quarantine); `usage_accounting.py` is the one-way policy layer above it (pricing, reservations, settlement, scopes, budget fences, imports, projections, admission), and the substrate never imports policy. The boundary means a pricing or budget-policy change cannot redefine valid ledger storage, and a locking or repair change cannot silently change what an attempt costs; compatibility events, state mirrors, task fields and UI projections may carry attempt ids and derived totals but never become a second charge source. `_usage_response.py` is the one NORMALIZER of a provider's usage block for accounting (its importers are `usage_accounting.py` and `loop_llm_call.py`), where a zero-usage body error settles at a confirmed $0 to release the reservation, so a provider storm cannot manufacture phantom budget exhaustion. Not the only READER of that block: every provider adapter reads the raw `usage` dict for its own response envelope, so a "consolidation" there would centralize a read that was never centralized.
### Caller-owned subscription model calls
`claudexor::<source>=<model>` selects a model operation in the Ouroboros-owned engine, not an Agent Run: `LLMClient` dispatches sync and async calls through `llm_claudexor.py` before the OpenAI-compatible filtering/retries, and Ouroboros keeps its SYSTEM, BIBLE, canonical messages, tool selection and execution. The raw-model adapter is Codex; connected Claude/Cursor and other harnesses keep their Agent capabilities, and direct API keys keep their existing routes. Every main-loop execution of one install sends the same Codex cache key (`llm_claudexor.cache_key_for_model`: per data root and model, projected to `prompt_cache_key` and the `session_id` header), so a new task, child or consciousness cycle is served the governance prefix its predecessors already cached; per-execution turn states ride `nativeContinuation` unchanged under that shared session.
`ClaudexorGateway` uses purpose-bound byte uploads, one idempotent operation ID, status/result reads, cancellation and explicit result acknowledgement. Claudexor performs at most one generation per operation: a lost local HTTP reply rejoins the same operation, and an unknown provider outcome never authorizes another generation. It shares its managed profile with Agent sessions — the adapter reads current access credentials while the official CLI keeps refresh ownership — and no shared external engine, relay process, copied operator auth store or public OpenAI-compatible endpoint. Metadata helpers use `read_owned_gateway` (owned-only discovery and handshake; they never install or start the engine); actual model calls and explicit lifecycle actions use `ensure_owned_gateway`.
Failure evidence: before creating an operation the host discovers the catalog's `captureFailureEvidence=true` query and freezes that choice for the invocation, same-key rejoin after control loss included (older catalogs keep the legacy request; the serving runtime pin remains the independent adoption boundary). A supporting engine puts every received failed-response byte and processing exception — incomplete responses and HTTP refusal bodies too — in the private result, which the observability CAS retains before ACK; usage and events carry only compact problem context, and a received prefix never claims the missing provider tail. A coherent provider terminal whose message cannot be normalized keeps its outcome, route and usage: the operation fails `response_rejected`, the host settles the reported usage and raises the non-repeatable `stream_rejected` class — no unknown-outcome recovery, no healthy account marked failed. `ClaudexorModelError.display_message` adds sanitized typed `vendorCode`/`parameter` details to the error event and terminal preview only after classification; unknown outcomes expose no provider details.
Money and bytes: physical attempts stay in the usage ledger, where exact zero cash, known charges, estimates and unknown cash remain distinct — a subscription is not evidence of zero cost. Claudexor removes terminal request bytes and acknowledged result bytes, an unacknowledged ready result expires after 30 days, and compact operation receipts survive body deletion, so an expired read cannot start a new generation. Model-purpose resources are not valid Agent attachments. Private continuation envelopes stay in canonical history but out of public/summarizer projections; their reuse is bound to the actual source/model/profile/account identity.
Image capability comes from the selected role/account's model catalog, not the global model-id overlay; main send preparation, browser image attachment, captions and registered VLM tools share that reader. Missing metadata stays unknown — image input is retained for the real call rather than silently omitted on a cold engine — while a confirmed text-only model keeps the caption/capability-gap behavior, and text-only messages and image-off mode never read the image catalog.
The send copy omits foreign provider envelope metadata such as `stop_reason` (canonical history retains it), preserves a nonempty `refusal` as assistant text even when the provider supplied no ordinary content, and normalizes all-text tool results so moving the message-side cache boundary never rewrites bytes already sent (non-text blocks stay intact). For the same reason the transcript is append-only between the sends of one execution: compaction is the one sanctioned rewrite, stamped at its seams, and `ouroboros/transcript_prefix.py` records every other break (a context-fit reprojection after a real overflow included) as a `prompt_prefix_break` checkpoint instead of blocking the send. Only the exact `model_request_invalid` create refusal proves validation failed before command admission; a generic HTTP error or failed status read cannot prove non-dispatch.
#### The live turn slot
Beside assistant-level history the engine keeps ONE live turn per model client session, and its opaque token is transport, not content: it belongs to the caller that is running, not to any stored message. So the caller owns a single mutable slot (`llm_claudexor.ModelTurnState`) that rides the ordinary `LLMClient.chat` parameters to the engine boundary, the only writer — a dispatched durable result replaces the value, a released `invalid_continuation` repair clears it together with message envelopes; other non-dispatched, unknown or legacy-shaped outcomes preserve it. Holding state never licenses another generation, and silence about a turn never ends one. WHY a slot rather than the transcript: the live turn survives compaction that drops the messages and outlives a body that fails after its headers, so the last stored assistant envelope would rebuild the wrong lifetime. One `run_llm_loop` invocation is one turn (rounds, owner steering, acceptance follow-up, forced finalization, quota waits and reprepares continue it; a next loop and a cold restart start empty), and a consciousness wake-up is one such ordinary turn keyed by its own task id. A dispatch that leaves this transport for a direct API or local route ends the turn at the caller and never revives it on return; the engine alone compares route identity. Opting in requires a serving engine at `config.CLAUDEXOR_MODEL_TURN_STATE_MIN_VERSION`, read from the last SUCCESSFUL handshake — a failed probe never un-proves it, so concurrent status polling of that singleton cannot change the shape a running caller sends — and an older or not-yet-observed engine keeps the legacy request shape instead of risking a schema refusal. The token is never usage, an event, a progress note or a task card; mechanism in the `llm_claudexor.py` docstring.
#### Roles, accounts and context
Model roles own account pins and context assertions through `OUROBOROS_MODEL_ACCOUNTS` and `OUROBOROS_MODEL_CONTEXT_WINDOWS`; ordered fallback entries keep their original ordinal, reviewers and configured actors reuse their route credential fields, and identical Main/Light or reviewer model strings never imply identical roles or accounts. Auto uses the largest advertised window of the exact account route, not CLI compaction thresholds; a manual value is a sizing assertion, not a provider unlock or reviewer authority; unknown capacity stays unknown, and input size, response reservation and capacity are separate. Each operation records its submitted options beside the engine's applied options (an absent applied report stays explicitly unknown), and the first changed reasoning effort of each model in a task produces one typed owner-visible notice. An account change rebinds preparation before another physical send; ordinary sends, prospective wrap-up payloads and forced final replies share one acting-role/account binding, and prospective subscription accounting uses the dispatch serializer, not an OpenAI-compatible approximation. The final header identifies the responding model/account; a paid review that ends incomplete retains its full report and custody with that classification, never switches to another delivery or inherits the previous route's window.
Host-authored sampling defaults stay separate from explicit parameters: review requests carry `default_temperature` as a hint each effective API route resolves at dispatch, raw model operations use their provider default, and an explicit `temperature` (zero included) is preserved and may be refused. Routes without structured output use the existing safety text-JSON parser, repair and refusal path — no new safety bypass. Ordinary confirmed provider errors keep each helper's previous handling, while quota, owner interruption and unknown physical custody propagate without becoming a helper result or verdict.
#### Quota and auth waits
`model_wait.py` holds confirmed quota/auth waits inside the live call. Typed owner actions use the existing mailbox and the `/api/decisions` family `model_wait:<task>:<wait>`; task-result rows are projections, not restartable stack checkpoints or a second attempt ledger, and metadata polling makes no generation. Auto may prefer the last successful same-route account unless it got a status-null or typed per-subject refusal this execution. An HTTP vendor quota refusal of a named account is re-asked with the same payload, no preference; a pre-dispatch verdict never is; only its `poolCause` proves the pool spent; an account chosen twice stops rotation, claiming nothing. Pin never rotates. A route may also answer with ANOTHER model than the one requested, with no error and nothing in the account's quota: the engine states that as a typed fact on its completed result, and the transport never accepts such an answer as the round. It is retained and acknowledged like any other answer, so the horizon it would have run under stays reconstructible, but it adopts no turn state, runs no tool call, and is replaced by a NEW operation that names no account — the engine, which deprioritises that account for a while, alone chooses where every redo lands. `OUROBOROS_SERVED_MODEL_REDOS` bounds those redos; a pinned account names one account, an already admitted candidate is bound to exact bytes, and an actor with no send left on a bounded budget would otherwise read that rail's own budget error instead of this refusal, and a spent owner window buys no further send, so none of those is asked again. An outcome nobody knows is never a substitution: the transport raises it before the rail sees a result, so the existing unknown-custody fence, not this budget, governs what happens next. Each discarded generation's retention and acknowledgement state travels with its durable row, so a failed one is visible rather than assumed. A discarded generation teaches the density witness nothing, because a witness belongs to the model that produced it. Once no redo is available the typed `model_substituted` refusal reaches the caller, which is an ordinary cross-model fallback trigger for a root and a typed failure for an exact-route child; it never cools the requested model, because one substituting account says nothing about the model. The first substitution of each model in a task also becomes one timeline row of that task's card. The rule covers this transport only: an Agent-session delegation is a separate engine capability with its own model evidence. Fallback keeps one attempt per model and a 120-second cooldown; on quota each configured route runs before the owner wait, opened from the retained refusal (no generation first); an intermediate fallback's own refusal remains the wait/retry route if a later fallback fails differently; nothing sleeps to a reset; inline Presence never waits (typed refusal; safe retry proof: §12). A temporary switch affects only the wait; Settings persistence, acceptance and task application remain distinct. Fallback Local remains a group-wide setting: a one-role persistence request may change model/account only when Local is unchanged, a different Local choice is task-only, and permanent group changes remain in Models. The current task-attempt identity selects actionable wait rows; a failed mailbox delivery retries the same accepted request without authoring a new one; terminal mailbox cleanup waits for post-task synthesis and pending accepted attachment promotion to settle.
`ouroboros/gateway/task_model_wait.py` owns model-wait decision effects and the bounded settings-writer receipts; `task_decision.answer_decision` returns `(status, payload)` to both the Web and authenticated Host Service transports (they supply the installation root — no synthetic HTTP request or parallel owner registry); `supervisor/task_model_wait.py` projects worker events and quota clocks through the existing supervisor, and `ouroboros/gateway/decision_contracts.py` owns the decision-family vocabulary. Quota pauses use union duration across simultaneous waits and do not consume internal execution time, while explicit calendar deadlines remain fixed; the worker and project-write ownership stay held (a one-worker queue waits), completed tools and reviewer results remain on the same live stack, closing a browser does not stop the wait, and stopping the Ouroboros process ends this continuation guarantee. Image tools retain the tracked VLM child: `vision_process.py` carries typed errors, physical capture, cancellation and operation checkpoints across that process boundary, and a lost child result remains unknown, never free.
A consciousness wake-up is an ordinary direct turn (`handle_wake_direct`): the same direct-turn registration, model wait, Stop and bounded interactive transport-death repeats as an owner's turn, an ordinary task result, and no background owner, locked memory, pause or stop event of its own. The alarm clock only decides when the next turn starts.
#### Streams and transport waits
Main remote completions stream (`llm_stream.py`), assembled inside the physical send closure before accounting settles: a complete stream yields the same normalized response shape as JSON; a partial stream never yields a usable tool call or answer, and its exact wire bytes and partial assembly stay in private CAS through the `physical_stream` manifest and attempt id. The assembler is strict about completeness, tolerant about form: on the Chat path identity scalars and metadata keep their first value, index gaps are forgiven and every forgiven fact is disclosed in the stream receipt (`anomalies`); the native Messages assembler rejects post-terminal shape (non-contiguous blocks, an unsigned thinking block, an incomplete tool block). A complete-but-unusable stream settles with its usage as a provider error (`stream_rejected`); unknown-outcome continuation is reserved for a stream that never reached its terminal frame (a lost socket, a malformed mid-stream native frame), and a complete response with absent final usage still has unknown money. Comments/pings do not define cognitive deadlines; every recovery candidate checks inherited calendar and quota-adjusted execution bounds before reservation, after preparation and at dispatch, and HTTP phase bounds stay distinct from the overall logical wait.
An ordinary managed task whose provider outcome becomes unknown stays in the transport-wait episode (`loop_transport.py`); the old attempt and its unreported cost remain unknown. After non-generating upstream observation, a user-role `[SYSTEM NOTICE]` supplies explicit recovery input for one new physical attempt under the existing budget, cancellation, owner deadline and absolute ceiling. An attempt that fails again, is unknown after dispatch or released before it — and a free redial that crosses dispatch and dies unknown — returns to the same episode: the backoff keeps growing (4→60 s), one redial is counted per wait iteration, an unknown repeat re-arms the probe's freshness bound at that latest failure and refreshes custody, one `continued` row names the transition (`continuation_outcome_unknown`, `continuation_transport_unavailable`, `redial_outcome_unknown`), and no cap on continuations — budget, deadline and Stop are the only stops. Finished tools are retained. Configured-session supervisors prefer their live-leaf hold (`delegate_hold.py`), then use this same recovery for their own model; reading a result or disposing a patch is not a condition for cognition. Direct turns keep their separate contract; manual Restart/Panic gains no resume authority. Recovery is proved by provider provenance, never by a local handshake: direct remote endpoints are observed through HEAD under the same no-proxy policy, reusing the ordinary connection allowance for every socket phase narrowed by the owner remainder and holding no cognitive in-flight lease; subscription metadata qualifies only when the catalog reports generic `provenance="provider_http"` with an original `observedAt` after wait entry for the selected source/model and effective profile/account fingerprint — cached reuse never advances that timestamp, and a static or pre-outage catalog cannot prove recovery. Capability discovery remains authoritative, not a new core provider table. Loss of the Claudexor control connection keeps reading the same accepted model operation across endpoint rediscovery, without creating another operation.
### Delegated subagents (Claudexor transport + the nanny)
Ordinary delegation requests no extra engine review panel; new ordinary runs on Claudexor 3.9.8+ default to no panel — a version boundary of the behavior, not a second release pin. The start receipt's `engine_version` is the handshaken serving version, distinct from the release pin; engine review outcome, execution success, the parent's integration decision and Ouroboros review gates remain separate facts. Delegated snapshot capture stays relative to its recorded baseline: committed bytes can be captured with a disclosed `head_moved`, the instruction still forbids committing, and the distinct `self_worktree` unchanged-HEAD check is preserved.
Children coordinate through `tree_note` and `tree_read`; read-only children can read/list the project-scoped knowledge store without writing its index; only the parent may use `override_delegation_constraint`, and a `review_requested` note carries an exact evidence reference/hash and wakes the parent without starting a paid cycle. Read-only and acting children alike hold the descendant-scoped `forward_to_worker`, `peek_task`, `cancel_task` and `discard_child_result` controls, and recursive delegation never widens filesystem, budget, depth, deadline, commit or owner authority. `forward_to_worker` also reaches the child's parent or a sibling (same parent and root) as `peer_task` provenance, its prefix naming the stamped `relation` (`sibling`|`parent`): context-only, relay refused, 8000-char bound. `delegation_budget` governs descendants (`may_delegate`, `may_fan_out`, additive depth provenance; a free-form intent note is never authority); persisted admission facts outrank later Settings changes, and a lower permitted depth is reported `capability_reduced`, never a silent flat tree.
**Registry.** `configured_subagents.py` owns `OUROBOROS_SUBAGENTS`: strict `{enabled, items}`, at most ten rows with hidden `subagent_id`, route-handle name (§3 Available subagents), any-language `recommended_use`, normalized API/session route, optional effort, account pin, access and row `enabled`. Session access defaults to `full` or the owner's `workspace_write`; snapshots lacking access retain workspace_write. Legacy env inputs fail closed. Row `enabled` defaults true, rejects nonbooleans and serializes only false, preserving fingerprints/receipts. Off rows retain configuration but leave catalog, alternatives, legacy matching, local autostart and reviewer resolution. Explicit selection returns `subagent_disabled`, distinct from global `subagents_disabled` and live unavailability. Tools take handle or stored id; roster edits refuse twins; records name their own engine. Captured snapshots never consult live settings. Descriptions guide the LLM, never host ranking/keyword routing; exact selection starts that route or refuses without substitution.
**Scheduling.** `schedule_subagent` requires `subagent_id`, a focused `objective` and `expected_output`; the remaining public fields describe child-local context, constraints, capability needs, write surface, narrower deadline, delegation budget and acceptance claims. `access=inherit` (default or omitted) preserves the configured session access; `readonly|workspace_write` can only lower it. API-model rows ignore this field with a result disclosure; `write_surface` stays the read/write authority. There is no model-visible lane/executor axis and no public `effort` override: the selected row is the complete execution choice, lineage/bounds/route/budget are host-derived, and omission never inherits the parent's acceptance claims. `subagent_runtime.select_subagent_snapshot` copies an immutable snapshot of the exact enabled row into the child task: an `api_model` row becomes an ordinary recursive API child on that exact model/effort, an `agent_session` row an ordinary recursive Ouroboros nanny on that exact external route. Burst cash (each sibling launched before the first sibling's first response pays its own prefix write on cache-write-priced routes) is disclosed in the tool description as an affordance, never scheduled by the host; so is an exchange of addressed turns, stated in `objective`/`constraints` (never a contract field or host section; no stage, topology or participant count fixed). The old lane/executor resolver serves only old durable records: for historical lane envelopes `schedule_subagent` reports the requested lane only (effective facts remain on the dispatched child record), and a task carrying `configured_subagent` goes straight to `subagent_runtime`, so legacy policy cannot reinterpret an active selection.
**Waiting on children.** `wait_task`/`get_task_result` return the full single-child handoff (verification receipts red/masked-first, exact omitted count); `wait_tasks` stays batch-compact: `task_id, status, child_result_sha256, outcome_axes, result, terminal_host_notice when present, trace_summary, capability_delta when the child has something to disclose, duplicate_of`, plus the nullable cost-finality pair `accounted_upper_bound_usd`/`cost_final` and, when the child's envelope carries one, `execution_evidence` (§11.1); the retired `cost_usd` spelling is tolerated only when reading stored rows, never emitted. Both use `task_status.SETTLED_STATUSES`, and a pending cancellation is the typed `cancel_state: "pending"` projection, never completion. A batch wait that expires with children still running returns the typed `wait_expired_with_live_children` block (live child ids, requested window, clamp ceiling): facts only, no advisory text, no host floor on the next window (how long to wait is the mind's call, BIBLE P13); an id this tree never minted stays `unknown_task_id`, never a live child. An optional `known_result_sha256` / `known_result_sha256_by_task` compares the join-ledger semantic result identity (`join_ledger.py`; `CHILD_RESULT_STALE`): an exact match returns `result_unchanged` without the repeated text while status, costs, outcome and custody facts stay; no persistent seen-state is inferred. `await_messages(timeout_sec)` holds the worker slot until an unread mailbox entry or its bounded window, delivers nothing and takes no lease: the in-flight tool lease spares the idle rail and its close stamps progress.
**What a delegated run costs.** Claudexor reports the amount in `summary.spendUsd` and its exactness in `summary.spendEstimated`; `delegate_custody.disclosed_spend` is the single reader of the pair, so the ledger row and the payload the nanny relays cannot tell different stories. Runs ask `authPreference: subscription` explicitly, because the engine default falls back to a paid key invisibly. Four cases:
| disclosure | ledger | finality |
|---|---|---|
| disclosed settled zero | `0.0` (`subscription_session` row) | `cost_final=true` — the proven free session |
| disclosed settled charge | the amount | final |
| estimated amount | the amount | `cost_final=false` — an estimated zero is not a proven free session |
| undisclosed | `cost_usd: null`, increments `unknown_unmetered` | drops `cost_final` for the projection |
An undisclosed spend contributes `0.0` to `accounted_usd` — inventing a conservative bound would fabricate a number the harness never gave (BIBLE P1) — so a `TOTAL_BUDGET` fence cannot stop spend it was never told about; the honest consequence is loss of finality, not a guessed charge.
**What a delegated run READ.** Harnesses count input tokens with incompatible conventions, so `settle_run` carries the engine's own normalized split, `summary.inputTokenUsage`, onto the `subscription_session` row as the optional `input_token_usage` object: `total_tokens`, `cache_read_tokens`, `cache_write_tokens`, each a nonnegative integer or `null` for unknown. `record_subscription_session` is the one validator and writer: all three keys are required together, and a partial, extra-keyed, negative or fractional object is unknown as a WHOLE, not repaired field by field, because a repaired counter reads exactly like a measured one. An engine that reports nothing leaves the row unchanged, and the object stays OUT of the idempotent row identity, so an engine that starts reporting never rewrites or duplicates a session already settled without it; the legacy `prompt_tokens`/`cached_tokens` axes keep their own meanings, compaction never folds these idempotency-bearing rows, and direct physical attempts take no copy.
**The nanny model.** An `agent_session` subagent is an ordinary recursive task-tree child acting as a **nanny** supervising at most one active bounded external leaf. The task node keeps lineage, authority, deadline, budget, acceptance, cancellation and descendants; the harness process stays a non-recursive tool leaf: session rows never flatten the task tree into harness processes. The nanny is the host: verification receipts stay host-authored, and harness output is a claim to check, never proof. A nanny-to-nanny chain through `schedule_subagent` is the host-attested realization of a nested subscription swarm, each level one metered supervisor task plus one free harness run; `schedule_subagent.requested_depth` is the typed ABSOLUTE request counted from the root, telemetry that never narrows the configured caps, while the legacy `depth_remaining` envelope keeps its narrowing semantics (both disclosed on the contract), and the root's `swarm_efficiency.depth` block reports requested, permitted and achieved. A nanny terminal names BOTH routes as separate facts: the host's own `model_execution` carries its model, provider and last typed error (`last_llm_error_kind`), each `terminal_runs` row the leaf's model, profile and selected actor, so a nanny that died on its own lane is never read as a verdict about the leaf, and a leaf half whose reconciliation was never persisted is omitted, not guessed. Harness-agnostic by construction: the row holds an opaque Claudexor target, Ouroboros asks for an access profile derived from task authority and lets Claudexor choose the mechanism; no harness-name branch selects a capability or fallback in core dispatch, and login-wire asymmetries stay presentation adapters in `gateway/claudexor_accounts.py`.
**Transport.** `gateways/claudexor.py` is pure transport (descriptor read, `/v2` handshake, the `config.CLAUDEXOR_MIN_VERSION` 3.2.0 floor). The daemon bearer token grants the entire `/v2` surface, so it never leaves this module, and the HTTP client runs `trust_env=False` so an ambient proxy cannot intercept the loopback control plane. Production starts obtain a handshaken owned gateway from `claudexor_daemon.ensure_owned_gateway` — exact reviewed engine/Node pins (`claudexor_runtime.py`; the reviewed pin IS the next-spawn selection — no mutable `current` pointer, no background updater — and `OUROBOROS_CLAUDEXOR_BIN` is the explicit operator override), never a PATH install. Keeping lifecycle above transport keeps account status side-effect-free and harness mechanisms out of Ouroboros (`claudexor_daemon.py`/`claudexor_runtime.py` docstrings; stop path, spawn latch and typed start failures: §9). Native child processes are contained by env token (`process_containment.py`, `OURO_PROC_CONTAINER_*`), because a surviving descendant can become invisible to parent-child traversal once its controller exits; an alive-or-undeterminable member is an honest hard-block answer, never a kill guarantee.
**Custody is durable, because the run is not ours to kill.** A delegated run lives inside the daemon and survives our worker, so the AUTHORITY is the durable `delegate_run_*` custody rows (`delegate_run_started` and friends) on the canonical/budget root (`ouroboros/delegate_custody.py`), read by every replay-class consumer through `custody_rows` — the process-local memo in `delegate_custody_memo.py` (the `_usage_rows_memo` shape: an ordered `(st_dev, st_ino, consumed, st_mtime_ns)` fingerprint of the folded chain prefix, only appended bytes folded later, a refold on any doubt, a bypass and never a cache while the chain is unreadable; the rows stay the authority, each process pays one cold fold, a durable compact projection is the next step); one compact incident projection, `<drive_root>/logs/containment_faults.jsonl`, exists because the event log grows without bound and a tail-bounded scan can bury an unresolved fault. A lookup answers OWNED, FOREIGN or UNKNOWN — collapsing UNKNOWN into "not yours" made a restarted owner indistinguishable from an intruder. An ABSENT custody log is a positively established clean state; an EXISTING-but-unreadable one audits as `delegated_run_state_unknown:custody_log_unreadable`, never cleanly reconciled. Every INTENDED start mints a fresh per-intention invocation UUID as the wire `Idempotency-Key`; the content hash is only the LOOKUP identity — a content-stable wire key would hand a deliberate re-run the finished old run — and reuse happens only by explicit token (`pending_invocation_id`/`retry_of`, replaying the STORED canonical body under the SAME key). The complete replay envelope lives unredacted in the existing private observability CAS, written before `delegate_run_start_requested`; the event carries `request_ref` and `prompt_chars`, keeping a large reviewer packet out of every custody scan (the memo swaps a legacy inline body for a locator `delegate_pending.request_body` re-reads). `delegate_pending.request_body` resolves legacy inline first, then verifies the CAS reference for both invocation readers. JSON values and the canonical request digest stay unchanged, including the thread fields needed to reproduce its wire projection; a missing/corrupt blob leaves the request unknown while preserving pending custody identity for start blockers, terminal audits and snapshot retention; recovery retains that invocation without POSTing. Known pending review tokens take the same missing-request path, while absent or definitely refused skill-review records keep their existing fresh-retry behavior. Pending scans resolve only surviving invocations, while legacy event/archive bytes remain untouched. `reconcile_orphaned_runs` visits every open run whose owning task left the live set and settles the terminal ones, but it CANCELS only behind a deliberate verdict: a durable owner result that is readable, truly terminal and finished by the task's own decision. A custody row carries its owner's kind, and task association confers no lifecycle authority: a run a review surface registered (`RunCustody.review_owned`, durable `source` under the review substrate) belongs to its panel, bounded by the slot window and its own `maxSeconds`, so the sweep spares it unless the owner task was itself cancelled — a `left_live` row names the panel, and a task that consciously finished under a running acceptance panel keeps its reviewer alive; a pending review invocation is likewise retained, never re-posted by the generic recovery, because the review substrate owns its rejoin. The owner-cancel kill boundary STATES the verdict it is about to write, because it audits custody before that write, so an owner cancellation stops the paid run at the boundary, not at the next sweep. A provider or transport death, a worker crash, a missing or unreadable result all SPARE the run, left live with a durable `left_live` reconciliation row: an undignified nanny death must not kill a healthy paid run, and "unknown" is exactly that case. `maxSeconds` is the damage limitation for a spared orphan, never custody. SETTLED is published before registration retirement, after the ledger row lands; settlement and registration retirement are separate durable duties, a failed retirement stays replayable on `project_owned` for the later sweep, and writing `settled` over a suppressed ledger failure would turn a lock timeout into a permanent leak. A start whose row did not land reports `started_uncustodied`: no supervision, no replacement, until the original run is proven absent or terminal.
**No terminal or cancel claim without a verified receipt.** `delegate_cancel` returns `confirmed` (read back terminal), `requested`, `failed` or `containment_fault_run_may_still_be_live`; the last two hold a durable CRITICAL containment fault until a receipt or settlement clears it — an overpowered run that may still be alive is an incident, not a reassuring string — and the state read decides, so a refused control is never a verdict about the RUN. One `daemon_says_absent` predicate decides everywhere that a 404 is the daemon ANSWERING that the resource is gone (scoped to the daemon that answered), never a failure to find out; custody closes such a run `delegate_run_closed_absent` (unreachable, not settled), inventing no terminal detail, usage or spend. Results are delivered, not severed: `delegate_wait` stages the whole terminal detail atomically under `task_drive/delegated_runs/<run>.json` with a typed `output_delivery` block, and cut fields are renamed `*_preview` so a partial read of head-truncated JSON fails loudly instead of looking like an answer.
**Four nanny verbs** (`tools/delegate.py`): `delegate_start`, `delegate_wait`, `delegate_cancel`, `delegate_answer`. There is no fake `hurry`: the transport supports cancel and answers, not in-place steering. `delegate_shared._fail` adds `ok: false` and `host_code` without replacing the payload; `_owned_run` retains its durable OWNED/FOREIGN/UNKNOWN boundary for wait/cancel/answer. Start selects an exact session `subagent_id` (API ids refused), or recovery-only `retry_of`. `subagent_bootstrap` starts the configured nanny's snapshotted leaf through `delegate_start(prompt="")` before its first model round, without waiting (`configured_session_started`); the model chooses `delegate_wait`. Recovery adoption precedes zero-run/unknown fences because a prior run may be live; a fence-wake outranks terminalization.
A definite refusal needs a typed cause, no custody handle and either producer `definitely_unrun` or `_DEFINITE_UNRUN_REASONS`; it ends UNRUN at $0. Ambiguity wakes the model: uncertain live work cannot become a zero-spend terminal. Unsupported readonly/payload geometry (`directory_execution_unavailable`) and unregistered roots refuse before start with durable start-blocked evidence; the nanny registers first and retires owned registrations at settlement (`delegate_registration_policy`). The payload selector names authority through fresh `ResolvedResourceBinding(skill_payload.write)`, never grants it. Actor replacement/zero-run fences exclude `review_owned` runs: those belong to review custody, while settlement, money, containment and registration still see all runs. New delegation readers join the consumer matrix in `tests/test_custody_owner_kinds.py`.
**Finite leaf continuation** (`delegate_continuation.py`): engine `maxSeconds` cancellation settles as `outcome_reason=wall_clock_exceeded`, replayed as `terminal_reason`. Explicit `continue_from=<run_id>` admits a NEW run/cap/key only for the caller's settled run with that confirmed cause, full output read to EOF, captured patch applied/rejected without ambiguity, and unchanged actor/route/access/mode/isolation/authority target. Stop/Panic, owner deadline, user cancel, failure and unknown endings refuse. `continuation_of` and the host block retain predecessor/cause/disposition and the prior run's recorded access, so a mutating run that captured no patch is described as having written in place (its effects already on the target), never as read-only; the model authors remaining work in `prompt`. No session state transfers; `delegate_recovery.NO_RESUME_CAUSES` stays unchanged.
**Execution evidence.** `delegate_evidence.task_execution_evidence` projects the custody rows read-side; `delegate_start_attempted` counts blocked and uncustodied attempts too, so a refused-but-obedient nanny is never disclosed as nudge-ignoring (`nanny_nudge_recorded`). `applied_access_profiles` is read off SETTLED rows only (empty = no receipt disclosed it, never "no access"); `acceptance_patch_dispositions` is the bounded section over `delegate_run_patch_verdict` rows (cap 20 with the exact omitted count, `unreviewed_delegated_apply` headline) whose absence means NO disposition was recorded, never "reviewed clean"; an unreadable custody log is the typed `evidence_read_failed` marker, never an empty-therefore-clean section.
**Work orders.** The compiler (`subagent_work_order.py`) sends the entire chosen assignment, preserving context and instruction roles without an arbitrary host-size cutoff: direct starts carry the normalized host contract once in `instructions` (`delegate_start_instructions.py`; the separately hashed coordination appendix is absent from the host pre-start) and the chosen assignment separately in `prompt`, and coordination context stays complete. Real native/HTTP limits return their actual failure with the original input and any pending invocation retained. Exact-source readers stay optional capabilities; incomplete source coverage never authorizes a terminal PASS or apply (`delegate_source_coverage.py`). Legacy partial starts keep their exact renderer digest, source-interval validation and stored-body retry; removing partial-start production certifies no incomplete old run and launches no duplicate after ambiguous dispatch.
**Supervision.** `delegate_wait` is model-visible as an event-only sleep, not a caller-sized poll: `delegate_supervision.supervised_wait` renews bounded transport windows in host code at zero LLM calls, and only a meaningful event (settlement, interaction, fault, addressed message, child signal, control, recovery judgment) becomes a coalesced durable wake, replayed across worker interruption until acknowledged; deadline, ceiling, budget and cancellation stay outer bounds. Every receipt and wake carries one host-rendered `coordination_context` (intent, time remaining, root-tree spend, active descendants, remaining paid acceptance capacity) (facts for LLM judgment, never thresholds), observed READ-ONLY, so a metadata-poor task reports `time.state = "not_set"` instead of latching an anchor from a poll. Polling writes nothing of its own beyond the canonical usage-ledger reader's bounded maintenance, identical for every reader: the torn-tail quarantine after a SINGLE crash mid-append (one verbatim row in `state/usage_attempts.quarantine.jsonl`, one `usage_ledger_tail_quarantined` event), the empty `state/` lock directory on a never-initialized root, and owner-aware `usage_attempts.lock` recovery (§1 Platform substrate); an absent ledger is that reader's known-zero; a crash inside the quarantine repair can leave the sink torn (a disclosed residual). A requested future inspection (`checkpoint_after_sec`) wakes once and is consumed by any earlier real event: no cadence, stall classifier or hidden polling. An observation the transport could not complete is a quiet renewal carrying its typed reason, never a refusal that spends a model round: `observation_read_timeout` is our own read bound expiring against a live daemon (quiet, nothing more), `daemon_unreachable` a socket that carried no answer, the only half worth an owner line, delivered once per episode with one recovery line on the supervising task's own progress surface. The beat stays three seconds with no backoff, durable counter or outage latch, and deadline, ceiling, budget and cancellation still cut a long unobserved stretch. On a wake the nanny holds its full tool surface and the parent's captured model/effort. Nanny economics are structurally quiet: only `BASELINE_RESET_TOOLS` (`delegate_start`/`schedule_subagent`) reset the burn baseline, and coordination verbs never buy metered silence (`nanny_pacing.py`).
**A run's question is the nanny's to answer.** Supervision wakes immediately with typed `status="waiting_on_user"` on a NEW `pendingInteractions` entry instead of burning the engine's answer timeout in dead polling; answer keys are echoed verbatim into `delegate_answer` (custody-gated like cancel, relaying the engine's typed outcomes: `subscription_window_exhausted` carries `reset_at`; a transport death or 5xx is `delivery_unknown`), and delivered interaction ids are acknowledged only after transcript injection, so a question neither re-triggers a round nor disappears across recovery; every waiting payload states `continuation: same_session` (an answer resumes THIS session; each turn paid). A question above the nanny's authority escalates to the nearest live ancestor. The codex lane has no mid-run channel: a terminal with `outcome_facts.reason=input_required` (`continuation: new_physical_run`) is answered by a plain new `delegate_start(subagent_id=..., prompt=...)`, never the engine's rerun verb, which would start a run outside this task's custody trail.
**Recovery is exact and cause-specific, not generic task resurrection.** Only a proven non-signal worker crash (or a planned self-restart's typed handoff) lets a successor adopt the exact run or pending invocation, before any LLM call or new start; anything ambiguous returns typed recovery-required, never a duplicate mutator; owner restart, panic, signals, deadlines and cancellation are explicit no-resume causes (`delegate_recovery.py`; `delegate_pending.py` replays a pending invocation under its original key and body). When no physical run exists and none can be started, the model may finish only with a typed zero-run receipt (`verify_and_record`) after custody proves nothing is open; a session actor's terminal is CLEAN only through its own physical leaf or that receipt ("completed direct child ⇒ clean" does not exist), and `physical_leaf_not_started` rides the terminal projection as an incomplete/unknown execution axis. An unknown-provider hold (`delegate_hold.py`) parks the nanny in the supervised wait until the leaf wakes and never resends.
**Configured dispatch is exact** (`subagent_runtime.resolve_configured_actor_dispatch`, resolved at dispatch, the last moment route availability is current): a bad choice returns a typed reason, current alternatives and any reset time with `host_fallback: false`; the host never ranks alternatives, converts session work to API work or picks the first healthy row. The WHY is recorded so no redesign undoes it: selecting an `agent_session` row IS the parent's typed LLM decision that this work executes on the harness, the FLOOR the host hardcodes (truth, money, authorship), while topology, decomposition and supervision judgment remain the model's CEILING (BIBLE P13/P5). At completion `actual_substrate` derives only from custody rows; a typed startup refusal never authorizes work on a different substrate: another session route or API fallback is an explicit LLM-selected start.
**Route health.** `subagents.route_health` is the one health reader for dispatcher, nanny and review slots, and it does not guess admission: the doctor `status` describes only the DEFAULT credential store while real accounts live in the engine's credential-profile pool, so admission is the engine's, whose start POST answers an empty pool with its own typed refusal (`credential_pool_exhausted` + earliest reset) at $0; the gateway types it as a timer-healing window, scheduled on its reset like a spent subscription window, only when `resetsAt` is dated; an undated (structural) pool stays a plain `ClaudexorUnavailable`. `enabled` is honored as `route_disabled` (the owner's switch, not an observation); the other structural refusals are `route_not_in_capability_catalog`, an access-profile mismatch, `engine_rejects_delegated_marker` and positively proven quota exhaustion. A fully-used ratio needs a valid future reset before it can refuse a route; an incomplete reading delegates admission to the engine and never certifies available quota; explicit active cooldowns bind. Review slots inherit "the engine decides", never a silent fallback onto metered API spend: on the `auto` lane a dead daemon keeps the native fallback with its visible marker, while "daemon alive, pool empty" is discovered at the engine and disclosed. The model sees a semi-stable facts-only catalog (`model_visible_subagent_catalog`); invalid, list-disabled or all-rows-off configuration projects nothing; dispatch is the live check.
Historical helper observations use `state/subagent_last_delegation.json`, owned by `subagent_history.py`. Task finalization joins configured API actors to their own attempts, preserving requested pins and observed routes/accounts; fallback success never certifies the failed route, nor does code/test failure condemn a provider. Sparse pre-response failures stay in usage/history/raw events; served-call trace references retain received response-backed calls, including incomplete responses. Foreground/recovered session terminals and typed start failures carry original effort/processing through `RunCustody`/`STARTED`/`START_FAILED` facts, without history-only archive reads; missing options stay unknown, distinct from captured defaults. Pre-invocation bootstrap refusals keep their source; reviewers keep separate receipts. Known occurrence and first observation differ: replay neither refreshes age nor replaces newer dated evidence; definite evidence may resolve an unknown start. `MAX_CONFIGURED_SUBAGENTS` bounds latest actor rows, with old single receipts readable. Dynamic Runtime context and owner UI consume history without changing catalog, admission, quota, dispatch or ranking.
**Read-only and mutating session rows share one nanny transport.** The only difference is the run shape, whose ONE owner is `subagents.delegated_run_shape`, asked one question — is this an acting child? — by `tools/delegate._derive_authority` (live `ToolContext`) and by `resolve_configured_actor_dispatch` (configured snapshot); a shape re-derived in two places drifts unsafely in exactly one branch.
| task authority | access | mode | isolation | `execution.delegated` |
|---|---|---|---|---|
| acting subagent (valid write surface) | captured `full` or `workspace_write` | `agent` | `live` | `true` |
| Ordinary root with a validated external workspace or a selected Project room | captured `full` or `workspace_write` | `agent` | `live` | `true` |
| top-level task selecting an exact skill payload (`root="skill_payload"`) | captured `full` or `workspace_write` | `agent` | `live` | `true` |
| anything else, including a fail-closed subagent or an invocation lowered to readonly | `readonly` | `ask` | envelope (default) | not sent |
WHERE a mutating run's changes are destined is the second, separate record — the host-derived **mutation authority** (`tools/delegate_integration._mutation_authority` / `_payload_mutation_authority`, re-exported by `tools/delegate`), never model-supplied:
| source | `target_root` derivation | `capture_mode` |
|---|---|---|
| `acting_constraint` | the child's own `task_constraint.write_root`, required to equal the genuinely ACTIVE workspace root | `delegated_snapshot` |
| `external_workspace_root` | the root task's validated external workspace or selected Project room | `delegated_snapshot` |
| `skill_payload` | the exact payload the fresh `skill_payload.write` binding resolved | `delegated_snapshot` |
| `readonly` | ordinary active root (nothing to write) | `none` |
A payload target gets a standalone private Git snapshot (`subagent_worktrees.provision_payload_snapshot`): the live payload is never initialized as Git, and capture trusts nothing under the child-writable snapshot's `.git`. Disposition (`integrate_payload_patch`) applies a live, index-free `git apply` in no-repository mode under a whole-payload content-hash CAS (drift = typed conflict; identical content = idempotent applied); `GIT_CEILING_DIRECTORIES` is pinned at the payload's resolved PARENT (git still searches the ceiling entry itself, and an ancestor Git worktree above the runtime data root could otherwise make git skip every hunk at rc=0), and reserved paths refuse the WHOLE apply as `blocked_reserved_paths` with the candidate preserved. The post-apply outcome set is complete: a live loader hash equal to the recorded RESULT hash is the success; a hash equal to the recorded BASELINE hash with a non-empty touched set is a provable non-mutation that RESOLVES the apply intent (typed `INTEGRATE_APPLY_NO_OP`, the `apply_no_op` arm: no success, no disposition, no reconcile queued, retry lane open); anything else is the ambiguous mismatch, whose intent stays PENDING and whose reconcile marker IS queued (the payload did mutate). A successful apply queues the extension reconcile (`request_extension_reconcile`) and the skill's review goes STALE pending fresh `skill_preflight`/`skill_review`; a run whose ONLY change is a mode flip is already refused at CAPTURE time as `unreviewable_metadata_change`, so a live hash still equal to the baseline means nothing was written.
**A mutating run normally executes in a PRIVATE EXECUTION SNAPSHOT** (metered children keep sharing the tree; their patches integrate through `integrate_subagent_patch`: sha256-bound, 3-way `--index`, protected-path gated, genesis refused, `coop_already_in_tree` a no-op). At `delegate_start` the host snapshots tracked, staged and eligible untracked state, deciding sensitive/credential vetoes BEFORE hashing (`git add -A` would put secrets such as `.env` in the shared object database). Git's binary verdict for the whole untracked inventory comes from one index-versus-worktree `git diff --numstat` over a scratch index (two when empty files need their attribute verdict; `workspace_patch_capture.untracked_binary_verdicts`, shared with patch capture; a failed batch falls back to the per-file verdict with a warning), never one process per file by design. The machine-wide worktree ops lock (`subagent_worktrees._ops_lock`) guards SHARED metadata only — the registry file, a target's `.git/worktrees` and its baseline pin — held twice, briefly (row FIRST, then the ref, then `worktree add --no-checkout`, so a crash after the row is GC-nameable); the acting `self_worktree` lane is split the same way (`worktree add --no-checkout -b` + branch under the lock, `reset --hard --no-recurse-submodules` and deletion outside it — the body's own `post-checkout` hook no longer fires at provision either), and only the boot-time `prune_orphans` sweep and the millisecond genesis `git init` still do their work under it; listing, classifying, hashing, populating (`reset --hard --no-recurse-submodules`: what `worktree add` runs internally, minus the target's `post-checkout` hook), copying the source's exact bytes over the checkout and one `update-index` re-recording their stat (a CRLF-converting checkout otherwise leaves every such file "modified" in the child's `git status`) run outside it (issue #1241: one 67k-file provision held the lock 40 minutes and every other mutating start timed out). A held lock refuses typed (`cause: lock_busy` + the holder's pid/task/op), a SIGKILLed holder is evicted by the owner-aware stale check, every refused snapshot provision is `definitely_unrun` with a durable `START_FAILED` row, and the start receipt discloses `snapshot` size/time. Disclosed residuals: the registry read-modify-write under the lock is O(registered rows, their per-file baseline maps included — moving those maps out of the row is a follow-up), and every eligible untracked text file is hashed into the TARGET's own object database, where it stays unreachable after the snapshot is removed until that repository's own gc. The baseline is pinned by `refs/ouroboros/delegated/` and checked out as a detached worktree; `scope.root` stays the authority target and `execution.workspaceRoot` names the snapshot. The host appends a separate typed binding after the immutable work order: the snapshot is writable and the authority is read-only until integration. Directory-copy runs use the engine-created copy and never the source folder. Full native access has no filesystem sandbox; the binding names the writable root but does not enforce it. Terminal capture records authority drift as evidence: ready-no-changes stays no-change with unknown authorship, while ready-with-changes keeps its private artifact and the locked baseline proof decides integration. Nested Git directories are excluded and disclosed; skill payloads use content-hash CAS. The binding is durable before POST and retry replays it; a GC-collected snapshot is a typed `execution_snapshot_missing` refusal. Worktrees live in `state/subagent_worktrees.json`; removal is explicit or custody-cross-checked startup GC, fail-closed on unreadable custody. The run still uses `live` from the engine view, so the scoped-HOME/`delegated` marker below applies.
At terminal, `delegate_wait` captures the run's diff against the baseline durably into the task's artifact store; NOTHING reaches the target automatically: the nanny explicitly applies or rejects through `integrate_delegated_patch`. Git and skill captures are whole-result operations: omitted `paths` and an exact empty list select the same captured result, disclosed when explicit; nonempty selectors stay directory-only. Engine-directory empty/subset semantics are unchanged. The staging substrate differs: a GIT workspace target applies under the repo git lock after PROVING no touched path drifted from `baseline_sha` (a plain `git apply` relocates hunks by offset; the touched-path set is read from `git apply --numstat` in BOTH directions, each naming only the paths it writes), then applies and STAGES, never commits. A SKILL-PAYLOAD target captures through the payload adapter over a parent-owned trusted index and applies LIVE into the non-Git payload: nothing is staged into any active root and no `.git` or index is created in the payload. The protected-path gate applies only when the target IS the Ouroboros body; a conflict (proven drift) is owned by the still-running nanny, with snapshot and patch persisting until explicit resolution or discard. Mutation rides an apply-intent protocol: a durable `delegate_run_patch_apply_started` row lands before any tree mutation, so on replay a pending intent without a disposition answers typed `INTEGRATE_DELEGATED_APPLY_AMBIGUOUS`, resolved only by explicit `acknowledge_ambiguous=true`, while the provably non-mutating outcomes (a lock error, proven baseline drift, a failed apply, a verified revert, a baseline-equal payload hash) RESOLVE the intent. `patch_verdict.py` is the ONE verdict writer for both pipelines: subjects are minted `run_<rid>` by the writer, never prefix-matched by readers, and each decision lands twice (artifact plus typed `delegate_run_patch_verdict` custody row), a failed artifact write disclosed on the row. `artifacts.delegated_capture_read_target` narrowly rebinds `artifact_store` READS for the owning task's own `delegated_runs/` prefix, and `delegate_shared.orphan_capture_read_target` extends the same read-only, one-directory READ to the terminal-owner ORPHAN the disposition rule authorizes (confirmed by `orphan_disposition_status`), so the actor that may dispose a patch can inspect it without wider write authority. A read-only child stays in Claudexor's default envelope: one transport with one derived difference, not a second pipeline.
**Terminal reconciliation captures only over PROVEN terminality, and `patch_captured` means "a usable artifact exists".** When the owner task is gone, a settled mutating run's diff is captured through `delegate_integration.capture_terminal_patch_for_drive`, capture only, never an apply: the decision stays with an owner, and the obligation stays visible (`delegate_custody.undisposed_patches`) until the durable `PATCH_DISPOSED` row clears it. The obligation is held by the THING (the run's custody rows and the target itself), not by the task that created it, so a terminal owner's run keeps its debt visible and disposable instead of locking the target forever. While the owner task is LIVE, only that identity may dispose its captured patch; once it is terminal, any live TOP-LEVEL task may. Apply requires the caller's active Git root or fresh payload binding to equal the run's recorded target, or, for an ORPHAN only, to CONTAIN it while both lie under the host-minted subagent-projects root (the swarm aggregator shape: `<project>/contributions/<track>` clones the host already checkpoint-commits), every apply-path guard unchanged. Reject requires only the owner's proven terminality: it releases a dead task's locks and snapshot without fresh target authority. The `PATCH_DISPOSED` disposition row records who did it as `disposed_by_task_id`, and the `delegated_runs_unreconciled` projection stays the evidence trail the sweep heals. An owner whose terminality cannot be proven (task result missing, unreadable, or still an unreaped `running` row) keeps the lock, with no time-based release. A run closed absent or left unreadable captures NOTHING (across the provisioning boundary the child may still be alive and writing; an eager capture would freeze an incomplete patch and serve it forever); the snapshot stays preserved for capture-on-demand, and `patch_captured` is minted only over a ready manifest so a failed manifest leaves every retry point open.
**The stored `delegated_runs_unreconciled` projection is healed only from the write side, at four seams.** `delegated_custody_unreconciled` is the disclosure a task carries when it wrote its terminal result while one of its OWN delegated runs was neither applied nor rejected; a review-owned run is never one of them, so a task whose only open runs are its reviewers audits clean and carries no custody notice (the review projection's `late_result_pending` discloses a still-running panel). The overlay is ADDITIVE (`outcomes.custody_debt_axes`): it merges an objective WARNING and sets the top-level reason code but never rewrites the derived execution, review, objective or artifacts axes (the paid verdicts the task earned survive); a rail truncation code from `BEST_EFFORT_REASON_CODES` keeps the single Reason line (a round-limit or budget-exhausted victim is more usefully labelled by its truncation). While the debt is open the task reads Done with warnings, not Failed (the raw code stays typed on the row); the debt list lives on `delegated_runs_unreconciled` plus the `delegate_terminal_reconciliation` envelope. The stored code stays historical; the owner-facing Reason line is resolved at render time by the Task lifecycle rule: host rows in `project_dialogue._custody_debt_reason`, the browser card in `log_events.taskReasonDetail`, both from the stored list the live `task_done` event copies from the row (`agent_task_pipeline._custody_debt_event_fields`), so no surface infers a debt it was never given. Disclosed residuals: a provider-death terminal carrying custody debt shows the custody code on its one Reason line while the debt is open and remains Failed (the devtools-only infra codes are not in the runtime best-effort set); a CLEAN derived outcome plus custody debt does not increment `evolution_consecutive_failures` (a warning is not a failure); and `project_dialogue` derives failed/degraded from axis STATUSES only, never `objective.warning`, so a custody-debt child reads as a plain success there. Readers serve the stored projection (projection-over-replay, no live custody join), so a run settled after its task's terminal write stays stale until one of four seams: the periodic sweep refreshes tasks named in its own reconcile outcomes; the boot backfill (`delegate_terminal.backfill_terminal_reconciliations`, once per generation) re-audits stored terminal rows still carrying a disclosure; the cursor pass (`delegate_terminal.refresh_recently_settled_terminals`, a durable byte-offset cursor `state/delegate_terminal_refresh_cursor.json` over the append-only custody log, 5 MB per tick) catches terminal-boundary settlements no outcome-driven refresh can reach; or a kill path clears a stale list. Every refresh is audit-only in both directions (unreadable custody proves nothing) and never rewrites `reason_code` or recomputes the frozen `delegated_runs_*` counters: counters stay a historical snapshot while `actual_substrate` and the envelope mirror are rewritten from live custody, current liveness lives in the `delegate_terminal_reconciliation` envelope, and patch debt survives every refresh as `patch:<run_id>`, never a blind clear.
Disclosed delegated-isolation residuals (deliberate): a live top-level task with a different active root may reject and release another dead task's snapshot; the orphan-apply widening by containment means the disposer's blast radius grows to exactly the host-minted descendants the checkpoint-commit already writes; an orphan disposed by a non-owner writes its verdict artifact and `delegate_run_patch_verdict` row under the DISPOSER's task while the capture directory and the `PATCH_DISPOSED` row keep the OWNER's task id (readers key verdicts by `run_<rid>`, but the owner's acceptance packet will not list that verdict); a GC-lost snapshot or permanently failing capture discloses in a typed refusal that its obligation can never be satisfied; the baseline is worktree-primary; the git lock is task-drive scoped, so two nannies integrating into the SAME external tree can interleave apply+stage (the drift check makes the loser's apply a typed conflict); and a credential-shaped file the CHILD creates fails the whole patch instead of shipping a partial diff.
**A mutating run requests its captured native profile, reads back what it got, and DISCLOSES the gap instead of refusing the work.** Full requests no filesystem sandbox; the private snapshot owns delivery, not OS confinement. The owned gateway creates a scoped full-access grant only when no trust record exists; an existing denial is preserved and addressable by lowering the invocation or capping the row. Grants persist per scope like stable project registrations; no automatic trust cleanup is implied. Workspace-write keeps its boundary disclosure. In place because Claudexor otherwise hands the harness the operator's real `$HOME`, which holds the daemon token (a compromised child could start its own runs at any access level). Four facts, one mechanism. (1) The `execution.delegated: true` marker rides in the same record as `isolation: live`, built from `delegated_run_shape` in one place, so one cannot be sent without the other. (2) The version floor is about the SCHEMA: `config.CLAUDEXOR_DELEGATED_MARKER_MIN_VERSION` (3.3.0) is the oldest engine whose strict `RunExecution` accepts `delegated`; below it the start is a 400 and no run exists (`route_health` blocks dispatcher and nanny identically before a token is spent). Read-only delegation sends no marker and keeps the 3.2.0 transport floor; an engine between the two floors serves read-only and refuses mutating. The floor is a schema fact, never a containment fact (the OS boundary is platform-dependent, a build declares one version everywhere); threat model and measured bands: `docs/DELEGATED_ADMISSION.md`. (3) What was APPLIED is asked of the attempt, never of the OS: `delegate_wait` reads `harness_home_isolated`, `confinement_mechanism` and the proven denied path from the attempt record (`gateways.claudexor.attempt_containment`, `delegate_containment.py`); `sys.platform` appears nowhere in the decision. (4) A missing boundary is disclosed (durable `delegate_run_unconfined` event, the child's instructions, the parent's terminal payload), not refused (the child already holds a shell in this worktree; cutting the lane on every boundary-less host costs more than it prevents); a home nested under the operator's own is disclosed-unconfined (`home_nested_under_operator_home`), never relabelled as isolation. A recorded FALSE is still a fault: `harness_home_isolated: false`, or a scoped home that IS the operator's own, cancels as a typed containment fault exactly like a widened access profile (those two exact facts are the WHOLE breach rule); a MISSING home fact is neither breach nor proof, so unproven is REPORTED.
**Delegated authority cannot widen.** Start exposes `prompt`, session `subagent_id`, `max_seconds`, lowering-only `access` (`readonly`/`workspace_write`), recovery `retry_of`, continuation `continue_from` and the payload selector; continuation refuses beside retry/payload selection. No mode/isolation/scope argument exists. Retry preserves route/root/access/permissions; host instructions remain unforgeable. Native process access never overrides task constraints or the assigned edit target. Reviewers remain readonly/ask even if an actor row allows writes. Every fetched detail passes `_containment_breach`, checking both access and harness HOME because Claudexor derives effective access rather than echoing the request. Wider access cancels as `access_profile_widened`; narrower access is valid.
**Read provenance on the accounts surface.** `GET /api/claudexor/status` carries a `reads` block (`ClaudexorStatusReads`: `catalog`/`accounts`/`quota`, each `ok`|`not_read`|`failed`) (the owned daemon starts lazily): an idle daemon serves empty collections under a 200, and "no account connected" must not be inferred from a collection that was never read: `ok` makes the matching collection authoritative (empty means empty); one parity-tested client reader (`facetReadState`) applies the same rule, and the aggregate `daemon.state` is never the negative answer. Login jobs: the daemon stays the sole process/fence authority and reconcile is an explicit POST, never passive polling (`gateway/claudexor_accounts.py`; the route inventory is its §1 row and §4).
### Git and commit review
`tools/git.py` owns repository writes, staging, reviewed commit, rollback or restore, tags, push and CI follow-up; its leaves are `git_plumbing.py`, `git_repo_edit.py`, `git_vcs_ops.py`, `git_review_cycle.py` (staging plus the advisory/triad/scope review and the reviewed-material binding the commit gate consumes) and `git_evolution.py` (campaign authority at the reviewed-commit and publication boundaries). File-edit tools validate their own atomic write shape. `mutation_attribution.py` captures the root-task baseline and projects the clean-at-baseline system-repository delta plus an explicitly adopted predecessor's exact retained changes — a changed pre-existing dirty path, a stale or missing baseline, or a failed scan blocks automatic staging; `commit_reviewed(paths=None)` stages only that attributed candidate, explicit paths must be a subset, an empty candidate returns `GIT_NO_ATTRIBUTED_CHANGES`, and managed update transactions keep their separate typed whole-tree authority. At startup a later independently initiated task may adopt exact retained candidates from its host-validated predecessor source in the same repository. `predecessor_adoption` preserves the observed dirty baseline and original source; terminal quiescence, unchanged path content and unchanged per-path base are required, while unrelated dirty work stays excluded. Missing fingerprints or size-only observations establish no transfer, and adoption grants no review approval or automatic new task. Commit preparation verifies the exact local working-branch ref before an unambiguous checkout: a missing ref refuses without changing the current branch, index or files (remote guessing and implicit branch creation are disabled), detached work retains the `checkout -B <branch> HEAD` recovery, and managed assisted merges retain transaction-owned precommit verification.
A reviewed commit is bound to one staged fingerprint. A cheap LLM-first advisory pass may run before the expensive gates; it is advisory, and skipping it never skips independently applicable tests, triad, applicable scope review, aggregation or exact-SHA binding. The hermetic preflight runs the candidate in a disposable worktree and data root; triad and scope inspect the same staged snapshot, aggregation preserves actor evidence and obligations, and any mutation stales the binding. Managed exception: a managed-update resolution commit reviews the declared M0→S subject (`tools/review_subject.py`) and the commit gate binds S to the exact index write-tree the fingerprint pins. External review wrappers report readiness but do not grant commit authority. The exact binding includes the `git write-tree` SHA, ordered `HEAD` and `MERGE_HEAD` parents, indexed VERSION, expected `v{VERSION}` tag, any existing tag target, and the binary staged-diff hash; after commit, tree, parents, VERSION and tag target are re-read before success or push is recorded, and an existing release tag is never silently accepted or retargeted. `release_sync.sync_release_metadata` projects version carriers during ordinary commit preparation; `VERSION_CARRIER_SPANS`/`substitute_carrier_spans` is the ONE span primitive the managed-update resolver and the commit-triad pack cut share, and `carrier_only_change` names a carrier changed only inside its declared version spans. Durable review state keeps attempts, obligations, readiness debt, raw actor evidence, and the final commit or tag binding; raw advisory output lives in `state/advisory_review.json` (selected by `snapshot_hash`/`ts`), not in `review_status(include_raw=true)`, which exposes raw triad and scope attempt evidence; `commit_gate.py` classifies review blocks, refuses an identical verdict, counts paid cycles against the ceiling and fingerprints the review contract. BIBLE supplies review authority, CHECKLISTS supplies criteria, and Development supplies the procedure; snapshot identity, advisory coverage or audited-skip evidence, deterministic results, actor evidence, and final Git identity must all describe the same material.
Material ordinary Advisory commit review returns the complete outcome and exact `review_reference` before commit, tag or push. The author may correct, accept unchanged bytes, request another permitted review or stop. Explicit `commit_reviewed`/`vcs_commit_reviewed` continuation with that reference and `author_disposition` binds the current attributed candidate, reruns independent required preflight/tests and exact Git checks, and dispatches no critic. The author record references the original attempt instead of rewriting its subject or paid facts. Settled handback is `reviewed/review_only`; actual pending custody remains `reviewing/late_wait` and collectible. Clean supported review keeps its one-call path; Blocking gets no author override. Evolution receipts record actual `triad_scope_status`, never infer PASS from a successful Git commit.
#### Commit review evidence
`review_evidence.capture_commit_review_evidence` freezes selected browser/vision calls and same-round automatic image-attachment observations, explicit unavailable-image gaps included, after cheap/free admission and before preflight/triad/scope. The borrowed loop trace and original call refs retain exact redacted arguments/results and the immediately following visible response in the same execution, excluding provider thinking; the original context binding is restored through the attempt's `_LoopExitContext` cleanup even if `tools._ctx` changes while the loop runs. Adjacency never attests visual inspection. One canonical task source handle holds the selected UTF-8 view, its initial exhibit within `_ACCEPT_NOTES_CAP` with counts, completeness and source identity: native reviewers read artifact-store ranges under the real canonical root, sessions receive byte-identical bytes in ignored `.review-drive/<task>/<sha>.txt`, and packet-only reviewers receive a bounded, explicitly partial view when needed; existing request evidence/refs and the preflight execution `evidence_source_ref` bind it. A pending rejoin restores the same source, a recorded empty selection included, and carries it from preflight into triad/scope without selecting new trace. Commit cleanup waits for physical custody; standalone or uncertain views remain retained without a new cleanup registry.
#### Hermetic preflight proof
The hermetic runner (`preflight_runner.py`) alone mints `ctx._preflight_test_proof`, after actual green lanes and containment; it binds the assembled checkout tree and installed index, effective pass specs, isolated environment (`test_environment.py` owns writable roots), and the invoked/resolved interpreter and active Node identity (a Python symlink's invocation path selects its venv). The candidate is ONE hardened raw-bytes `git diff --binary … HEAD` capture, because a staged+unstaged diff pair cannot faithfully materialize unmerged entries; a capture/apply failure is the typed `PREFLIGHT_CANDIDATE_ASSEMBLY` block, never a test verdict. The node lane runs first (`preflight_node.py`: bundled signed node then PATH, floor 20.11; `PREFLIGHT_NODE_MISSING`/`PREFLIGHT_NODE_TOO_OLD` hard blocks and `NODE_TESTS_FAILED` on a red suite, never a silent skip), then the two-pass parallel/serial pytest split under one budget (`LANE_EXCLUSION_EXPR` is the marker-lane SSOT); a dead xdist worker or a missing required plugin is a distinct named block, never a retry or silent serial fallback. Equivalent ordinary and managed checks reuse that process-held proof only after their distinct baseline checks; a phase label alone cannot establish equivalence. Every workload binds HEAD, because even an unmarked default-lane test may inspect committed files or history: repeated preflights can reuse an unchanged subject, creating a commit changes HEAD and requires a new run, a skip or no-suite `None` is not a green proof, and a restart loses the proof and reruns the suite. Creation and reuse emit `preflight_test_proof` observations through the existing event log, naming the task, HEAD, tree, index and workload fingerprint; an unavailable log falls back to an explicit diagnostic, never a test failure or a new authority source. Managed `tests_evidence` is forensics, not reuse authority. Unproven lane teardown retains the temp tree and Git registration (`_release_hermetic_tree`), including raised lanes; deleting inputs beneath surviving processes would destroy custody.
#### Commit advisory cycle
`preflight_review(commit_message="...", deterministic_only=True, source="worktree" | "index")` returns a release-only diagnostic report before sync, staging, custody/state access, provider checks, tests or fingerprints. `commit_admission.release_metadata_diagnostics` owns source acquisition and reports all independent applicable `findings` plus `unavailable` sources; `release_sync.release_metadata_findings` reuses the carrier validators, release grammar and P9 history counters. The report identifies its source and status (`clean`, `blocked`, `unavailable` or `not_applicable`), gives no review freshness, and writes no review history. It checks the selected VERSION, not a future release. VERSION/README absence and failed reads are unavailable evidence; an absent optional older carrier stays optional, while a readable malformed carrier is a finding. Coverage is release metadata only, not syntax, structural size, tests or critic review; size policy stays warning-only locally.
Standalone advisory retains automatic carrier sync and worktree reads. Prepared advisory follows the commit gate's whole-index applicability, regardless of a narrower paths hint, including partial staging and the version-neutral carve; standalone documentation-only scope stays exempt. A present malformed VERSION is now a finding, not a silent skip. Optional[str] wrappers format the complete release findings, separating unavailable evidence from defects. Author continuation uses the shared name-status formatter and unavailable classification. Changelog prose, history trimming and version allocation remain deliberate.
Advisory availability is evaluated from the current configured slot and route, never inferred from a stale stored verdict: a disabled advisory slot is an audited bypass, an `api_chat` row requires provider credentials for its RESOLVED model, an `agent_session` row a resolvable session route. `claude_advisory_review.preflight_review` (callable `advisory_review` alias) owns that admission policy: an `api_chat` row rides the native tool-round episode (§6 Review stack), an `agent_session` row the session review executor; a native episode that ends on its own transcript bound — keyed on the structured `native_transcript_cap_exceeded` code, never on message text — reaches the caller as the typed non-blocking `ADVISORY_SKIPPED` with reason `native_transcript_bound_exceeded`, carrying the bound, refused chars and paid rounds, and every episode exception keeps `failure_custody()` as the advisory meta's `usage`, never an empty `{}`. The retrieving advisory brief carries a touched-path manifest instead of duplicated file bodies and applies the shared span-only release-carrier cut over HEAD→working-tree, with the same `PACK EXCLUSION NOTE`; omitted bodies remain readable through its own tools. Governance comes from the shared tiers (§6 Governance delivery). If the commit advisory is unavailable, the commit gate runs its compensating hermetic preflight only when tests remain independently applicable (not explicitly skipped, diff not documentation-only). Readiness projections receive only an exact repo/hash-matched advisory record and keep its failed status and freshness.
The reviewed-commit cycle (`git_review_cycle`) checks authorization and unresolved prior work before mechanical preparation and staging, binds the exact candidate before free-cycle/budget admission, runs a needed preflight inline with the full rebuttal and applicable tests, and revalidates the index and worktree snapshot before triad/scope; it owns entry-point resets and pending/blocked finalization, and its custody check joins current reviewer facts with the strict durable advisory record before index cleanup, an interruption before local metadata was updated included. `AdvisoryRunRecord.execution_pending` preserves physical custody for history retention and the external wrapper; `blocks_preflight` separately tracks logical admission. An explicit audited bypass releases that admission while retaining the original task, invocation, source and unknown physical outcome; late results update their original history row without superseding the newer bypass. Free advisory replay still checks freshness and buys no automatic preflight; stale coverage requires an explicit audited skip with applicable compensating tests preserved.
`AdvisoryRunRecord.execution` binds delegated preflight intent and candidate to the existing invocation token before POST and retains terminal usage/receipt evidence on failure; `delegate_custody.invocation_record` owns the immutable request, full prompt included — no second prompt store. An exact pending rejoin restores that request through the shared executor: a delivery-kind change cannot replace an unresolved session, same-session model/profile changes still rejoin the canonical request, and changed, foreign or lost identity is refused before preparation, naming the mismatched fields; `review_status` projects the recorded rejoin intent and execution identifiers through the public secret redactor without exposing the private reviewer prompt. Failed checkpoint writes cannot dispatch; unresolved rows survive history trimming. Only a corroborated `failed_definite` invocation discharges a stranded logical checkpoint automatically — age, missing identifiers and unreadable state cannot prove never-started work — and an audited skip never posts a replacement preflight. Released history stays eligible for an exact delegated rejoin but does not lock a later explicit standalone request for different intent or evidence; an unchanged audited bypass can satisfy the freshness shortcut without a new model call, naming the actual bypassed status. Native preflight uses its executor's end-event and monetary custody, with no separate advisory operation checkpoint or native resume protocol; a returned unknown physical outcome remains failed evidence on the existing refusal and audited-skip paths, and a hard process death has no native exact-rejoin handle. The external wrapper runs the same cycle, exports the real advisory record to `advisory.txt`, and preserves a pending checkout without restaging.
### Review stack
`review_cycles.py` owns the shared paid-cycle ceiling `OUROBOROS_REVIEW_MAX_CYCLES` (a positive integer or `unlimited`; anything else fails closed to the shipped default). One number, four gate meanings: paid reviewer-panel cycles per task for plan review; paid panel runs per task for acceptance, independently of ordinary author responses; paid triad-plus-scope cycles per ROOT task for the commit gate, counted from the attempt ledger at DISPATCH so the ceiling counts money and only undispatched attempts stay outside it; paid panel dispatches for skill review. Byte-identical material replays each gate's recorded verdict for free under that gate's identity. Physically dispatched technical failures consume capacity; undispatched refusal and collection do not. The last allowed panel still permits author reaction. An explicit task-local `max_improvement_passes=p` retains its p author responses and p+1 paid ceiling; without that override N creates no N−1 response limit. `unlimited` removes only the paid count, not ordinary task rails.
For skill review, convergence guidance to the reviewer and retry coaching to the author use the group's `review_round`, so revising payload bytes does not reset the coaching. `snapshot_attempt` remains the separate ordinal shown in history and headers; free replay keeps its exact material and contract identity. A paid-ceiling refusal names author finish only where the existing enforcement predicate permits it.
Review waiting has six independent axes:
| axis | bound |
|---|---|
| transport | dead-socket read bound |
| operation | typed active-operation lease |
| logical | slot/task deadline |
| policy | budget / cancellation |
| absolute | task ceiling |
| aftermath | late-result custody |
#### Physical custody
`review_custody.py` is the small worker-lifecycle seam under `review_substrate.py`: it schedules no tasks and keeps no second timing ledger. Parallel slots hold independent operation ids; the active-operation map only keeps a live physical call from reading as idle — deadline, budget, cancellation or ceiling still wins. Delegated-session expiry uses the verified cancel path; API/thread calls disclose `in_flight` and reconcile a late answer before the same retry identity can dispatch again, and a settled terminal API error stays in the cycle's replayable actor roster, so a sibling can neither erase its terminal fact nor buy a duplicate physical call. Physical custody is proved by the capture/operation state, never by the synthetic operation id, in `usage_accounting`'s vocabulary: only `reserved`/`released` proves pre-dispatch, a positive capture outranks a contradictory synthetic label, and `dispatched`/`unresolved` without a typed terminal HTTP status stays custody-lost/no-resend. Custody never infers pre-dispatch provenance from Python's implicit `__context__` (a fallback raised inside a prior provider handler inherits that earlier attempt): only an explicit `__cause__` or typed transport metadata can release a paid row. A retry token without a durable invocation is custody-lost on every delegated review surface, and a durable token is valid only for its recorded surface, slot and operation. Retry identity is the explicit material/cycle identity when supplied — mutable prompt or history is deliberately not part of it, while a changed snapshot, owner intent, reviewer route or admitted cycle mints a new one.
Timeouts compose, never override: send-time VLM captioning keeps its direct 90-second provider cap and Anthropic's direct route its 120-second provider default — neither is a generic review-reasoning cutoff — and the whole hierarchy is narrowed inside the owner deadline and finalization reserve before dispatch. A returned provider response or typed terminal error is settled even with an empty body, so bounded repair/retry may apply; a dead socket after dispatch is `provider_outcome_unknown` and cannot trigger another paid route. A spent owner window yields a typed `$0 not_dispatched` row before fan-out; under blocking enforcement an in-flight triad row stays pending rather than becoming a final quorum verdict, and a primary call at the deadline boundary enters the local finalization rail (`deadline_local`), not a provider-outage relabel. A reviewed commit has no independent outer tool cutoff: the foreground caller keeps custody until settlement.
#### Paid stamp and owner custody
The `paid` fact behind the cycle ceiling is recorded WRITE-AHEAD at physical dispatch (`review_dispatch.py`): a gate installs one once-only `ReviewPaidStamp` for its wave, and each delivery stamps at its own boundary — a packet api row at the usage ledger's physical-attempt transition, a native episode on each paid send, a session before its replayable `START_REQUESTED` row. Assembly-only refusals exit before that seam, so a $0 attempt stays outside every ceiling, and neither a worker outliving its logical caller nor a crash after dispatch can lose the durable fact. The stamp is idempotent across the concurrent triad and scope callers, so no side's transport starts before the paid fact is durable; commit review and task acceptance treat a failed write as fail-CLOSED, skill review only when resuming a wave, other callers keep fail-open accounting. Task acceptance binds a strict exact-hash tree-wallet claim to that same stamp on every delivery its rows run — one idempotent claim per panel, because the paid identity is material, not route; a wallet or cancellation veto at the claim releases the reservation and no reviewer transport proceeds on any row (the launch floor is evaluated once, at loop admission: §6 Task lifecycle). Disclosed residual: a compatibility transport that raises with positive physical capture without having entered the canonical marker is stamped after the send; if the tree's last paid cycle is consumed concurrently at that late stamp, the wallet refusal replaces the captured exception and the substrate may resend.
Before either parallel surface starts, one locked write records `paid=True` plus both complete slot rosters and operation ids in the commit-attempt row; a delegated slot patches its reserved row with `pending_invocation_id` before the provider POST. Exact resume preserves rows and tokens; a missing or mismatched operation is `custody_lost` under every enforcement mode. The stamp records the process-custody server session and pid owning the reviewer threads (`review_owner_custody.py`): owner-loss is proven by pid death, never elapsed time — a TTL would convert waiting into resend authority, and starting another Agent is not evidence this owner died. Tokenless rows settle as typed infrastructure failure only after a death seam confirms that exact pid gone, and a legacy row without owner identity stays fail-closed. Plan review re-enters a recorded in-flight cycle only through exact live custody or complete operation-addressed CAS for its recorded dispatched rows; a delegated poll that loses transport after a run id exists preserves the exact durable invocation for retry custody rather than cancelling a healthy unknown run.
Per-row delivery is configured only through the structured reviewer-slot SSOT (`reviewer_slot_config.py`); the older per-row route envs are retired and ignored. The one shared session target is `OUROBOROS_REVIEW_SESSION_ROUTE` (`review_execution.REVIEW_SESSION_ROUTE_ENV`), an opaque `harness[=model][:effort]` spec that falls back to `OUROBOROS_SUBAGENT_HARNESS` when unset, so one delegated route is configured once; a set but unparseable value turns review sessions OFF with a warning rather than silently re-routing them onto the subagent route.
#### Surfaces and money admission
The review surfaces:
- Advisory pre-review (`claude_advisory_review.py`) — a cheap, staleness-aware error-finding pass; the audited skip covers only advisory admission, never authoritative review, test policy or snapshot binding.
- Triad diff review (`tools/review.py`) — configured reviewer slots cover the Repo Commit Checklist with JSON findings under `config.adaptive_quorum` (the same SSOT as scope/plan/skill/acceptance review).
- Scope review (`tools/scope_review.py`) — intent/scope/coupling across the whole repository by retrieval on every row, in every context mode; the exact candidate, required-source manifest and diagnostic reading coverage stay distinct from substantive findings and configured enforcement (§6 Scope review by retrieval).
- Parallel orchestration (`tools/parallel_review.py`) — two-phase admission: the triad packet is assembled and fit-checked and every scope brief is prepared before any reviewer is dispatched, so a deterministic assembly block anywhere dispatches nothing and spends $0 everywhere (a lane paid while its sibling was doomed is real money for a half-verdict). Money admission (`tools/review_admission.py`: `commit_gate_paid_seats` / `admit_commit_gate_wave`) is all-or-nothing and scope-first: every PAID seat of the wave — the blocking scope seats FIRST, then the triad — is priced with the reservation math `reserve_attempt` will apply (its exact first send, including native work-order and tool schemas, plus its own output reservation; later native rounds reserve themselves), under the usage scope its substrate sends under (`review_usage_category` + slot, so a warm split of the caller's transcript never stands in for a seat's cold prefix), against every fence `reserve_attempt` enforces — the global `TOTAL_BUDGET` remainder (unbounded when the owner set no finite global budget) and the task's CURRENT root fence — and admitted as ONE wave through `review_wave_budget_gate(surface="commit_gate")`. A wave that does not fit is a typed `$0 not_dispatched` record on every seat plus one `review_wave_budget_insufficient` event carrying `binding_axis` (`global` or `root`) and both remainders, and the block names the binding fence, the shortfall and the knob that moves it — never a half-dispatched panel whose non-blocking seats hold the money the blocking reviewer needs. Admission is a read-only pre-check without a wave-level hold: the per-seat reservation stays the enforcement, so money a concurrent task takes on the global axis between admission and a seat's reservation can still refuse that seat (disclosed residual). A paid seat is any api row (packet or native episode); an agent-session row rides the owner's subscription — its ledger row is written at settlement, never reserved — and is neither priced nor waited for. Every executor transition on the way to a seat runs under `contextvars.copy_context()`, so the usage scope the wave was admitted with — its bound root fence included — is the scope the reviewer rows reserve under; a settings reload that changes the per-task limit mid-turn changes neither side. An admission that raises is fail-open like the gate's own unknowns, but loud and typed (`review_wave_admission_unavailable`; the wave dispatches unadmitted). An admitted wave submits the scope seat first and starts the triad only once the scope seat's OWN reservation is on the ledger (nothing else releases the hold), bounded by `NESTED_SETTLEMENT_MARGIN_SEC` and by the scope seat itself; a hold that ends without observing it emits a typed `review_scope_lead_unobserved` event. Admission touches no transport or logical timeout contract.
- Shared helpers (`review_helpers.py`, `triad_review.py`) — packet/brief building, checklist loading, JSON extraction, usage events, obligations scaffolding and reviewer actor records.
A free refusal never wears the form of a verdict: a refusal that spent nothing is recorded as a typed `not_dispatched` fact plus its reason, never as a DEGRADED panel, a synthetic actor or a verdict. That is the one shape across every $0 exit — a plan-review locator the evidence policy cannot attach is a named omission row the panel is dispatched with, an acceptance packet that overflows takes the ladder rather than reporting a verdict, a truncated or self-pageable row names its cut, and a request never sent is one seat with `operation_state='not_dispatched'` that stays in the denominator.
Rationale: diff reviewers catch line-level mistakes; the scope reviewer catches cross-module contracts; running both on the same staged snapshot prevents one result from hiding the other. The managed-update exception reviews the same declared M0→S subject in both lanes. Task acceptance is a separate root-owned post-delivery system specified once under Task lifecycle; it shares `adaptive_quorum` and these transports, has no acceptance scope actor, and stores its verdict separately from the terminal result.
#### Session identity and advisory parsing
Session reviewer identity comes from `gateways.claudexor.final_attempt_facts`: the unique `final_attempt_id` row in the engine-owned `final/telemetry.yaml`, bound to the requested run id, and model, harness and credential profile are read from that SAME attempt. `summary.model`/`harnesses` echo requests, and summary route/auth projections may borrow earlier-attempt facts; none supplies a missing observation. Reviewer usage, last-execution views, delegated terminal payloads and settlement preserve known facts or explicit absence; the original custody/billing route remains the chosen authority, separate from the observed actor. No quorum rule or model-name mapping follows from this disclosure.
Advisory row parsing owns its `PASS|FAIL` and `critical|advisory` values at `preflight_review_run._is_checklist_array`, canonicalizing case and surrounding whitespace once for every consumer. An unknown verdict, or an unknown or missing FAIL severity, rejects the whole array into the bounded extraction rail; unresolved output stays `parse_failure` with its full source retained. PASS without a severity remains compatible and the separate genuine-empty-clean predicate is unchanged; an optional array validator lets the shared canonicalizer honor this surface contract without changing triad, object-verdict or report semantics. Canonicalization never changes the reviewer's judgment by searching the repository for words or identifiers.
#### Structural gates
Structural smoke gates are a deterministic BIBLE P3 codebase-size component. `ouroboros/review.py::iter_gated_modules` is the one source inventory for smoke, `codebase_health`, census and the UTF-8 byte gate (Python everywhere plus first-party `web/**/*.js`, vendored/minified excluded). `ouroboros/size_ratchet_manifest.py` is a generated, data-only debt register consumed through AST literals: exact module debt above 1600 lines, exact function debt above 300 lines, the 1001–1500 band with rationale authority, and exact byte debt above 200,000 UTF-8 bytes. `validate_size_ratchet` proves the live and staged manifests exact against their trees and shrink-only against the merge-aware committed authority — no first-parent history replay: the previous manifest resolves from `HEAD` or any of its parents, and a checkout with no committed manifest anywhere bootstraps from its own tree, so a fork whose local line predates the manifest is never condemned by inherited topology. Official-repository CI runs the blocking `size_ratchet` pytest lane while every local surface reports the same findings as warnings; two disclosed residuals — pairwise validation covers only the base→HEAD interval, and the official block presupposes branch protection. Within a validated pair, debt can shrink but cannot be swapped, re-entered, grow on the byte axis or survive stale. `MAX_TOTAL_FUNCTIONS` (`ouroboros/review.py`) remains the coarse runtime ceiling. `codebase_health` and review readiness also report informational headroom from that same inventory: runtime function count, LF-normalized module lines/UTF-8 bytes and function size against current limits. The overview prioritizes ordinary near-limit modules before registered debt; readiness selects touched paths. Debt and omitted rows are labelled. Positive capacity stays separate from warning findings and adds no gate or setting.
A deterministic hot-store growth invariant sits beside these gates: `agent_startup_checks.py::hot_store_growth_notes` (surfaced by `context_health.py::build_health_invariants` and once per worker boot) stats eight hot stores plus the `archive/chat_*.jsonl` aggregate — `state/usage_attempts.jsonl`, `logs/events.jsonl`, `logs/tools.jsonl`, `logs/supervisor.jsonl`, `logs/task_reflections.jsonl`, `logs/progress.jsonl`, `state/scheduled_tasks.json` and `state/skill_review_root_tasks.jsonl` — against justified byte thresholds in `ouroboros/context_budget.py` and emits a WARNING with a remediation pointer.
Three gen/verify inventories ride the same discipline as the size manifest (generator `scripts/regenerate_inventories.py`, verify `tests/test_generated_inventories.py`, staleness = red): the frozen-contracts inventory (`docs/inventories/FROZEN_CONTRACTS_INVENTORY.md`, a machine extraction of §11.1 with every owner/anchor path resolved against the tree and the `ouroboros/contracts/` package-coverage gap pinned), the data-layout inventory (`docs/inventories/DATA_LAYOUT_INVENTORY.md`, every entry of the §1 "Data layout" tree probed as a tracked repo path or a runtime-source literal, so a durable file renamed in code while its tree row survives turns red), and the facade inventory (`docs/inventories/FACADE_INVENTORY.md`, the AST-derived `noqa: F401` re-export surface with per-leaf domains from `ouroboros/domains.toml`).
#### Prompt size, density and windows
`review_helpers.REVIEW_PROMPT_TOKEN_BUDGET` bounds assembled inputs: triad packets, per-slot plan packets, task-acceptance packets and skill-review chunks with their own headroom. Scope and deep review assemble briefs, not repository packs; native sends use `review_native_transcript_bound` and reach other sources through tools. Every surface scales its output reserve to the route's usable window (`reviewer_window.window_scaled_reserves`); a reserve sized for a much larger window can otherwise consume all input room before a request reaches the provider.
Sizing is density-calibrated, not keyed to model names. `usage_accounting.execute_physical_attempt` records settled `(prompt_chars, prompt_tokens, route_fp)` witnesses in `capability_evidence.json`. Main uses the newest fresh exact-route witness, then exact-model witness, then neutral density; reviews use the densest fresh exact-model witness, which can undercut `COLD_START_TOKEN_DENSITY`. Stale, absent or cross-model evidence keeps the conservative floor. The witness TTL and bounded retention preserve both its densest support and recent observations. `calibrated_input_token_limit` combines the shared cap, measured density and absolute margin per call; triad, plan, acceptance and native transcript sizing consume it rather than freezing an import-time result.
Only the triad packet has a cold-start probe rung: `review_admission.density_probe_before_size_refusal` invokes `capability_evidence.cold_start_density_probe` when an irreducible packet exceeds a cold route's cap. One bounded slice of the actual prompt is sent on that exact model under `physical_attempt_limit(1)`; its witness buys one resize and fit pass. It never probes a warm route, retries the probe or runs when the packet fits. Progress and `review_density_probe` retain the attempt; a paid-ledger refusal is `budget_refused` and leaves the size refusal unchanged. This breaks the loop in which a request rejected before dispatch can never teach its tokenizer density.
`reviewer_window.resolve_reviewer_window` is the route-specific sizing source, with metadata probes serialized per route and limited by evidence TTL. There is no model-window table or window-authority floor. `ReviewerWindow.sizing_window` has a full-window sizing default for an unknown API window; raw subscription routes keep no numeric unknown-window assumption. `scope_window` adds its designated-default or conservative fallback and distinguishes confirmed, asserted, stale-unverifiable, designated-default and unknown provenance. `POST /api/owner/capability-ack` records an asserted window bound to one route fingerprint, invalidated by a route change; it sizes sends and grants no reviewer authority. A failed send supplies no verdict, while a small window alone does not remove a responding reviewer.
#### Scope review by retrieval
Every scope row retrieves in every context mode. A native API row preserves its configured identity, model, route, account pin, effort, processing preference and local-route flag while the surface selects inspection delivery; a session row uses its configured delegated session. The first review where a stored bare API row acquires episode semantics emits `review_scope_delivery_migrated` and writes `state/scope_delivery_migration.json`: changed cost and duration are disclosed before becoming an unexplained bill. Failure to write the marker repeats the notice later rather than failing review.
`review_admission.prepare_scope_review` builds each brief through `scope_review_session.build_scope_session_task`, without dispatch. It carries intent, rebuttal, history and open obligations; the touched-path dispositions and candidate sizes; the repository index; shared governance tiers and scope checklist; calibration/output contract; task evidence; and the required-source manifest. Stable governance precedes change-relative context for cache reuse. The full staged diff stays inline while the first send fits, measured by `native_first_send_chars` including schemas and wrapper. Otherwise it becomes one exact paged source in the task artifact store, with a byte-exact `.review-drive` view for a delegated session. An unavailable store is disclosed and keeps the diff inline. Native inline sizing uses the landing threshold within `scope_first_send_bound`; sessions use `SESSION_INLINE_DIFF_CEILING_CHARS` because the harness owns their context selection. A managed-update subject is the authoritative M0→S resolution delta, never a substitute `git diff --cached` over the two-parent candidate. Invalid roots, missing subjects and failed assembly remain typed pre-dispatch failures; only an irreducible first send reaches `native_bound_below_first_send`.
When context construction fails before dispatch, preparation best-effort appends one `scope_review_preparation_failed` row to the existing events log, whose normal sink forwards it live. It names the task, slot, final model and original `context_unavailable` reason. This is row-local preparation evidence, not a panel verdict; logging failure preserves the same typed result, and model-control exceptions still propagate before logging.
`scope_required_sources.py` declares the change-relative reading minimum: touched protected runtime, prompts and frozen contracts; protection-owned families (`GIT_OPS_FAMILY_PATHS` and tool-dispatch leaves); and declared Python/browser contract twins, with both sides included. An ordinary touched file needs no whole-body obligation because the diff already carries its complete change. Each readable row binds raw-byte `source_revision`, normalized-text `complete_sha256`/`complete_chars`, range basis and candidate tree. Deleted or renamed sources retain their exact baseline preimages; paged diffs join the same manifest; unavailable sources remain explicit rows. Byte-identical governance already inline is satisfied by delivery without another read. The final rows are shared by brief and policy, and `required_sources_ref` plus `SCOPE_REQUIRED_SOURCES_POLICY` bind the actual manifest and review-contract fingerprint, so old replay authority cannot survive a changed contract. This minimum never restricts the reviewer's further reading.
Coverage records what could be observed: native receipts match source identity, opened root/path and delivered character intervals; sessions fold weaker journal evidence over the same rows. States are `complete`, `incomplete` with missing ranges, `declared_empty` for an explicitly empty manifest, and `unobserved` when measurement is unavailable. A changed source is a `source_gap`; availability or inline delivery never proves understanding. `review_context_atlas.repository_index` supplies orientation only: compact tracked-path dispositions, collapsed excluded classes, touched-file and direct-importer facts (size, digest, language, symbols and imports, with bounded totals disclosed). It renders no bodies, spends nothing, writes nothing and never limits what the reviewer may open.
The scope reducer keeps received findings and counts responding reviewers through `config.adaptive_quorum` independently of coverage. Reading gaps remain actor and panel diagnostics: they cannot change status, block a commit or trigger another paid review. The agent judges whether a particular gap needs more reading. Substantive critical findings still follow `OUROBOROS_REVIEW_ENFORCEMENT`; an unanswered reviewer does not count, and an all-`not_dispatched` panel retains its assembly/admission failure. Oversize requests, route failures and candidate-binding failures never become PASS; budget, custody, deadline and ownership keep their separate contracts. Retrieval keeps the reading obligation proportional to the change while exposing the actual evidence, instead of letting total repository size decide whether review can occur. If the change cannot fit a reviewer, split the change rather than weaken the reviewer.
#### Governance delivery
`tools/governance_context.py` owns the shared tiers for triad, scope, advisory and deep review. Tier 1 carries the applicable CHECKLISTS section, BIBLE.md and CHECKLISTS_ARCHIVE.md inline in a stable prefix. Tier 2 selects the review-protocol chapter, DESIGN for `web/` changes and DEVELOPMENT chapters naming touched files, within `runtime_limits.REVIEW_GOVERNANCE_INLINE_SHARE` of the usable window. Tier 3 gives physical-source book navigation for ARCHITECTURE; packet triad rows can additionally inline relevant sections within the same share. Each non-inlined document has a manifest disposition and navigation pointer. Retrieving rows can follow those pointers; tool-less packet rows receive an index of material not delivered, never an instruction to call unavailable tools. A whole document with the same source identity is credited as `delivered_inline` without a duplicate read.
#### Guaranteed-fit ladder
The triad packet retains its one-pass fit sequence in `review_admission.fit_triad_prompt`: shorten the optional evidence projection, replace duplicated post-change snapshots with a touched-path note, then re-render the same pinned subject at `-U0` while preserving all file/hunk identities and added/removed lines. The cold-density rung above can precede this sequence once. Advisory briefs and triad packets share the disclosed span-only release-carrier cut; triad additionally omits governance bodies byte-identical to its inline prefix, while managed subjects retain full texts. Every omission is named, and an irreducible oversized packet is a typed refusal. Scope uses its paged exact subject instead; plan review sizes its typed specification and declared evidence per slot, retaining the constitutional material self-modification plans require.
### Planning, deep review, reflection, memory
Plan review, task acceptance, commit review, and deep self-review answer different questions and never inherit one another's authority: planning judges a proposed approach before implementation, task acceptance the delivered objective, commit review a staged self-change, deep self-review the whole system. Post-task reflection and memory persistence learn from execution but approve none of those boundaries.
#### Plan construction and review
`plan_task` reviews an INTENTION before the work starts — the same organ for code, research, deliverables, and actions (BIBLE P3). The obligation is constitutional and the finalization gate is structural (`owner_hurry.force_plan_decision`); the order in which a task asks, explores and plans is the mind's judgment, and no prompt choreographs it. The envelope carries the goal, the plan prose, and a typed domain-neutral SPEC: `in_scope`, `non_goals`, `acceptance_claims`, `invariants`, `decisions`, `deferred`, `affected_paths` (REQUIRED — the files the work will CHANGE, `[]` when none), `affected_resources` (the same question in words: systems, services, projects, people, never resolved as paths) and `evidence`. A spec without `affected_paths` is refused before any dispatch with the typed `PLAN_RESOURCE_FORM_REQUIRED`, which records nothing. `ouroboros/tools/plan_spec.py` validates and normalizes the full operative content without shortening or dropping anything, hashes it, and mints the only valid `breaks` targets (`goal`, `claim_N`, `invariant_N`, `decision_N`, `deferred_N`) positionally, because a caller-chosen id could shadow another target and corrupt what a blocking finding `breaks`. Governance documents always come from the system repository; declared targets and evidence resolve against `active_repo_dir_for(ctx)`, and a path escaping the active subject or an unreadable root is a named omission, never a silent gap.
ONE structural fact tiers the governance pack: `constitutional` is true iff a declared `affected_paths` locator resolves under the Ouroboros system repository, whether the file exists yet (creating `ouroboros/new_module.py` IS self-modification). Nothing else buys the pack: `affected_resources` is prose, and an `evidence` locator is something to LOOK AT — reading a repository file is not changing it — so system-repo reads are only named in the disclosure note. A constitutional plan carries BIBLE.md and ARCHITECTURE.md in full (inline on an `api_chat` row; mandatory full reads on a retrieving `agent_session` row); that packet is not tiered, so assembling it without either is a typed failure, never a disclosure. Every other plan carries the runtime heading-derived navigation maps (`context_layout.generate_doc_nav_map`, never a copy) plus resolvable pointers. There is no plan-kind taxonomy, no agent-declared `plan_class`, no planning scouts, and no assembled repository pack for a plan.
The ONE caller-facing strength axis is the envelope's optional `reviewer_effort`: the panel's effort for THIS order, placed on the default rung of each row's ladder (`reviewer_slot_config.row_effort`: an explicit per-row effort, then a compound Cursor/Agy route slug, then the declaration, then the owner's `OUROBOROS_EFFORT_REVIEW`) and passed as an ARGUMENT of `plan_review_slots` only, so every other review surface keeps reading the untouched rows. Effort is roster identity: a different declaration re-dispatches a paid panel within `OUROBOROS_REVIEW_MAX_CYCLES`, the same one replays free, and it is recorded per wave (`reviewer_effort`; `declared_effort` in the last-execution projection, apart from the row's saved effort). Against BIBLE P1 this is the review panel's strength for one order — the owner's setting is the default, Ouroboros may order stronger or weaker per envelope — never the core's own model, effort or horizon. Disclosed residual: a cheap panel's closed GREEN is earned authority for that envelope, and a later stronger order does not reopen it.
Declared evidence is resolved by `ouroboros/tools/plan_evidence.py` against exactly two roots — the active workspace and the system repository — with the shared sensitive-name policy on locator and target and the disclosed bounds `EVIDENCE_PER_ITEM_BYTES` / `EVIDENCE_TOTAL_BYTES`: every refused, missing, truncated, oversized, binary or URL locator becomes a typed omission row in the manifest (the host never fetches a URL); a locator may carry an exact range (`::lines=A-B`, `::bytes=A-B`, `::tail=N`, `::symbol=Name` for `.py` sources), and an oversized source is attached head-first with the cut named. A `need_evidence` locator a reviewer names is attached by the host on the next cycle through the same policy; what it cannot attach becomes a `[reviewer-requested]` omission row the panel is dispatched with, never a reason to run no reviewer. The manifest hash joins the spec hash and `constitutional` in the wave fingerprint, so changing what the reviewers can see changes the identity of the review — an attached locator makes the next envelope a new fingerprint, never an idempotent replay — while the root exploration log stays outside that identity. The evidence continuation uses a fresh full-packet dispatch only when no exact artifact reference exists, disclosed per slot as a `capability_delta`; an unreadable referenced artifact instead fails closed with `plan_review_exact_artifact_unavailable` and never mints replacement authority.
Own-room dialogue is automatic evidence: `ouroboros/dialogue_evidence.py` reads the complete retained room through `Memory.read_chat_generations` (rotated and consolidated history included) under the shared `project_dialogue.room_membership` predicate, plus the progress stream and addressed `owner_mailbox` entries — both speakers, options, quiz recommendations and answers, attachment names (the canonical manifest `label`), typed outcomes and sender/relay provenance. The winning quiz answer is written to canonical chat history as its structured block (after `record_answered`; an answered `quiz_closed` reply heals it), so eviction and mailbox cleanup never erase it and History shows it in the original card without a second bubble. Missing history, room binding and unstable capture remain explicit coverage gaps; acceptance freshness stays with the acceptance directives ledger, and planning consumes no separate directive section.
`ouroboros/tools/plan_dialogue.py` captures the redacted full room as an immutable `task_source` in existing artifact custody, bound to the author-request identity (goal, prose, spec, declared evidence, the constitutional fact). Identical author inputs and free collection reuse the recorded snapshot — progress, panel-settled messages and other room growth cannot buy a new panel — while a genuinely changed author request (a changed plan or an explicit evidence request) captures the current room. The wave fingerprint includes the snapshot digest; `dialogue_source_ref` and the author identity survive hot-index compaction and child promotion beside `spec_source_ref`; every result names the original capture and excludes later messages. `chat:<id>` and `chat:<id>@<sha256>` locators address the room and its recorded snapshots through `plan_evidence.resolve_evidence` with the same selectors, redaction and omission accounting; runtime files do not become general path evidence, and related Main/lineage rooms (bound siblings included) remain labelled pointers, never unsolicited inline neighbours.
`ouroboros/tools/plan_packet.py` keeps the current objective, spec and prose complete, and own dialogue bypasses the ordinary per-item/total evidence bounds. Delivery follows the per-delivery source contract: Packet/API slots receive all of it while their calibrated route window, governance and output reserves permit, otherwise the newest text plus the exact immutable omitted byte range; native retrieving slots fit under `review_native_transcript_bound` and the mandatory-reading declaration — native sizing measures the complete first request (`native_first_send_chars`, schemas and wrapper included) — with full artifact access through `native_data_root`; delegated sessions receive a mandatory full-read instruction and the complete redacted artifact outside `active_root`, their harness owns context selection (no host-invented numerical window or 1M claim), and an actual overflow is disclosed as exact coverage/omissions. File availability never attests that a reviewer read it all: plan review declares no required-source manifest to fold, so coverage stays declared/unobserved and a retrieving reviewer is recorded `host_file_read_attestation: unobserved`.
Slots come from `reviewer_slot_config` over the `review_execution._review_route_executor` seam. Reviewers return ONLY a typed findings array — `blocking` with a `breaks` id, `note`, or `need_evidence` with a locator the host attaches (or a spec id in `breaks` as a question to the author, who answers, escalates or defers it openly in the disposition); the HOST validates membership, demotes invalid findings with disclosure, keeps failed slots in the quorum denominator, and computes the aggregate — `GREEN`, `REVIEW_REQUIRED`, `REVISE_PLAN`, or the honest `DEGRADED` — through `config.adaptive_quorum`. No reviewer emits GREEN as authority. Premise criticism and simpler alternatives are optional brainstorming notes whose adoption is Ouroboros's; no mandatory competing plan, and repetition alone cannot promote advice to a blocker.
`plan_review_state` v2 inside the root task result is the bounded durable index: each wave records a hash-bound `task_source` reference to the full operative spec, the hashes, `constitutional`, validated findings, aggregate, dispositions and paid state (full waves, specs and manifests: `plan_review_artifacts.py`; review evidence stays redacted). State reads share the task-result writer's lock without rewriting JSON; source resolution starts after lock release. Current-authority readers restore the full spec before comparing plans or binding its `acceptance_claims` (`contracts/task_contract.effective_acceptance_claims`); compaction and child promotion retain both references, and an unavailable recorded source is a typed infrastructure failure, never empty claims or a new unpaid cycle. A v1 record is read-only (`legacy_v1_projection`; an OPEN v1 wave projects `legacy_open_requires_resubmission`, never auto-closed). A corrected full goal/plan/spec can be selected by an explicit `review_disposition.author_action` with author disposition and the referenced review fingerprint. `current_attempt.author_subject` holds its exact source and separate critic reference; `plan_review_artifacts.current_author_plan` restores it without a new wave, GREEN or paid cycle. Advisory finish can select its claims as `acceptance_claims_source=author_plan`, after ingress claims and through the existing receipt-support checks. Blocking can save and stop but cannot use the plan as approved; `closed_plan_review_wave` remains literal critic-closed authority. The gate and terminal disclosure follow the author's explicit `review_fingerprint` for historical critic outcome and pending custody as EVIDENCE INTO A GAP — a revised plan with no wave of its own stops reading `unavailable` — without granting that wave's closure to the revised plan: the enrichment fills only an ABSENT outcome and never moves the attempt's own lifecycle, because overwriting the status with `open` took the release back off an exhausted cap and off a degraded rail (an author permitted only to finish honestly was pushed back into a cycle it cannot buy), forcing `closed=False` contradicted `closed_plan_review_wave`, which still bound that wave, and replacing a current wave's own aggregate would let a historical verdict stand in for the one this plan earned. The projection therefore carries `historical_critic`, and both of `owner_hurry.py`'s renderers read it: the terminal disclosure labels that outcome as the earlier plan's instead of showing it as this plan's own, and the blocking reminder answers it with its own branch, because its outcome-keyed advice would otherwise send the agent to dispose a wave that cannot approve the revised bytes. A rail that forced finalization keeps its own sentence with the author clause appended, since returning the author decision alone hid `the cap is spent; the task ends blocked`; an author STOP carries `allow=True` under every enforcement, so it is excluded from the advisory-continuation sentence rather than described as work that proceeded — and because the projection's stop REPLACES the status, a rail that fired under a stop is not separately named, the stop being the fact that ended the task. The open-wave status reads `cycles_exhausted` from the attempt and the wave, because an author finish at a spent cap stamps only the attempt (`_apply_author_subject`, the one writer with no in-flight guard) and the gate otherwise held a task that may only finish honestly; that read is guarded by `custody_pending`, since a paid slot still working could yet close the wave and releasing then would trade an honest hold for a premature blocked terminal. Missing criticism stays unavailable, and author finish/stop stays separate. Disclosed gap: `web/modules/review_presentation.js::planReviewGroupFromTaskDetail` still selects the critic by the author's `review_fingerprint` and renders that aggregate as the group's own verdict, so the CARD can show an earlier plan's GREEN where the terminal disclosure says the revised plan has no verdict; the same projector also reads `rail_degraded` from the attempt but `cycles_exhausted` only from the wave, so an author finish at a spent cap on its own wave now releases the gate and reads blocked in the terminal while the card still shows the wave's aggregate. Both are one presentation gap needing its own change with browser verification, not a backend fact. Old-wave collection never replaces the selected author source. The first exact wave also stores the request policy and per-slot prepared inputs/fit sizes, which collection reuses unchanged — `native_mandatory_read_chars`, the review contract hash and the physical-operation binding never move with live exploration or owner clarification — and a historical wave without them is a typed source-unavailable result, never guessed values.
A fresh dispatch returns control at the dispatch barrier (`ReviewRequest.drain_deadline`, drain window 0): the wave is recorded open with typed `pending_dispatch` rows for the workers still running, whose custody stays with the recorded wave in the process (`review_custody`). A paid actor still physically in flight keeps the wave open as `DEGRADED` with `review_late_result_pending` even when the settled rows meet the arithmetic quorum, so a late blocking result cannot arrive after a false GREEN. When the last released slot settles, the settlement thread writes ONE system frame into the task's mailbox (`owner_mailbox.write_task_message`, provenance `system`) so the mind wakes exactly as for a child result; it never closes or aggregates the wave — collection is the sole wave writer (`plan_review_collect`). Task acceptance's twin announcements (its own quorum, completion) carry each reviewer's own verdict; the aggregate still comes only from collection.
Collection is the existing `review_disposition` mode addressed at the recorded wave (`items` may be `[]`): it reconciles process-local custody with drain window 0 over the wave's own recorded inputs — never re-reading the evidence, never changing the wave's identity — closes or advances the wave, pays the cycle once from the settled rows, and never waits; waiting longer is only the identical envelope resubmitted. A reconciliation re-records the packet that was physically DISPATCHED (each slot's `request_messages`, `session_task` for a retrieving row), never a rebuilt one, so a directive that arrived after the dispatch reaches the reviewers only in the next physically sent packet. A NEW review envelope over an in-flight wave first collects what has settled at $0, then supersedes it (reconcile-before-supersede) instead of being refused; late settlers attach as historical supplements, accounted once without another send. A wave with custody pending is never compacted out of the hot index.
An answer that has not arrived is a gap, never a failure or a verdict. The STORED wave is the fail-closed floor: an unanswered slot stays an `ok=False` row no quorum counts, and the open `DEGRADED` + `custody_pending` pair holds every gate with no fifth aggregate for older builds to refuse. Renderers read `plan_review_runtime.plan_wave_slot_census` before wording a slot: `awaiting` is `operation_state=pending_dispatch` only; `unresolved` is `in_flight`/`custody_lost`, named as unresolved on the line and by raw state in Reviews, never called waiting (`late_result_pending` is true for both); `uncollected` has a settled supplement not collected yet; `not_dispatched` is a $0 refusal, not a failure. Verdict and failure wording and the task's `degraded` stamp belong to collected slots only: a finish over an only-awaited wave or panel keeps its disclosure and records `awaiting` (`execution.plan_review`, review axis; `owner_hurry.plan_wave_only_awaited`, `_outcome_receipts.review_runs_only_awaited`). `owner_hurry.force_plan_decision` makes ONE free collection of the current pending wave in every enforcement mode, hurry too, then projects the returned state (`plan_review_collect.collect_before_gate`). Advisory and hurry keep their local release; collection updates disclosed facts, never permission to proceed. Context health reads the canonical wave, never collecting: `PLAN REVIEW WAVE OPEN` names its fingerprint, pending-slot count and snapshot timestamp (not a physical start, not a liveness lease) and disappears once collection clears custody.
Closure follows the finding class (`plan_spec.closure_after_disposition`) at initial synthesis and at later dispositions: GREEN and note-only REVIEW_REQUIRED close immediately in either enforcement mode; outstanding `need_evidence` closes through a disposition-only `plan_task` call (no model call, no cost: accept = answered, reject, defer = deferred openly; an answer reaches the reviewers on the next paid cycle, a revised envelope supersedes the wave with its open requests); a below-quorum blocking finding stays open; REVISE_PLAN is never closed by disposition — a subsequent paid delta review may evaluate a changed spec or a justified rejection when another cycle is available. A closed note-only wave still accepts voluntary `review_disposition` annotations through the same writer, spec/verdict and paid-cycle count fixed and the immutable predecessor retained: optional reasoning history, not a new plan gate, never another panel.
Paid cycles are bounded by the shared `OUROBOROS_REVIEW_MAX_CYCLES`; `cycles_paid` and the typed `cycles_exhausted` state count PAID cycles only. A wave is paid iff at least one reviewer slot was physically dispatched, and the proof is a settled row: a slot released at the barrier is $0 until its row proves the send, so the cycle is paid at collection, never at the barrier, and a nothing-dispatched wave of typed $0 skip rows stays unpaid and never replaces a paid predecessor. A barrier-dispatched panel is nevertheless committed money: while its wave is custody-pending and the cap has no room for another panel, a revised envelope first collects EVERY other custody-pending wave at $0 (`plan_review_collect.collect_before_supersede`) and, if the cap still has no room, is HELD with the typed `PLAN_REVIEW_IN_FLIGHT` refusal (`plan_review_collect.in_flight_hold`) so the current wave stays collectible (the per-case cap arithmetic: `plan_review_collect.py` docstrings). An identical envelope replays the recorded wave free, with one exception: an open wave whose blocking findings all carry valid reject dispositions may dispatch exactly one subsequent paid delta panel when another cycle is available.
Before fan-out the engine captures one panel health snapshot (`subagents.route_health`, route-level evidence): a slot with positive structural evidence of a spent lane becomes a $0 typed skip row that stays in the denominator; unknown health dispatches (fail-open). The wave records the health epoch and reviewer-roster fingerprint, and a recorded DEGRADED wave replays free only under an identical envelope, matching epoch and unchanged roster — otherwise another paid panel needs remaining cycle capacity (replay/epoch casuistry: `plan_review.py` docstrings). When the wave's own typed rows prove the quorum structurally unreachable, it carries `quorum_unreachable` plus the earliest recorded reset, and under blocking enforcement the finalization gate RELEASES while the review stays open: the agent may finalize `blocked_with_evidence` (reason `plan_review_quorum_unreachable`), wait through a one-shot `schedule_followup`, or ask the owner — the host adds facts only, never an answer template.
Under blocking an open wave otherwise holds implementation and an exhausted cap escalates as typed `review_cycles_exhausted` (free disposition and exact pending custody keep their rules); under advisory the agent may proceed with the wave open, after typed owner-visible `plan_review_advisory_open` events (one per announced outcome: the dispatch snapshot, then each settled failure with its code and reported cause) and with the open wave recorded in the task's state and result; the model-facing guidance, the `plan_task` description and the cap head say that the agent's own answer is where the open wave is stated, never that a host narrator will state it for it. A forced rail hands its ONE model call the same limitations as typed `key=value` facts inside `_prepare_forced_prompt`, so they are priced by the existing wrap-up reservation probe and no no-call rail gains a call. A hold the agent was told about that the owner's hurry then released reaches it once before its last word; every other gate transition already arrives through the agent's own `plan_task` result, the settled-wave task message, or that forced prompt. Unavailability, invalid state, budget refusal and deadline rails stay typed non-authoritative attempts, never substitutes for GREEN, and a reviewer request the host cannot attach is never refused before dispatch: the only $0 exits are those attempts, `not_dispatched` slot rows and the `pending_dispatch` rows of a barrier-released wave. Every disclosure states the wave's CURRENT state, names the late-result fact while a slot can still settle, never calls an open wave ended, and renders a rail with no recorded reason as absence, not an internal token.
#### Deep self-review
The retained report (`memory/deep_review.md`) is historical evidence, not a current-tree verdict: it records its actual `generated_at` time in the provenance header, and `source_revision=unknown` is explicit because these deliveries capture no frozen reviewed Git revision — neither a live HEAD nor a file mtime can fill that gap. Context quotes a disclosed 8,000-character excerpt, labels unknown legacy provenance, and leaves the stored report intact.
A task with `type=deep_self_review` bypasses the ordinary tool loop and calls `deep_self_review.run_deep_self_review` on its configured `deep_review` row. An absent row is synthesized as API from `OUROBOROS_MODEL_DEEP_SELF_REVIEW`; the tool and agent read the effective row. Every row retrieves, with two deliveries under `ReviewRequest(surface="deep_self_review")` and the shared execution seam:
- **Native inspection**: every `api_chat` row, bare or configured-subagent, runs `NativeToolRoundReviewExecutor` with the repository root and real runtime data root. The task carries the role/method, BIBLE.md and standing disclosures inline through the shared governance tiers, relevant rules within the inline share and reference-book navigation. Bounds are the route-sized working view, owner deadline and logical window, and paid ledger; exhaustion returns the retained draft with its actual incomplete end.
- **Delegated session**: `AgentSessionReviewExecutor` receives the same task and report contract (markdown, not an output schema), and reads through its harness. Inline delivery is recorded separately from the weaker or unobserved tool-reading provenance.
Both deliveries inline the seven-file memory whitelist byte-exact: identity, scratchpad, registry, WORLD, full knowledge index, patterns and improvement backlog. Every entry carries `inlined`, `missing`, `empty`, `oversized` or `read_error` in the task, `deep_review_memory` usage and the report header. Memory is the subject of this review and is supplied directly, not left to discovery. Matching inline governance is `delivered_inline`; additional native coverage comes from executed receipts matched by opened root/path and source identity, with `read`, `partial(fraction)`, `missing` and `unobserved` stated honestly. Reading gaps remain diagnostics and never invalidate a finished report.
Every delivered report has a sanitized host provenance comment and human line naming delivery, model, memory and omissions, coverage, incomplete status and attestation; native rows additionally carry rounds, calls, receipts, end reason, transcript and landing facts. `incomplete` reflects the delivery's actual interruption, independently of diagnostic reading gaps; session completeness stays unobserved. Each outcome (responded, empty or exception) records the row's last execution and memory fact, and `persist_call` retains prompt/response custody. Typed `deep_self_review_unavailable` or `deep_self_review_error` failures return `execution_status=infra_failed`, reach task/error records and leave the previous `memory/deep_review.md` intact; `BudgetExceeded` propagates to the agent's budget-pause rail.
Availability follows `deep_review_route`, never a window floor. An API row needs credentials for its actual routed model: a stored `openai/<slug>` can resolve to direct `openai::<slug>` when direct credentials exist and `OPENAI_BASE_URL` is unset, while a product `-pro` slug resolves to the direct default. Session rows require a healthy delegated route. These reviewers have no mutating tools and run no plan, acceptance or commit reviewers: the report is diagnostic memory, not implementation or publication authority. Retrieval enables a targeted whole-system survey across successive views without confusing an assembled packet with measured consideration of the whole system; exact sources remain accessible, and the record distinguishes supplied inline content, observed reads and unknown coverage.
#### Post-task reflection
Typed root post-task triggers decide whether a run warrants Experience Review. `reflection.generate_reflection` sends the Light route one open prompt with the EXACT initial text (never a prefix) and its host-recorded `task_inputs.run_origin` beside it (provenance, never by itself the accepted requirement) plus tool-use, error, review and child projections and the same frozen non-final cost snapshot the task summary uses; it runs outside the tool loop, records its own usage, and its failure never erases the delivered result or changes a review verdict. Its execution trace is the ALL-CALLS listing (`build_trace_summary(all_calls=True)`): every call in order with every argument, identical consecutive calls folded into one `×N` row with their rounds, the first line of a failed or repeated call's result, and one header count of rounds whose every call was non-ok — no positional window and no literal cut, because the consolidation seam fits the call to the Light route whenever that route's window is known (an unknown window sends the prompt unchecked — the accepted residual of enlarging it); the STORED `trace_summary` (task card, parents, children) stays the bounded two-argument preview. The trace row carries the `round_id` of the model round that issued the call (absent when unknown), when the listing really cut an argument value or a failed/repeated call's answer, the redacted per-call record is retained through `retain_memory_source` and named in the prompt as OPTIONAL reading (never a required source), and its claim states the STORED bounds on both axes, not completeness: an argument already passed `sanitize_tool_args_for_log` (an oversized value carries a marker with its length and sha) and a result is the stored actor-visible cap — more than the listing, which shows only the first line of a failed or repeated answer — a partial one naming its own `FULL_RESULT_SOURCE_JSON` or `FULL_RESULT_SOURCE_UNAVAILABLE`, with a call's recorded manifest named only when it has one — claiming results "in full" or an unconditional manifest overstated a cognitive artifact. Cut detection reads the ONE shared marker list (`artifacts.SANITIZER_OMISSION_MARKERS`), because a width test over already-sanitized args measured the widest argument in the task as a small one and retained nothing at all, while a hand-rolled subset missed the `_repr` and `_error` shapes whose arguments survive only in the call blob. Unavailable source retention is disclosed, error details group by full redacted content before display clipping, and post-task synthesis — the reflection, its Pattern Register update and the episodic summary — thinks at the owner's Task / Chat effort (`settings_scales.resolve_effort("task")`), never a literal. Admission to the Pattern Register is typed, not a word scan: it opens on a call the loop recorded as errored (its stamped `tool_result_code`, or the recorded status for a legacy row), on a producer fact naming a failure the ok status cannot carry (a preserved commit whose post-commit tests failed publishes `post_commit_tests`), on typed codes already stored with an entry, or on a genuinely FAILED child — a cancelled, soft-landed best-effort or degraded child is not a failure. Those failed-child classes reach the root through the child evidence the synthesis walk already collects and make the run error-bearing for both the trigger and the prompt's error details: children do not reflect, so a short clean root that delegated the work is the only place its child's failure can be learned from at all. Deliberately not "a reason code exists", which would open the register on every terminal.
A reflection lands where it durably belongs: a non-project root appends the full entry to the canonical `logs/task_reflections.jsonl`; a project-scoped root appends the full entry to its project drive and the canonical log receives only a bounded pointer row — full project text never enters the canonical log, which feeds future global context. A project-bound task's context includes a bounded labeled tail of its own project's reflections; the headless mirror drive of a split root is never the reflection home, and the Pattern Register update stays canonical in both cases. Every entry carries task identity, evidence, lessons, backlog candidates, and validated memory actions. `MEMORY_ACTIONS_JSON` permits only `scratchpad_append`, `knowledge_write`, and `identity_update_candidate`, at bounded count and size. `apply_memory_actions` routes accepted actions through provenance-preserving memory and knowledge APIs. An `identity_update_candidate` is recorded in the scratchpad for review and is never auto-written to `identity.md`. For a project-scoped task, reflection applies knowledge actions only (the project store by default, explicit global allowed); its scratchpad and identity-candidate actions are skipped because this automatic Light pass lacks the conversation's full view — the conversation writes identity and scratchpad from any room through its own tools. Reflection may propose a future campaign or backlog item, but it cannot enqueue, review, commit, or enable one.
The Pattern Register writer (`reflection._update_patterns`) REPLACES the whole document, so every decision input it reads is complete — the full current register, the exact initial text (`goal_exact`, beside the bounded `goal` display) and its `run_origin` and the whole reflection text; a prefix of a decision input can never authorize the rewrite, because a clipped clause can record the inverse of what the reflection concluded. `patterns` is a reserved global-only topic alongside `improvement-backlog` and `overview`: whichever room writes it, the register has ONE home on the canonical drive, the only one its writer and its readers (context assembly, deep self-review, the headless copy) address. Re-read, exact compare, history append and atomic replacement are one critical section outside the Light call; a register that moved under a losing writer is preserved and that task's learning is DROPPED with a warning naming the task, never retried against a source it did not decide from. A row's count is bumped once per observed episode, so two roots of one owner request bump it twice — the number counts episodes, not distinct requests.
The advisory rows a reflection or summary reads are ATTRIBUTED. Advisory runs are scoped by repository and several tasks legitimately review one checkout, so `collect_review_evidence` renders as `recent_advisory_runs` only the rows this task owns plus legacy rows carrying no owner (which stay unknown, never re-attributed); another task's rows reach the prompt under their own heading naming the owning task ids, and every row carries its owning task, attempt, phase and complete snapshot hash. Repository readiness (`current_repo`, open obligations, commit-readiness debt, the exact-snapshot match) stays repository-scoped, because that is what "can this checkout be committed" means. The split is keyed on ROW IDENTITY, never on a scope key: an empty repository key widens the candidate list to every advisory run on the drive, and the installation's history is not one task's record.
Only roots synthesize; `root_phase_checkpoint` makes paid synthesis at-most-once across restart, while children contribute evidence. Durable-result persistence owes the answer as `final:<tid>:<digest>` in `supervisor/terminal_delivery.py`'s bounded outbox (§5; normal/cancel/reap). `send_message` delivers immediately; the retained buffered copy shares its ID for durable dedupe. Replays use bounded backoff. Exhaustion/eviction preserves full text on disk, emits `terminal_delivery_exhausted` and a chat notice; external delivery remains at-least-once. Buffered `task_done` stays last to retain the slot/child drive during synthesis; a hung-synthesis reap need not lose the delivered answer. Project roots keep early answers in Project. Their canonical row and deferred Main mirror use `terminal_projection.settle_terminal_projection` via task-done/checkpoint/startup/maintenance; §3 "Main rows and host-stamped card rows" owns readiness, retirement and limits.
Synthesis receives a sealed final package from the durable result — the submitted final text, its artifact manifest and completion_observations. Full redacted action observations live in the canonical artifact store (`task.budget_drive_root or drive_root`), in the write-once `source_handles/context_checkpoints` store with verified `task_source` refs, before compact publication and outside deliverables and inferred readiness; their native reader `get_task_result(include_completion_source=true)` returns complete length/hash first, then explicit `source_start_char`/`source_end_char` ranges (`artifacts.text_source_range_projection`, the shared work-order range contract), with bytes, kind, path containment and SHA checked before any excerpt. Packet-only summary/reflection receive per-send-tool counts, each family's latest recorded return, and task-related skill readiness with coverage; full-source references are for later readers, not evidence the synthesizer has read. Positive observed facts correct error-trace impressions, while tool success does not prove owner receipt, empty material does not prove absence, and skill readiness does not attribute an owner's action to the task. Before context cleanup, `agent_task_pipeline.emit_task_results` also freezes `review_evidence.task_inputs` through `post_task_synthesis.capture_task_inputs`: `run_origin`, the existing task-local owner corpus, intact question/answer provenance and the canonical split-root verification-receipt union. Summary and reflection receive the same complete redacted content through `reflection.task_inputs_prompt_section`, separate from bounded trace/review excerpts. A zero return code is positive evidence; an unrelated later pass cannot resolve another check's failure. Peer proposals stay attributed, and unavailable input is not evidence that approval or verification never existed. Recovery uses these stored observations and inputs, not a later conversation. Summary trace pointers name the task and existing archive-aware reader (`ouroboros tasks watch <task_id> --jsonl`), not guessed flat log paths.
Pooled workers retain their slot until root post-task synthesis settles, for API-only and subscription tasks alike; early final-answer delivery keeps the response independent from that queue timing. Ordinary native post-work, including an inline Presence turn after its durable result is returned to the adapter, stays on its registered actor thread without a pooled worker slot: its `TaskModelWait` owner remains available through `POST_TASK_SYNTHESIS_INFLIGHT` after ordinary dialogue admission closes, detached server post-work binds its own live owner in the same registry, and the task mailbox stays available until the terminal post-task checkpoint. An open phase remains finalizing rather than appearing completed. Typed quota/auth waits resume only the unsettled call, stop or unknown outcomes degrade the phase without repeating finished stages, and restart recovery degrades an indeterminate `running` phase rather than replaying a possibly paid request.
#### Project registry and lease
A project is a focused working room, not an isolated sub-mind (§6 Durable memory and project focus). The projects registry (`projects_registry.py`, `data/state/projects.json`) owns immutable project identity, canonical chat id, optional working directory, lifecycle/tombstone state (`active|deleting|tombstoned`), routing generation and activity revision; admission persists the resolved project id in the task itself. An id minted from a DISPLAY name collapses dash runs and carries a short deterministic suffix when the name held characters the slug could not keep, so two different non-Latin names cannot share one project; the normalizer is unchanged, so existing ids are never re-slugged and stay reachable by their explicit id. `project_lease.py` serializes assignment of pooled roots by Project while allowing their own subagent trees; it is not a physical-folder lock and does not withhold tools from ordinary conversation. Binding/history files support routing and presentation, not the lease. Delete closes routing, cancels/quiesces the tree, and tombstones only after settlement, preserving everything for recovery.
#### Project binding by task and by origin
The durable binding is the SINGLE truth about a task's project: the in-task scope guard and `project_facts.resolve_project_id` read it FIRST, ahead of `task["project_id"]` and a worker's in-memory `ctx.project_id`, which are copies a mid-run conversion never reaches — a guard reading only the copy is how a task already bound to one project can mint a second, empty one. The same store also answers BY ORIGIN (`projects_registry.project_id_for_origin`, keyed on the ingress-captured `(chat_id, client_message_id)` stored in every binding's `source_ref`), because one owner message spawns several task ids — the turn that received it, the root it promoted, the timeout retry that replaced that root — and the project its WORK has belongs to all of them: a promote with no explicit target inherits the promoter's binding, then the origin's, before the in-memory scope copy, and its admission resolves the origin once more. A timeout retry is bound by the reaper inside its retry admission transaction (`worker_promotion.bind_retry_to_origin_project`, called by `task_reaper._run_retry_admission_transaction`): the predecessor's binding, then its origin, answer, and the retry reuses the stored origin by value, so retried work keeps its room instead of arriving in Main as a second convertible unit. That bind happens only after cancellation has lost the admission boundary, because a binding is immutable and a bound-but-never-admitted retry id would answer `project_id_for_task` forever; a retry suppressed by a cancelled or already-terminal root is never bound, and a refused bind or unreadable store leaves the retry unbound and discloses `project_binding_failed` rather than holding up the retry.
Every IMPLICIT claim — the UI conversion, that admission, the reaper's retry admission and the in-task `ensure_project_scope` — holds one process-local claim lock (`projects_registry.origin_claim_lock`) from the read of which Project the origin already names through to its own durable bind, while a naming model call stays OUTSIDE it, so two cards of one message clicked inside a naming window yield one Project and two bindings. Explicit `project_name`/`project_id`/`route_to_project` remain the model's choice (P13): a sibling that names a different room gets it, and the message's existing Project is never renamed from a task that does not belong to it. Legacy state where one origin names several active Projects resolves to the one whose task is still live, else the latest binding, disclosed as `project_origin_ambiguous`; continuations that carry no owner origin (`schedule_followup`, direct auto-resume, Presence promotion) stay per-task. An unreadable bindings store is disclosed once and read as "no binding" on every path: the refusal is reserved for the measured case, a readable binding to another project, and on BOTH conversion paths it precedes every side effect — no project row, lease mark, broadcast or announcement survives it: a UI conversion that named a DIFFERENT project id answers 409 naming the bound project by id and display name, the one-click conversion adopts that project instead, and a bind refused after the lease mark restores the lane with the same 409, so no lane keeps a project the binding does not name. One residual survives: a conflicting bind landing between the final re-read and the bind leaves behind the empty project row the request had already created, because the registry deliberately has no primitive that removes a row (delete tombstones the id permanently, which would cost the owner that id and display name forever); the row holds no task and no binding, and the owner deletes it like any other project. The lease mark is fill-only, except that a conversion which already holds the binding it is about to write moves the lane onto that binding.
#### In-task project scoping
`ensure_project_scope` (`tools/control_delegation.py`) can create or bind the current root to one project mid-execution: it persists the durable registry binding first, then marks the live queue/lease surface under the queue lock so the lease recognizes the running task as a lane occupant; it is idempotent for the same project, and a task BOUND elsewhere is renamed, not re-scoped — its scope call carries the requested display name to the project it already belongs to and creates nothing. The act rides the same receipt rail as the other routing verbs (§6 Task lifecycle) under its own synthetic `agent-steer:<token>` id: the tool scopes itself in memory at once (`journal_write` targets the project meanwhile) but its RESULT states only the durable outcome the supervisor recorded — `delivered` with the binding it re-reads, a typed `rejected` (bound elsewhere with the rename outcome, a refused bind, a registration failure), or `unconfirmed` when no receipt landed within the bounded wait; `persisted` on a routing receipt is true only when a row was written. A late bind after a lost acknowledgement is discoverable from the durable binding, which the next same-project call reads as already scoped. A planning obligation stays with the task on ensure. A project-SCOPED but unbound run keeps the older refusal, since there is no durable project to rename, and a child still cannot escape the inherited scope.
#### Durable memory and project focus
`context.py` assembles static governance, semi-stable memory, and dynamic task evidence without treating truncation as forgetting; the recent-activity sections are each task's OWN newest rows (progress 50 rendered; tools 20 selected, 10 rendered and 20 scanned for review markers; events 200 counted by type) through the bounded reader `jsonl_tail.py` (`Memory.read_task_recent`: a doubling live tail plus at most three newest archives), never a global tail filtered afterwards (issue #131), and their header's coverage line names the rows, the window and any unopened archives while `read_file` pages the rest; a subagent child gets the same three windows beside its `## Working sources` block, its tools and events read from its own execution drive (its worker rows; host-side rows such as waits stay in the canonical log, as the header says) and progress from the canonical log; the Development context matrix and `context_layout.py` own which reference form is resident. When the rendered scratchpad exceeds `SCRATCHPAD_SECTION_BUDGET_CHARS`, `context.py` keeps the newest whole blocks that fit and drops the oldest behind an in-band gap marker naming `memory/scratchpad.md` as the live source; no block is retired by a context build, and scratchpad replacement keeps its explicit summary and source-journal provenance.
`consolidator.py` publishes a block and advances its generation-aware cursor only after complete room draft and correction; a missing generation appends `[MEMORY GAP]` instead of resetting. `context_fit` measures Light against fresh route/account capacity and calibrated density; `llm_local` owns the local output reserve, and absent evidence stays unknown. `room_consolidation.py` processes each room separately and assembles sections deterministically. Both knowledge stages receive the entire current note and source episode; the corrector's complete read, not the draft's, binds revised entries. Older episodes cannot negate newer facts; model judgment governs supported corrections. Source range reads and CAS preserve old/new history. Startup compaction is a no-op; earlier digests cannot be reversed. A failed correction withholds the chunk and cursor.
Oversized sources split without clipping, including within an entry. `consolidation_retry` records source hash and a smaller same-route bound, invalidated by source/route/capacity/reserve changes. Era compression regroups rooms deterministically; failed eras retain blocks and legacy provenance stays unknown. `last_consolidation_error` clears after a failure-free advance. `pending_knowledge_nominations` records each source entry BEFORE note publication; an unrelated successful batch cannot remove one. Legacy `last_unpublished_nominations` persists. Health shows three distinct abbreviated source+position IDs and omitted count; full proposals live in `knowledge_history.jsonl`. Unreadable meta preserves debts and warns in Health; Nano pressure records no-progress instead of aborting Main. Invalid legacy receipts warn separately. Old digests and debts need explicit resolution. Spend remains nullable; model-control errors follow `propagate_model_error`.
Consolidation labels every chronological source message through `dialogue_provenance.RoomLabelResolver`, using the actual `chat_id`, never lineage `project_id`; one read-only registry snapshot supplies the window. Main is named only for the actual Main id, a resolved project uses its current registry name and stable chat id, and missing, unknown or ambiguous rooms stay explicit. Ephemeral formatter offsets carry the original room/author/direction/transport header into split continuations without parsing message bodies or duplicating their bytes. Room draft and correction prompts require meaningful decisions, approvals, outcomes and unresolved commitments of that room, retaining source distinctions (who decided, what was authorized, what stays owed) and one first-person Ouroboros voice. Length adapts to content within the existing output-token ceiling; no per-room word quota, semantic gate or absent room is imposed. Labels establish provenance, not summary success. The mixed Main recent view opts into the same labels; focused Project rendering, membership and explicit `chat_history` retain their existing behavior and bytes.
`knowledge.py` owns note reads, revision-checked writes and indexing; `tools/knowledge.py` exposes them. `knowledge_write(mode="edit", old_str=..., content=...)` replaces ONE body occurrence under the current revision, preserving other bytes and frontmatter. Missing/ambiguous anchors, absent notes and stale revisions refuse before writing. Character/heading deltas reach results, bounded there but complete in history. Edit adds storage capability, not a semantic writer policy; overwrite/append stay unchanged. Blank revision creates only a missing note. Malformed legacy preambles stay readable with uncertain metadata; missing linked notes remain unwritten.
Ouroboros remains one identity across Main, project rooms, and Background Consciousness: unified dialogue memory remains available to the one agent, while an executing project task preferentially receives its own thread, journal, workpad, and project knowledge. `project_facts.py` routes project facts to `projects/<id>/knowledge`; subagents inherit the root's resolved project id and never derive a new one; identity and the scratchpad are one canonical pair written from every room, with no per-project copy. Project `journal.jsonl` records curated milestones and `workpad.md` retains active working context (`tools/project_journal.py`); focused context includes the workpad in full and recent journal rows with a visible pointer to older entries. On root completion, only high-signal blockers, questions, and interface contracts are mirrored once from the ephemeral task-tree ledger into the durable journal, and a finished root whose effective working tree is not the registered `working_dir` writes one typed "work lives at <path> @ <sha>" journal row from facts the task record already holds. A project digest gives consciousness a concise completion signal without pretending to be the raw project memory.
### Skills and extensions
Skill capability grows through independent gates, each with its own module: discovery and manifest parsing (`skill_loader.py`, `contracts/skill_manifest.py`), content-hash-bound review (`skill_review.py` / `skill_review_runner.py`), owner grants, dependency reconciliation (`skill_dependencies.py`, `marketplace/isolated_deps.py`), enablement, readiness (`skill_readiness.py`) and execution (`tools/skill_exec.py`). Discovery establishes identity, source, provenance (`.self_authored.json`), conflicts and hash; it confers no trust. Review status, grants, enablement and dependency health stay independent durable facts under `data/state/skills/<name>/`; the gate sequence, payload buckets and marketplace/hub layer are §13's. `skill_readiness_for_execution()` composes review, hash, enablement, grants, dependencies and peer conflicts into phase-specific next actions and reads the mode-aware `gate_for` projection (`executable_review`), never a raw verdict string. The owner may attest their own skill, or a hash-verified official hub payload, to skip the expensive LLM review (`skill_owner_attestation.py`); the deterministic preflight floor still gates it. The model-facing catalogue (`list_skills`, the per-turn Installed Skills section) projects the same live extension facts as `/api/extensions`; each enablement change best-effort appends a typed `skill_enabled_changed` row to `logs/events.jsonl` naming the actor (empty for writers not yet labelled), and an append failure is logged and never blocks the enablement change. Cyber review failures remain visible advice; visibility, successful execution and review PASS are still different facts.
A skill may declare a reviewed `presence:` behavior profile: instructions, knowledge topics, runtime defaults and portable capability requests that installation-local selections resolve to exact targets. Admission requires the bound behavior skill installed, enabled, freshly reviewed and complete for every required request, then compiles one immutable positive capability ceiling copied through `task_contract`, so mutable skill/Settings state cannot broaden a live turn. An optional owner-local `workspace_root` (chosen through `configure_presence` or the Skills settings form, validated by `workspace_admission.validate_workspace_root`, frozen into the task's workspace contract) changes the file/process target while retaining canonical shared memory; it neither derives Project scope nor forks the data drive, and later edits and promotion/follow-up tasks keep the admitted folder rather than rereading mutable settings. A folder that has become unusable refuses the promotion typed (`workspace_unusable`, the repair in `detail` naming the profile) and writes no chat row — the public conversation's synthetic chat id is never spoken to.
`extension_loader.py` and the isolated-dependency layer load only a ready, hash-matching extension — in-process via `PluginAPIImpl`, isolated-dep/native ones as child-process proxies — and a review PASS alone proves neither dependencies installed nor a widget loadable. Per-skill health at `data/state/skills/<name>/health.json` is process-qualified: the server's observation is authoritative, a worker's is a handoff-qualified view.
`skill_lifecycle_queue.py` (§13) exposes queued/running/succeeded/failed plus stale metadata — stale is recovery evidence, not a fake unlock of a still-running thread. Scheduled work is reconciled by `resync_skill_schedules()` and runs only after `skill_readiness_for_execution()`; schedule evaluation is a DST-aware system on the shared cron/timezone contract. Evolution remains hard-blocked in `light` runtime mode (§5). The existing payload binding (`skill_payload_binding.resolve_skill_payload_base`) fills an omitted skill name or bucket from a valid selected normal/repair `TaskConstraint`; explicit selectors still have to resolve to its same physical payload, and collision, native-mutation and child-profile restrictions remain independent. A selected-skill repair (`skill_repair_admission.py`, §13) is admitted against an immutable `base_content_hash` and verifies the observed `expected_content_hash` before each payload write, advancing it after opaque process work without claiming authorship; drift makes the repair STALE — no long shell lock, no rollback, other writers never blocked. These separations let skills expand capability without turning discovery, a UI toggle or old review state into execution authority.
#### Skill publication
The passive installed-skill projection never launches Betterleaks and never claims the current bytes are publication-ready (`skill_publish_eligibility.py`). Selecting Publish calls `POST /api/skills/{skill}/publish-preflight` (`gateway/skill_publish.py`, read-only), which resolves one current payload, captures and scans its bytes, recomputes review staleness and returns exactly one backend-authored state (`ready`, `warnings`, `needs_attention`, `repairable`, `hard_block`); the browser only renders those facts, and only `hard_block` prevents task creation. The authoritative flow is:
`passive index (no scan) → selected preflight → explicit confirmation → ordinary managed task → immutable current capture with separate review provenance → payload scan → GitHub read-only planning → derived-output scans → first GitHub mutation → validated same-skill pull-request receipt → ordinary acceptance`
Publication validation and immutable capture share `skill_publish_eligibility.publication_author_acceptance`: fresh publishable critic authority or qualified current Advisory author authority. Blocking keeps fresh critic requirements; owner attestation alone is not that independent-review/author chain. The host PR checklist states original critic and current author/published hashes, referenced review basis and author rationale, without claiming a second independent review. Recorded findings retain their actual severity, including critical and partial-quorum findings; catalog hashes and PR receipts never become review PASS. Passive and selected preflight use the same authority, with current captured bytes authoritative in selected preflight.
Every outbound byte derives from the capture (`skill_publish_snapshot.py`); the mutable live payload is neither reread nor rehashed to authorize the transaction (no time-of-check/time-of-use drift). Outside Cyber Pro, literal Betterleaks `high` findings (`skill_publish_scanner.py`) block the current outbound call; lower or unknown confidence is a redacted warning. Cyber keeps scanner findings/failures and review status as independent advice; an unavailable scan does not invalidate an observed publication. Packaged installs resolve the bundled `betterleaks-standalone`; source checkouts resolve the managed runtime installed explicitly with `python -m ouroboros.betterleaks_runtime install` — Publish never downloads it. A top-level `skill_publish` task (`tools/skill_publish.py`) is accepted only when pre-truncation metadata holds a validated pull-request receipt for the requested skill and configured Hub repository; the receipt proves the observed publication and never manufactures review PASS. A definite failure before any submission-branch request fails an explicit Publish objective even under degraded review; a last-confirmed stage never proves the next request unsent, so a previous unknown branch/PR attempt or valid same-target receipt is preserved.
A failed envelope (`skill_publish_result.py`) names its cause beside the stage: `reason_code` (the failed stage — `fork_sync_failed` and siblings), a `repair_hint` chosen from producer evidence, and the transport's sanitized `error_detail` with `github_status`/`github_operation` read only from gh's own error shapes, never inferred from prose. Fork synchronization is mandatory before any branch, commit or PR mutation; ambiguous PR settlement is a read-only exact lookup. Publication metadata is read from the producer JSON before the shared ToolResult host-note separator, so appended route or safety notes cannot hide diagnostics or a receipt.
Owner lifecycle actions share `skill_lifecycle_actions.run_skill_action`: grant/toggle run in the lifecycle lane, local delete uses `skill_uninstall_state` (uninstall tombstones and authorized local deletion keep separate retention), and attestation keeps its deterministic floor in `skill_owner_attestation`. The host supplies actor identity and checks exact resource/revision plus an existing member chat message, answered quiz or owner-mailbox record when owner intent is needed; the model interprets that source, and a generated edit-and-review request confers no grant, attestation, deletion or enable authority. An enable with no explicit chat/quiz/mailbox source resolves the original owner source itself (a client `allow_enable` flag authorizes nothing, Presence ceilings come first, a newer owner disable in `enabled.json` wins), and a Repair review never auto-enables — the model chooses the ordinary `toggle_skill` call. There is no permission ledger and no HTTP impersonation; the auto-grant policy stays separate from enablement, and an explicit owner disable survives review and free replay. `skill_exec` returns the revision captured before its script launch, and extension tool receipts carry the descriptor's `content_hash` from the same publication as `extension_generation`, only after physical dispatch; neither is review PASS or a semantic test verdict.
Marketplace update keeps the owner's selected version through retries; `install.PayloadRollbackSnapshot` captures payload/environment and the affected lifecycle-state quintet for update and adopt alike, and the single restore path verifies the required reload before claiming `rolled_back`, preserves independent enablement/history, and never deletes an already restored payload after a failed state write. Catalog updates disclose exact version strings and infer no ordering. The atomic publication-record owner (`marketplace/provenance.py`) also clears locally: it compares the displayed `published` object before setting that section to `null`, preserving unknown siblings, and changes no GitHub PR or installed bytes.
### MCP and browser-facing external tools
`mcp_client.py` owns configured HTTP/SSE and local stdio MCP discovery and invocation (HTTP surface `gateway/mcp.py`; tools are named `mcp_<server>__<tool>`). HTTP/SSE entries validate URLs and auth headers, and URL userinfo is masked through the shared secret projection (`secret_masking.py` owns the exact MCP token placeholder shapes). Load-time placeholder repair runs before environment precedence, so a real environment credential is never mistaken for a wire mask, and only for recognized top-level Settings secrets — nested MCP values are never silently migrated. Stdio entries pass one executable `command` and an exact string `args` list to the MCP SDK without a shell; optional `cwd`, literal `env` and `env_from_settings` (environment name → existing setting key) resolve from the same saved configuration for discovery and calls. Unknown fields are retained with a visible not-applied warning while a valid server stays usable; invalid known fields or references produce `MCP_CONFIG_ERROR`; the response-only `auth_configured` flag never becomes configuration. Discovered tools join the selected initial capability envelope beside enabled, granted extension tools, behind the same `network` resource gate as on a managed task (`schemas()`, `get_schema_by_name` and `execute` agree on that), and a discovery failure is an explicit capability omission through `list_available_tools`, never a silent removal. The catalog lookup (`MCPManager.resolve_tool_name`) precedes the paid safety check: an unlisted name is `UNKNOWN_TOOL`, a disabled server or an unlisted catalog (health unknown) `MCP_UNAVAILABLE`, an allowlisted-out tool `ACCESS_BLOCKED`; none runs safety, transport or a refresh; hits recheck settings before Safety to honor revocation. Hints name only the addressed server's `raw → callable` pairs and an exact naming-rule identity (`mcp_client.naming_rule_matches`); nothing is ranked, aliased or dispatched, and raw names stay out of schema descriptions. Cyber metadata-target and configured `allowed_tools` filters do not veto an existing MCP call; descriptions and results remain untrusted data; every call still crosses registry, resource, safety, timeout and result-handling policy. Web-tool prohibition and network prohibition are distinct: `web=false` alone does not disable configured MCP or extensions, while an explicit `network=false` and tool disables are enforced at discovery and dispatch.
Browser tools are stateful and thread-sticky because Playwright sessions and greenlets have thread affinity; they cannot be scheduled as ordinary parallel stateless calls. A stateful-tool timeout therefore RETIRES the browser generation: the shared `browser_state` slot is replaced and the close runs on the owning thread once the hung call settles. Retired sessions are bounded in-process at `_RETIRED_GENERATIONS_MAX` per task, after which another session is a typed `BROWSER_BACKLOG_RETIRED_SESSIONS` refusal; generation isolation is best-effort under concurrent replacement (a fully closed class needs a process-isolated browser worker — disclosed future design). Model-driven in-page evaluation runs the supplied expression without substring guesses through `_evaluate_bounded`: Playwright's `evaluate` accepts no timeout, so the expression is raced against an in-page rejection, which bounds the ASYNC class honestly and no further — a synchronous event-loop block cannot be interrupted from inside the page, and the outer tool timeout remains the backstop. The caller's timeout also becomes the session default (`page.set_default_timeout`), floored on the action path so the five-second action default cannot strangle a capture. Chromium is the default; WebKit and device descriptors are targeted tools for a real Safari/iOS risk, not a universal acceptance matrix and not a claim that a narrow Chromium viewport is Safari-equivalent. First-party PR helpers are normal built-ins whose mutating operations stay subject to selected-root policy, runtime mode, delegated-child constraints, credentials and reviewed-publication authority.
Browser target policy is `browser_policy.py`: ordinary owner-control admission uses the actual HTTP request and three-valued service identity — never JavaScript substring guesses, no per-network-request LLM, no second policy engine. Cyber root/acting requests retain target and control reach; explicit read-only assignments keep their restricted contract. A live binding in `state/server_port.bindings.json` proves an endpoint ours; a service's missing snapshot can mean an older installation or failed publication, so its recorded facts still name what is EXPECTED — the integer `state/server_port` (the launcher's `state/server_process.json` can prove main), the Host Service configuration beside an expected main, and the local model's custody row with its live argv — and a matching expected endpoint whose process cannot be verified is refused as unknown, never treated as foreign (`server_process.runtime_service_identity`). Only a live snapshot of the same service supersedes its legacy expectation: main cannot vouch for Host Service. Unmeasurable POSIX identity never counts as proof. Every other port is an ordinary target, an unrelated application reusing an `/api/owner/...` pathname included — the owner-operation request shapes apply only at a proven or unknown Ouroboros endpoint — and a restricted target that DNS cannot classify is a typed `BROWSER_POLICY_UNAVAILABLE` refusal before navigation. Chromium and WebKit follow HTTP redirects natively and Playwright route callbacks see only the first URL of a chain, so `tools/browser.py` re-checks each document's `redirected_from` chain with the same target predicate before any page result: an allowed→blocked→allowed redirect withholds its content and refuses actions on that document, while the forbidden hop's request has already been dispatched — an owner-accepted, disclosed residual, not a pre-request or DNS-rebinding guarantee. Cookies, POST semantics, service workers and native redirects are untouched: no proxy, CDP fork, pre-probe or refetch.
### Budget tracking
`swarm_efficiency.fanout_count` counts observed fan-out emissions and `fanout_interval_sec_total` sums the wall-clock gaps between them, intermediate parent work included; new events emit `fanout_interval_sec`, and readers tolerate the retired `wave_count`, `inter_wave_latency_sec_total` and `inter_wave_latency_sec` without rewriting stored bytes. These observations never infer semantic work waves or child-wait time.
The time/cost/intrinsic pacing checkpoints carry `resource_facts`: incremental per-tool call/error counts and producer-reported duration intervals (which may overlap; unmeasured calls stay unknown), plus own-task, tree and delegated-tree ledger buckets (own and delegated explicitly overlap the tree total) and the shared unreserved global remainder. No argv, stdout, sleep/poll classification or new stop rule selects behavior, and prompt and checkpoint show the same facts.
`usage_accounting.py` owns monetary policy over the physical-attempt ledger. Each send gets an ID and `reserved → dispatched → settled | unresolved`, or pre-dispatch `reserved → released`. Only typed proof of no sent bytes (connection/pool failure) permits `dispatched → released`; timeouts/unknown errors remain unresolved. Each retry is a new attempt; SDK/stream boundaries are labelled opaque. Main/direct/child/scout/review/safety/synthesis/reflection/consciousness work, retries and opaque SDK calls are covered; root scopes count tree and post-task/review work once. External scripts/extensions with model credentials stay unknown/unmetered at host-observed opaque boundaries absent authoritative settlement.
`cost_projection.py` owns producer amounts/openness: `accounted_upper_bound_usd` = settled+reserved+unresolved, null is unknown, finality is never invented, and `COST_OPENNESS_FIELDS` accompanies amounts. Retired `cost_usd[_with_children]` stays read-only. `cost_presentation` binds amount/facts to one bucket (§11.1): `reconstruct_task_cost` owns own scope; root terminal/heartbeat/checkpoint/synthesis paths use root-tree, never falling from null subtree to own zero. `_usage_rows._summary`'s weighted `priced_rows`, `tracked_nonfinal_rows`, `accounting_open_rows` survive compaction. A price or retained bound evidences zero; empty/unpriced rows do not. Settled unpriced rows leave the known subtotal exact; estimates (zero included), retained bounds, open attempts and integrity gaps stay nonfinal. Terminal refresh/copyback, synthesis, heartbeat and history preserve scope; unknown never becomes sticky-final. Sums and `cost_final` are unchanged.
A proven abandoned send closes administratively as `settled` with `settle_reason="abandoned"`, `cost_usd=None` and `cost_final=false`: its reservation bound stays accounted, not confirmed spend, and the price remains non-final. A real late receipt may correct the same attempt once (`late_receipt`); a positive never-started receipt may instead release it through the existing `before_dispatch_failed:` seam. Ordinary terminal settlements and releases remain immutable. Full and incremental ledger validation preserve the same correction eligibility; compaction retains unresolved and abandoned chains for that exact-id join, while existing historical baseline groups remain aggregates (the full storage contract is [Usage-ledger compaction](../USAGE_COMPACTION.md)).
After cancellation custody, existing maintenance calls `server_maintenance._reconcile_abandoned_usage` for settled tasks with no live physical owner/open post-task synthesis; review attempts retain their review owner. `llm_claudexor.recover_model_attempt` reads retained terminal CAS or the exact operation via owned-only discovery, retains bytes before ACK and never starts/restarts/cancels bookkeeping work. Missing identity, live work or unreadable evidence defers. The money lock rechecks each row so real concurrent receipts win; no age cutoff, timer or ledger is added. Projection reconciliation includes compacted-baseline attribution, retaining retries after failed result writes/native late receipts. One bulk breakdown feeds task/root owners; foreign canonical budget roots stay separate. `events_task_done._refresh_terminal_task_cost` changes money fields only, without completion events or invented synthesis.
Summary/reflection/consolidation share a frozen ledger snapshot: settled subtree cost, reservations, unresolved bounds/counts, unknown exposure, integrity and non-final state. Final terminal checkpoint alone owns final cost; read failure is unavailable/null, never zero. No reconciliation LLM or parallel ledger exists. Cache-hit health uses only rounds carrying `cached_tokens`, and stays absent with insufficient samples: unmeasured is unknown, not zero, so providers lacking telemetry cannot dilute measured ratios. Unknown spend remains the existing `unknown_unmetered` field, not a third accounting status.
`root_phase_checkpoint.accounting` retains a cumulative ledger observation: root ID, accounted/reserved/unresolved amounts, non-final/unresolved counts, unknown exposure and integrity. Separate from own-cost and read-time `TaskCostBreakdown`, it survives public detail/recovery and canonical replica protection. The checkpoint owner refreshes it, including explicit null/unavailable on failure. These facts prove neither an invoice nor local paid-work closure; phase/execution ownership is independent.
**Planning threshold** (`task_pacing.py`). One unreserved `CostCeiling` (`disabled`, `active`, `exhausted_soft_land`, `unknown`) owns runtime `in_task_cost_ceiling` disclosure and loop checks. Roots use min(global-remainder percentage, root cap) minus margin; uncapped roots use starting wallet. Enabled descendants inherit that number, not later percentages. Disabled profiles disable only pacing; monetary admission stays independent and surfaces name the binding bound. Rooted tasks compare subtree spend/holds even without root caps; own-spend fallback is a disclosed lower bound.
`loop_budget._finish_tool_round_budget` and `_finish_no_tool_round_budget` share `_check_budget_limits` after unfinished spending rounds. No-tool tails, including cold Resume (`resume_point.budget_tail`), neither advance nanny baselines nor arm tool controls. Candidates survive; READY returns before this check without an extra paid final. Eligible actors pause, even before work on a soft threshold. Ineligible actors retain the priced terminal rail (`budget_wrapup_unaffordable` if it cannot fit); global rejection before work buys no call. Tool requests remain incomplete even beside `replace`.
Affordability probes reserve nothing; full-cap ledger admission decides. Unknown spend is not zero; unknown price fails open. `_usage_cache_splits.py` reuses settled provider/normalized-route/review-surface splits: cache loss reprices full write, expiry may under-reserve one write, competitors may consume room. Proxy stops/native images require exact transcript-copy pricing before service finalization. `task_pacing.main_loop_wire_options` owns wire options with no later payload additions. Identity drift refuses pre-send (`forced_candidate_drift`); terminal handling retries once unpredicated, still ledger-priced, so drift cannot lose the answer.
#### Exact budget pause and Resume
`budget_pause.py` pauses pooled/direct work under the SAME ID on global exhaustion, graceful ceiling, either last-fit stop, soft landing or refused dispatch, without a paid final. A missing task id/root/continuation owner keeps terminal behavior (`exact_pause_unavailable`). Increases never wake tasks.
| Owner | Ordered contract and reason |
|---|---|
| `request_pause`, `usage_accounting.reserve_attempt` | Close `DispatchFenced`; review POST/replay and Light extraction also refuse (`budget_pausing_no_send`, `budget_pausing_no_extraction`). Persist `pausing` before waiting (direct RUNNING stub if absent) so crash custody cannot replay ambiguous work; the reaper derives the crash-retry fence from its one pause-row read (`worker_health._complete_exact_budget_pause_after_death`), failing closed on unreadable evidence. |
| `local_producer_observation` | Hold the worker nonterminal until sent reviews settle through `review_custody` AND timed-out tool futures finish their late settlement callbacks (`register_tool_future`/`hold_tool_settlement`); `done()` alone is insufficient. Failed writes/producers publish `budget_pause_hold` on change and retry, never claim a pause or buy a final; a publication failing after quiescence retries the SAME prepared snapshot, discarded if a producer revives. Only Stop/Panic/deadline/cancel ends the hold as `abandoned`, via the loop's model-wait control rails (a no-call terminal, never a task exception); controls read the canonical budget root on split drives. |
| `observe_task_runs` | Re-read custody (the loop side resolves its own custody root; the grant reads it directly); unreadable is `custody_read=failed`, not empty. Pre-terminal subscription coverage is unproven, so unsettled runs request `cancel_and_verify`, retaining typed outcomes. Unknown stop permits no second writer. |
| `owner_wait.continuation_state` | Retains cognition, opaque acceptance/preparation, usage/rounds/clocks. `resume_point` keeps the budget tail and unanswered call IDs; Resume inserts execution-UNKNOWN host rows, never replays tools. Source/row precedes `BudgetPauseRequested`; no task_done/result/Main final. |
| `events_budget.install_exact_budget_pause` | Validate event/source; park RUNNING or `parkable_direct_task` in PENDING as `_budget_pause`, persist snapshot, then mark `paused`. Root scope fences the tree. Park, projection and grant share the queue lock; pause-id/state CAS turns late publication into `budget_pause_park_superseded`. Direct records keep `_is_direct_chat`, omit inline image bytes and release the local fence once the actor unwinds (the durable row owns the hold). Resumed direct work uses a pooled worker; `log_addressing.address_task_event` keeps the lane. Projection carries ledger cost planes; a live-pause canonical row outranks the child replica and the queue mirror in `task_status.effective_task_result`. |
| `worker_health._complete_exact_budget_pause_after_death`, `queue_snapshot._park_pausing_running_rows`, `workers.kill_workers` | Complete a saved pause after death/restart; shutdown leaves exact carriers/rows untouched (never `pending_parent_interrupted`); boot re-parks. Revoke an UNCONSUMED grant against its current identity and re-park; a CONSUMED grant takes terminal crash custody, never ordinary retry. Failed revocation keeps source/grant in a nonterminal hold (memory-only if unpersisted). |
| `budget_pause_restore_refusal`, `events_budget.budget_hold_fact` | Pauses restore at any snapshot age/stamp validity without wake. Missing/unreadable/mismatched/terminal records or sources, root acceptance or malformed snapshot fences retain `_budget_pause` under `_budget_pause_hold`, never drop/cancel. Regrant revalidates. |
**Owner grant.** `POST /api/tasks/{id}/resume` → `queue_transitions.resume_budget_paused_task` → `budget_resume.grant_exact_budget_resume`: require authoritative positive global/root headroom, readable source, clear cancel/restart/Panic/deadline/lifetime rails and an unpaused ancestor root; re-read failed custody. Root headroom is ONE fresh strict ledger read (`usage_accounting.refresh_root_accounting(strict=True)`): a cached snapshot after a failed read, an unreadable ledger or a degraded tree refuses typed; display readers keep a bounded-stale fallback. Pause ID/generation and `resume_generation` bind a single-use grant. `revoke_exact_budget_resume` re-parks after lost money/restart, keeps newer identities and retries unwritten revocation before regrant. `model_wait.execution_elapsed_seconds` subtracts quota-union and cumulative `paused_duration_sec`/`budget_paused_sec` from original `started_at` across continuations, owner waits included; None stays unlimited. Money/round/review wallets never reset.
**Descendant selection.** `events_budget.hold_root_resume_descendants` makes children eligible, not runnable: exact rows stay paused; zero-dispatch/fence-only siblings get `root_fence_lifted_pending_selection`, rebound to the current root grant on each Resume. `resume_child_task` (`budget_resume_child`, recorded `selected_by`) selects within lineage: root selects stored descendants, an intermediate parent only direct children; owner selection shares the path. With a live latch, `budget_fence_selected` binds one member to its fence generation, leaving siblings fenced; legacy root Resume selects only root; `reserve_attempt` admits the member selected against that exact fence. New root pauses invalidate pending child grants; assignment rechecks root grant and fence for exact grants and zero-dispatch selections (stale → unselected hold). Other roots/cancelled/completed/stopped members never revive.
**Consumption.** `resume_paused_loop` rejects spent/revoked/foreign grants and discloses drift/custody before effects; a consumption publication that fails HOLDS (`resume_grant_consumption_unwritable`; identities kept), a grant revoked underneath re-parks; never FAILED. Graceful Resume refreshes ledger-authorized headroom via the same strict tree read and admits ONE fitting reservation (`_second_reservation_fits`, `budget_resume_last_fit_admitted`), avoiding the early-margin re-pause. Hard exhaustion needs owner increase; full-cap admission binds. Browser/services/hidden state and crash recovery are not restored.
#### Monetary authority and projections
**Live global limit.** Agent readers share one resolver: saved settings first, re-parsed only on file change; failure (including refused benchmark pins) falls back to environment without caching failure. Workers refresh environment only at task start, so live file reads make every reservation/wallet follow current `TOTAL_BUDGET`; planning thresholds remain separate. Absent means shipped default; non-positive means unbounded and silences the loop global axis. Disclosed pre-existing gap: supervisor startup/reload still parse raw settings and treat absence as unlimited.
`pricing.py` uses bounded best-effort exact normalized-route lookup and provider catalog fields, never manual tables/prefix inheritance/fallback prices/admission allowlists. Unknown price admits while known spend fits: reserve None, settle from reported cost or later exact price, otherwise None/non-final. Only structural pre-generation evidence with zero usage permits confirmed-zero rejection/release, preventing phantom exhaustion during provider storms (`_usage_response.py` normalizes; adapters retain raw usage). `review_wave_admission` uses the same estimator and tighter global/root remainder as `reserve_attempt` before skill/plan/task/commit review; managed-update assisted apply reuses it before destructive merge. Unpayable slots bypass without model substitution. Unknown-priced slots disclose uncertainty without disabling priced siblings' admission.
Reservations retain applied `global_limit_usd`, explicit unbounded state and `global_limit_source`/`global_limit_revision` through transitions. Overrides inherit no foreign revision; legacy unknowns remain unknown. `usage_compaction.py` archives original rows beside the substrate, preserving sequence; commit requires identical NON-MONEY projections and decimal-identical money because six-decimal display rounding is not monetary equality. Policy abort emits `usage_ledger_compaction_skipped` with cause once per process/cause; name-tier refusal keeps `usage_ledger_compaction_refused`, making size warnings diagnosable.
`usage_ledger.py` owns one short cross-process lock for validation/reservation/transition/append/fsync; accounting imports it, never vice versa. Network I/O stays outside. Torn tails quarantine loudly; validated prefixes remain readable but degraded/non-final because paid work may be missing. Failed settlements retain dispatched/unresolved attempts; durable root-budget refusals clear on Resume only after proving no paid dispatch. `_usage_rows_memo.py` incrementally replays validated records, changing cost of reads, not meaning; append repairs a torn newline-less tail without losing earlier rows. `_usage_rows.py` owns pure arithmetic.
For a root task, `GET /api/tasks/{id}` derives `cost_breakdown` at read time from the same ledger (own, child, unattributed, disclosed delegated, subscription sessions, unknown/unmetered, finality, authority); it is never persisted, is not a third sum, and an unreadable or unattributable ledger omits the whole object rather than returning a confident zero. One physical reviewer send produces exactly one `llm_usage` row, emitted by the review substrate with that reviewer's wave and slot attribution (`skill_review_usage.py` is a read-only projection per `(review_skill, review_wave_id)`, not a second ledger), and a delegated session row reports its own route provider and resolved model, never an inferred one. `state.json`, task results, `llm_usage`, `/api/state` and `/api/cost-breakdown` are compatibility projections only; startup's resumable importer (`usage_legacy_import.py`) records source hashes, imports only attributable usage, and represents ambiguous history explicitly without rewriting source logs or fabricating attempts.