ouroboros/docs/architecture/06-agent-core.md
Ouroboros 6a2ee44ac2 docs/tests: narrow the transport-death repeat wording to Host-executed effects; state the 409 condition exactly
Reviewer P2: a lost response proves no Host-executed tool or correspondent action from the failed attempt, not that nothing ran anywhere — priced read-only provider-owned retrieval (server-side web search) may rerun on the repeat. The 409 and owner notice apply when the round ends unresolved without a further permitted repeat (deadline or a finalize control can refuse the first repeat). Test comment on the owner-notice writer corrected; chapter budgets untouched.
2026-09-26 02:00:57 +03:00

303 KiB
Raw Blame History

6. Agent Core

This chapter maps execution and cognition inside a worker: task lifecycle, tools, context fitting, safety and runtime mode, reviews, delegated custody, planning, deep self-review, reflection, memory and project focus, skills, external tools and budgets. Pooled tasks and direct turns, including consciousness wakes, enter OuroborosAgent's tool loop (loop.py); reviews and post-task operations use separate executors. Shared contracts do not make these paths one loop.

Task lifecycle

A queued task enters through a reviewed transport, is admitted by the supervisor queue (§5) and runs in OuroborosAgent; each direct-chat turn runs on its own in-process agent, tracked by the process-local DirectActivityRegistry, and creates no PENDING/RUNNING queue record (§3). The root pipeline (agent_task_pipeline.py) captures the task contract and immutable context core, runs the loop, preserves a delivery candidate, stores the result (task_results/<id>.json, every write stamped _schema_version: 1; an unstamped, future, malformed or retired-key row is quarantined by task_result_schema.py with its id kept occupied) and artifacts, emits lifecycle and usage evidence, performs the root-only post-task work (post_task_synthesis.py, checkpointed by post_task_checkpoint.py) and publishes the typed outcome. Queue admission proves only that asynchronous work was durably accepted; completion, objective satisfaction, artifact finality, verification and review acceptance remain separate facts. The loop's tally rides loop_outcome.usage while the top-level total_rounds/prompt_tokens/completion_tokens are the ledger's answer (reconstruct_task_cost); an internal exception is failure.kind = "runtime", never a fabricated provider failure, and a capture the crash lost is reported unknown, never as zero counters.

The separate review and post-task executors (review_execution.py, the post-task pipeline) use the same evidence and context contracts. supervisor/cognitive_operations.py tracks in-flight LLM, review, VLM and tool work as typed cognitive_operation leases for the idle rail; deadlines, budgets, cancellation and absolute ceilings remain independent.

Between the sends of one loop execution the transcript is append-only — each send a prefix extension of the previous — because OpenAI-family caches reuse a request only when it is a byte-prefix of the next, so a replaced tail or a rewritten earlier message discards the whole conversation cache. transcript_prefix.py records and never blocks: unsent_in_previous_send allows _append_or_merge_user_content to merge only a tail absent from the last observed send; an absent slot or missing observation means unknown, so the producer appends a new row, compaction seams stamp sanction_rewrite, and a break is the prompt_prefix_break checkpoint fact (kind = system_rewritten | tail_replaced | rewritten | shrunk, plus sanctioned_by; a context-fit reprojection after a real overflow is an ordinary break). The observation follows a usable ordinary model response; it is not a ledger of every physical send, and existing image eviction is unchanged. A FINAL ANSWER: line is latched every round as the latest typed candidate, so review, nudge and forced-finalization paths never erase a structured answer; marker prompting is gated on task_contract.answer_protocol="final_answer_line" (answer_protocol_active) while the latch and extractor stay unconditional, and outcomes.extract_final_answer refuses the outcome-tier ledger identifiers as answers — internal enum vocabulary is never a deliverable.

model_execution is a compact projection of the last ordinary solve response the loop accepted (text or tool calls), selected from marked llm_call_refs; it keeps the initially requested model/local route, the route used and the provider's own label separate, and empty responses, forced finalization, review and post-task calls cannot replace it. Result, terminal frame, history and card share it without changing dispatch or claiming final-answer authorship; the terminal task_summary keeps it when the full result ages out, and a current result or live card outranks that fallback. After tool work, host-driven fallback or model-wait model changes record an authoring handover. The successor's first tool-less response gets one recovery round with delivery-control JSON held; a second yields execution=degraded and authoring_handover_incomplete. Ordinary no-tool turns and explicit model changes stay unchanged. Verification and skill checks still run; later tools clear only this warning, preserving incident and recovery history. Each handover gets one recovery prompt under ordinary stop, deadline and budget limits.

DeliveryCandidate is retained before verification or review so a later notice, reviewer failure, deadline or provider outage cannot erase a useful answer. outcomes.py combines the execution, objective, review, artifact and child-absorption axes without converting one into another — the terminal custody overlay (outcomes.custody_debt_axes) is an instance of that rule, not an exception; verify-before-done receipts and exact artifact references are host-attested evidence, and declarations or answer prose are no substitute. Individual failed tool calls alone do not degrade a delivered answer: unresolved calls retain their typed signal/exit evidence in execution.unresolved_tool_errors, cosmetic errors keep their separate bucket, and either kind adds residual_tool_errors_without_review when the canonical objective is not_evaluated, including after delivery/child-state normalization. Host acceptance or qualified Advisory author completion establishes objective success; unresolved FAIL/DEGRADED verdicts, empty or failed execution, provider failure, deadlines, incomplete delivery, deferred children and failed artifact verification keep their own outcome rules. A forced exit may publish the best current candidate only with its typed rail and evidence-freshness disclosure, and lifecycle may remain completed while the objective or review axis records a best-effort or unaccepted result. Custody debt heals from the write side while the stored reason_code may not be rewritten, so the owner-facing Reason line resolves the code at RENDER time — the custody warning only while the row's own delegated_runs_unreconciled list is non-empty, else the execution reason, both when both are real; a healed debt is never restored, and the debt is a warning beside the rail cause, never a replacement for it — as is every standing limitation of the same answer (deferred child results; a plan review still open at delivery, worded by the outcome class recorded on execution.plan_review when there is one), which the normal rail types on the result exactly as the forced rails do, so one clause per fact reaches the card, the durable row and Ouroboros's own memory.

Host plan/orphan disclosures ride beside the model answer as terminal_host_notice; delivery sends the model answer alone and the disclosure stays a field of the result (an owed pre-upgrade notice row still replays as the untyped System row it was), and the custody audit is its own typed row (terminal_custody_notice on the send event → system_type="custody_notice", card_row="timeline", its own owed delivery id), so the open-delegation fact is a row of the task's card and the unreconciled-runs note never enters the assistant text. CLI gets one host-labelled status (terminal_host_notice_text); external Presence speech excludes host notices (chapter 12). Stored answer bytes do not change: the answer keeps its hash and its real PASS or FAIL when only the notice changes, and equal text cannot revive a superseded verdict. Every reader that hands a result on — synthesis, parent handoff, get_task_result/wait_task/wait_tasks, the task:<id> plan-evidence reader — carries the notice as a separate field, and the child-result and plan-evidence hashes include it, so a changed child limitation invalidates an old parent disposition or an old review. A host-salvaged terminal is labelled, not hidden: the durable row carries Preserved intermediate output (not a final answer): plus the bounded excerpt every terminal row uses, the untruncated copy stays with get_task_result and the stop receipt, and only a row written to the chat the stop receipt itself reached (cancel_receipt.delivered_chat_id, recorded after a successful send) reduces to its label beside the pointer; the two durable writes are not atomic, and no writer trades its only pointer for the label. One event is disclosed once, at the layer that owns it: the forced orphan note leaves a child to its own terminal row only where that row reached the reader the note addresses, and a rail that ended a routing turn is named by that turn's own row, never inferred from another layer's stamp.

The delivery-control protocol is resolved here and only here (DEVELOPMENT keeps the rule and points here). The candidate carries sticky loop-local provenance that its lineage has seen a host-issued delivery-control episode; without one, exact JSON is ordinary text. In a marked lineage both the ordinary and the forced resolver intercept recognizable whole-body envelopes and balanced trailing protocol attempts — valid keep resolves to the retained candidate, valid replace to full_answer, anything malformed preserves the retained candidate. Both strip one whole-body fence and treat a balanced protocol object at the very END of prose as a protocol attempt (utils.extract_trailing_json_object + loop_delivery._parse_delivery_control_body); the trailing-object rule deliberately refuses substring scanning, because quoted protocol literals mid-prose are legitimate text, so a control object quoted mid-prose stays prose and a truncated trailing fragment remains prose. During an ordinary acceptance continuation, finalization_control="acceptance_feedback" makes a complete revised answer ordinary prose and keeps keep/replace/pending_review optional. Prose resets the pending-review choice to wait; it never means finish. An outstanding effect, owner-revision or child-action control retains its stricter rule, and unread owner source still requires acknowledgement. Empty or recognizable malformed control bodies retain the candidate rather than becoming its replacement. The ordinary resolver takes one repair round then degraded-preserve; the forced resolver resolves purely and never re-loops — malformation preserves the retained candidate with the typed delivery_control_degraded reason, which is this forced rail's own code, while the ordinary repair path records invalid_delivery_control_after_repair. Every degradation carries the cause it computed: outcomes.derive_loop_outcome falls back to delivery_control_degraded only for a degradation that reports no cause and publishes the loop_outcome.degraded/degraded_reason pair the benchmark ledgers read. A malformed attempt in a marked lineage never leaks JSON, even after the transient latch clears. Presence forced calls (turn or root, not ceiling-only child) always arm presence_finish; original bytes reject duplicate keys (chapter 12).

The child-absorption gate is an action gate: while undispositioned direct children remain, the loop HOLDS the candidate (child_absorption_or_revision_required) instead of arming the JSON-only control instruction — the hold-vs-arm split exists so the model never receives two contradictory instructions in one round. A typed keep cannot close the gate; after the one bounded reminder it forces the best-effort children_unabsorbed rail with a current id [status] sha256 listing. The absorption digest's ## child header carries the child's typed custody debt (delegated_runs_unreconciled, bounded, with a get_task_result pointer) as visibility only — the parent's authority over that patch is exactly the orphan rule, and a child's debt never relabels the root card (DESIGN §4). Finalizing over an UNDISPOSED OWN delegated patch is deliberately NOT gated: the consequence is disclosed where the decision is made (the integrate_delegated_patch schema, the apply receipts) and lands as the additive Done-with-warnings custody overlay rather than a hold; a pre-finalization reminder and propagation of child custody debt into the acceptance-subtree snapshot remain disclosed deferred gaps.

Provider death is the one forced rail that is NOT a best-effort completion: it salvages the best available text but stamps infra_failed, so the task terminalizes failed with the typed provider_unavailable reason and an immediate "provider outage — NOT completed" owner notification; a waited-out transport outage reaches the same rail through its deterministic no-resend branch (transport_unavailable_no_resend, or provider_outcome_unknown_no_resend when the round still holds a transport-death repeat record). The rail makes its one forced model call only while a call can still land: a transport that spent its same-model retry wall stamps _llm_retry_wall_exhausted, a no-call gate beside context_overflow and provider_outcome_unknown, so the forced rail never re-pays a second retry window over a proven-dead provider (disclosed residual: the marker is a last-invocation bool on the shared usage dict). Terminal delivery preserves producer authorship: terminal_provider_notice is host presentation beside the raw result, and receipts and secondary System incidents carry the wait duration, finalization cause and unknown-outcome warning and never recommend a blind rerun. Every forced rail stamps one closed-vocabulary producer word at the single forced-finalization sink: model_final for complete model text (host-authored notices stay separate), host_notice for host-written terminal text (an owner System row, never automatic external Presence speech), and host_salvage only on the provider-death rail, where managed/direct delivery replaces the text with the enriched outage receipt and the full bytes stay in task details. Missing origin stays unknown; agent.py resets it on host text replacement. Presence synthesis seals the recorded adapter body and internal result separately (chapter 12), never a delivery receipt.

Finalization controls are typed owner-mailbox entries, not injected owner prose. The supervisor may request one bounded tool-less answer, salvage the last persisted assistant text, and retain a full canonical copy when a preview would truncate it. A grace episode has one durable control and can be revoked atomically when the task itself resumes; descendant activity does not count as the task's own progress. A process that cannot be killed remains visibly running, and custody checks prevent another runtime instance from reaping work it does not own.

Disclosed cancel-lifecycle residuals (deliberate): a cascade over a tree with no resolvable lineage chat whose typed handoff-row append ALSO fails still settles; an empty-intent release can add liveness noise to a foreign claim's forensic trail (bounded by the generation fences); cascade postcondition timing can flake under heavy load (the watchdog re-feeds — one retry, never a lost teardown); and the cost projection of a task whose delegated runs stayed open may read cost_usd=0/cost_final=true while a run is still live — the disclosure line names the open runs.

Task acceptance

Host-enforced task acceptance is a root-owned completion coach, not the P3 commit gate. off disables it; auto and required review observable effects, typed deliverables/criteria, and an explicit root task_acceptance_review nomination, read-only research included. Queue membership alone does not qualify; ordinary conversation, exploration or cognitive-memory updates alone do not qualify in auto, and no prose or tool-count classifier decides their meaning. Child reviews remain advisory evidence superseded by the root decision.

The explicit call nominates the complete ready result and returns deferred_to_host_acceptance, authoritative=false; after the whole tool-result block the host advances the same acceptance operation ordinary final delivery uses, and early feedback does not seal the task. Main authors the effective criteria; the paid subject is result bytes, criteria and material effects, so a status question preserves a running review while a new criterion can buy review of unchanged text. A reviewer panel is advice for its author, never a signature on bytes the reviewers did not read: the wave wakes the original Main through the mailbox at its quorum and again when the last slot settles, each wake carrying each reviewer's own verdict (acceptance_settlement.announce_acceptance_settlement). The optional answer control is prepared before either ready-feedback shortcut, so a settled panel or queued wake skips parking without losing control provenance. Mailbox readiness is distinct from owner authority: transport and waits wake for all entries, while acceptance source capture, acknowledgement and final sealing use the drain's typed owner/control boundary; system, descendant, independent-task and peer-task messages alone imply no owner revision, while owner/principal messages, quiz answers and typed owner controls keep their handling. Only a still-pending panel may park the turn; acceptance_settlement.awaited_panel_has_settled lets the next round, control repair included, run once it settled. Final delivery without re-nomination uses the same task-panel feedback whether ready or pending; delivered feedback stays on its exact review_runs row, never in phantom pending state. For a pending panel Main waits (the default; the only option under blocking; Cyber Pro: Main's final response is its decision) or consciously finishes through the delivery control's pending_review key; a panel that settled PASS on the earlier revision accepts the task on the reviewers' word (previous_revision_accepted, said on the owner row) only when nothing but the answer text changed — same owner source, criteria and material evidence — while a changed subject or any other settled verdict hands delivery to the ordinary path with the collected verdicts in its dialogue history. Before a new panel or any capacity refusal, every recorded still-pending panel of the same root is collected at $0 over its recorded request and roster (review_dispatch.reconcile_pending_acceptance_runs), so a re-authored subject cannot discard verdicts the tree already bought; they enter the next panel's dialogue history. Stop, missing custody and unfinished work keep their observed outcomes.

The host-collected repo_diff uses one bytes capture (repo_diff_capture.capture_repo_diff) for preview (review_evidence_sections.collect_turn_diff) and source (artifacts.materialize_repo_diff_evidence); a legacy partial preview without its capture stays unavailable, never certified by a later read. git writes to a private spool under a real subprocess timeout; memory holds a bounded prefix per section, the rest stays on the spool for retention. Every gap — nonzero exit, timeout, unreadable repository, memory cut, undecodable run — is stated: an unreadable diff is a gap, never an empty (clean) tree. A root PROVEN plain (.git and HEAD both absent by lstat of the root alone; any entry, even a broken gitfile, or an unreadable or missing root stays Git-required) is applicable=false, complete, without a Git call — no baseline, never a clean tree, no source_unavailable partial — and its preparation identity is a bounded content hash (files by content, symlinks by target, no traversal, nested .git skipped; special, unreadable or over-bound material is unknown); Git retargeting variables are scrubbed and the root's parent is the discovery ceiling, so no foreign repository is captured. Two never-interchangeable identities leave it: the EXACT bytes, streamed into the private observability CAS (retain_private_capture → repo_diff_capture.write_blob_stream, 0700/0600, host-forensic, never published; a failed retention is disclosed, not implied), and the REDACTED text projection, decoded and redacted WHOLE per section before any presentation cut so no credential straddles a cut (marked diff lines redact by BODY, file headers kept; a PEM block in a hunk is masked WHOLE by the existing secret_masking pattern), carried by the packet, the actor-readable repo_diff_source_ref and every owner surface. The packet's repo_diff_capture publishes applicability, digest, size, gaps and whether raw bytes were retained, never bytes or a handle; readable hunks of binary, non-UTF-8 and mixed patches stay beside located gaps, an undecodable run being a located mark, not content. Git stderr, exception diagnostics, failed or timed-out stdout and bound-cut sections are withheld from public projections (a failed process can end inside a secret) though privately retainable; complete sections and gap metadata are redacted before cuts. Small complete captures are retained privately by collect_turn_diff, larger ones by the exact-source consumer (or locally without one); spools are released on projection or later-section failure, and a half-created spool pair closes what it opened.

Local packet preparation failures have their own material identity (acceptance_preparation.preparation_source_identity), fixed before packet budgeting, subject hashing or assembly: semantic criteria, canonical receipt content, artifact bytes, child results, service-finalization facts and the repository or plain-folder content inventory — never the paid delivery fingerprint, tool history or indices, repeated writes, staging of unchanged bytes, candidate prose, owner-message counts or exception text. A proven unborn symbolic HEAD uses the empty-tree base; other Git failures stay unavailable. Incomplete or unreadable material is explicitly unknown; lost availability keeps the last known identity, failed attempts and spent retries (a transient unknown never reopens preparation), and size/mtime never certifies equality. Changed material or criteria reopen it; otherwise one more attempt needs an explicit source-bound acceptance_retry (material_change, repair_evidence, owner_retry) after the current failed attempt was exposed: the host resolves the current owner_source_sha256 or an existing verification receipt's exact content (index = locator only), Main judges substance, rationale and basis are no grant identities, a spent source grants nothing again, and a retry beside a terminal author stance is refused. The paid identity is unchanged.

Feedback exposure binds the responding Main request, incident and attempt. Informed Advisory finish or Blocking unfinished stop retains the answer without rebuilding evidence, bound to that attempt, answer and owner acknowledgement with material rechecked; only root auto/required acceptance qualifies, never child/off lanes or outstanding action gates. Reconcile, dispatch and application keep separate stages and paid/unknown custody; preparation failure authorizes no resend. loop_delivery._delivery_evidence_state and its observed_delivery_evidence reader serve retention, nomination, post-tool, controls, publication, forced exits and stance merge. Failed reads yield UNKNOWN (empty), preserving the last known fingerprint and revision. An answer retained over unknown evidence remains one candidate across unchanged repeats, with keep allowed but no verified subject: evidence_current=false, evidence_status=unavailable_local_preparation, empty subject hash, unaccepted binding. Unknown subject rebinding revokes current approval; authoritative UNKNOWN cannot qualify for keep. Forced exit preserves text unaccepted with an unreadable-evidence notice. Publication hash failure has the same projection. The host acceptance pass accounts the incident; retention never opens one.

Local decisions carry origin=local_acceptance_preparation without rewriting reviewer decisions. Publication follows §10: settlements remain visible; delayed snapshots cannot erase newer warnings or resurrect resolved ones. Binding and admission alone are not dispatch; a processing failure after handoff or around existing custody carries origin=host_acceptance_processing beside the real runs, never as fabricated critic findings. The incident rides the existing acceptance decision and review_projection.acceptance_incident, panel or none; derive_loop_outcome projects an unresolved preparation terminal as unfinished/blocked in Blocking (forced rails included) while Advisory keeps best effort with stronger rails, unfinished stops and earlier critic failures visible. Cards and history show the concise warning and expanded detail (completion stays the only notification); resolution clears the warning, keeps history, and a first-attempt success on repaired material resolves the previous warning under its original incident id and attempt count. The final cause states both the local failure and the terminal rail without inventing reviewer rework.

A completed reviewer from an older plan wave is attached as a historical supplement through the locked task-result writer and exact producer CAS: it never rewrites the original verdict, aggregate, closure, author dispositions or current-wave pointer, settles its historical cost without another cycle, and reaches a terminal parent without a new model turn. Task acceptance has the same twin: a panel that settles after its task is terminal is collected at $0 over the recorded operation, republished on the task's review projection with a host-composed late_settlement note (the verdict, the revision it covered, settled after the terminal; honest about a reviewer whose physical outcome is still unknown) and announced once in the task's room as a System row stamped card_row="reviews" (acceptance_settlement.attach_late_acceptance_settlement, deduped by delivery id) — a timeline item of the card, read by the next turn from chat history; no model turn starts.

Before an eligible panel runs, supervisor/queue_transitions.py closes subtask admission under _queue_lock (transport: §10 invariant 25) and task_status.find_child_tasks proves the subtree quiescent; a fence that refused or never answered buys no model round — the panel runs on the disclosed rail admission_fence_available=false. Early settlement releases the fence; final delivery seals ingress by the queue's typed answer (loop_delivery._seal_admission_before_delivery: the worker's own earlier seal, a real owner follow-up, or no answer — delivered with admission_released=false and, where blocking reviewers approved, noted admission_close_unconfirmed); a changed subject reopens review, keeping the earlier one. Reads use the canonical budget_drive_root. The reviewer packet carries verbatim owner directives, the full contract and criteria, canonical deliverable identity, terminal child state, verification receipts, artifact references, touched-skill lifecycle facts (visibility, never a gate) and an explicit omissions manifest; a required component that cannot be assembled makes the affected actor DEGRADED, never a silently smaller prompt. The packet is SIZED against the quorum's real windows (the reviewer_window/review_synthesis.quorum_input_token_limit seam the triad and plan review use), resolved once per task so the bytes cannot drift between the binding build and the staleness rebuild. Non-core sections shed through a DISCLOSED ladder — predecessor authority envelope, trajectory tail and its results, artifact previews, agent-supplied evidence, last a diff preview that keeps the durable repo_diff_source_ref — each shed a row in omissions_manifest. A slot whose window cannot hold the rendered prompt is a typed $0 not_dispatched row while the rest of the panel reviews; a packet still overflowing after every shed stamps __immutable_core_overflow__ naming the oversized sections and refuses the panel without spending. __unresolved_partial_artifacts__ withholds packet rows only for a tool result whose exact source is genuinely source_unavailable (a retrieving row reads the source itself); a budget shed with a durable, actor-resolvable source ref is an omission, never an unresolved partial.

The packet is delivery-conditional; the FULL packet is not. Every triad row reaches the panel as configured (reviewer_slot_config.triad_delivery_slots): a packet (api_chat) row receives the assembled packet; a retrieving row receives the route-owned work order acceptance_retrieving.acceptance_retrieving_work_order writes onto ReviewRequest.slot_session_tasks — the same task-stable contract and output contract (review_execution.review_output_contract as policy["output_contract"]), absolute pointers to the task's ACTIVE workspace (review_repo_dirs_for's subject root, never the governance repo, as session_root) and its result, artifacts and receipts, and the packet in the form its delivery can use. An agent-session row gets the FULL packet (its run is unobserved by the host, so the packet is its only attested view) with the disclosure that access outside the workspace is not guaranteed and a refused read is absence of evidence, not of the artifact; a native inspection row gets the packet WITHOUT its freely degradable tail (tool-trajectory rows and artifact previews, manifested as retrieving_delivery omissions) plus the real data root (policy["native_data_root"]), as its episode reads those sources itself (host_file_read_attestation: host_observed). Reviewer evidence_refs resolve against the FULL packet on every delivery, never the rendered projection, so a retrieving row citing a real receipt is clean like a packet row. The wave budget gate prices API money only and DECIDES on one work-order send per paid row (a session row rides the owner's subscription, unpriced); that admission is the whole money rule (no rounds multiplier, no second read-only pricing pass), so a panel's cost is bounded at dispatch by the per-send wallet binding, not predicted; only a floor that does not fit is refused review_wave_budget_insufficient, and a packet row's second physical send for format repair never buys a retrieving row a second episode or session. Child-task and off-mode acceptance stay advisory and run packet rows only.

Paid identity binds that semantic subject together with substantive nonempty obligation dispositions (acceptance_paid_identity); forensic source hashes and ingress counters alone do not buy a panel. A resubmit with the same paid identity reuses its recorded verdict for free: a clean replay can authorize acceptance, a non-clean replay keeps its verdict and the identical_acceptance_refused outcome — no repeated payment, and no cosmetic edit needed when real criteria or evidence change.

The configured slots are independent actors with adaptive quorum (config.adaptive_quorum: 2-of-N for N≥3, both for N=2, a single reviewer as loud single_reviewer_no_diversity; a fewer-responded shortfall stays a loud infra quorum failure); each receives one substantive interaction on its bound route — at most two physical sends for a packet row, one bounded episode for a native row, one delegated session for a session row — and a retrieving verdict is equally authoritative. Transport status, parse status, semantic verdict, criterion support, route, quorum contribution and binding hashes stay distinct, so an unavailable or malformed response cannot masquerade as a negative judgment; a panel that refuses before any transport projects not_dispatched on every row and on the panel, and a slot released at the dispatch barrier projects awaiting for transport and parse until it settles — distinct from success, timeout and provider_transport_error, never a failure or a verdict. PASS/FAIL/DEGRADED are reviewer verdicts; the host-owned completion decision is separately accepted, revision_requested or finalized_unaccepted, written only by loop_acceptance._set_acceptance_decision. A clean quorum supplies critic approval; an informed Advisory author finish supplies separate current-author authority. Neither changes the original verdict. Material FAIL and typed unavailable outcomes reach Main before it chooses how to respond, including after the last paid panel; a no-quorum outcome with real minority findings retains that partial feedback. A terminal technical failure keeps reason=review_degraded when no permitted author completion follows.

A clean criterion is evidence-resolved, not only well argued: reviewer evidence_refs must be exact members of the packet's enumerable reference vocabulary, and a claim id resolves only through acceptance_support_refs linked to a passing host receipt for that claim. Agent-supplied, declared-intent, unattested and non-resolving sections never certify success; an OPEN plan wave binds nothing — its claims are disclosed as acceptance_claims_source='none_open_plan_wave' beside a non-binding plan_claims_exhibit inside DECLARED_INTENT_SECTIONS, so citing it never resolves and the task is distinguishable from one that never had claims. An unresolved reference keeps the actor's record for audit but removes its clean contribution (criteria_refs_unresolved). This total, fail-closed resolver is why the task cannot certify itself by echoing its expected outcome.

Actionable findings enter the durable obligation dialogue with stable identity: fix, rebut with an evidence-bearing disposition, or ask the reviewers to declare the issue unreachable or a stable disagreement. A re-raise must name an existing obligation id or is disclosed as new; a valid rebuttal retires the row, an invalid one reopens it with both positions preserved, and the reviewer's reviewer_rebuttal_response rides into the next panel's catalog so it can tell "already answered" from "never answered". Each panel receives the bounded acceptance_dialogue_history OUTSIDE the hashed evidence material — so reading the history cannot mint a fresh paid binding.

Every received review outcome permits ordinary author inspection, correction or rebuttal, including the last finite panel. loop_acceptance.merge_agent_acceptance_stance binds author_action=finish|stop to the actual outcome exposed in Main's request and its current tool/owner-directive state; only queueing feedback or predeclaring a decision is insufficient. loop_acceptance_review._finish_advisory_author permits informed Advisory completion on unchanged or revised work after criticism or disclosed review unavailability, even when paid capacity remains. The current author hash stays separate from the critic hash/verdict. Under Blocking, corrections may be retained and the attempt stopped, but disputed advancement still requires fresh reviewer approval; a later independently initiated task can resume under its own admission, never an automatic root or budget reset. stop means unfinished work and grants no approval. The earlier-revision PASS rule for final task-message delivery above remains distinct; it grants no new commit, plan or skill authority. Typed dialogue_status votes retain the reviewer's position: invalid votes abstain, material-free continue is continue_without_findings, and zero valid votes is inconclusive. A terminal opinion may close another critique of that case, but cannot suppress feedback or choose the author's stop. Auto and Required share these enforcement rules once eligible; explicit task-local response limits and real task rails remain. Cyber retains its own BIBLE P0/P3 authority without rewriting verdict, cost or custody.

Pacing predicts no review duration: a task has three host-owned rails — a deadline, a paid-cycle cap and a wallet — and a panel starts iff the cycle cap has room, the wallet can buy one work-order send per paid row, and MORE than the configured floor (OUROBOROS_ACCEPTANCE_REVIEW_EST_SEC, never below 200 s) remains above the finalization reserve; that floor applies only to a new critic. Author reaction uses ordinary remaining task time above the existing finalization reserve, budget, cancellation and round limits, without a reviewer-sized floor or adaptive multiplier. The two refusals remain review_skipped_deadline_reserve and improvement_window_inside_reserve for their respective owners. task_pacing owns both predicates (review_launch_allowed, improvement_pass_allowed); the launch rule is evaluated ONCE per panel, at loop admission before a panel is built — the paid dispatch claim (review_dispatch.task_acceptance_paid_dispatch_stamp) checks cancellation and the paid-cycle wallet only, and no other surface evaluates time. That claim is minted in the locked task_results.task_acceptance_review_accounting at first physical reviewer dispatch, and a claim without a recoverable terminal host run is UNKNOWN, never permission to re-dispatch — the double-spend fence. Once launched, a review is clamped to the owner deadline and the task ceiling with the per-send money fence: a panel the deadline cuts is a typed DEGRADED outcome, never a free skip. Disclosed residual: a panel whose evidence build consumed the margin after admission still dispatches and may be cut, and packet repair/retry sends and native rounds mean the total is NOT bounded to one floor wave. Panel durations are task_acceptance_review_timing telemetry (delivery, deliveries, native_rounds, native_rows) that no gate reads back; the structured review axis is mirrored as top-level review_status. A revision row names the pass it starts and the causes the wave recorded, never the aggregate word as its own explanation — a DEGRADED wave that still fed an improvement capsule is not the no-quorum terminal that shares the word.

A forced turn sends the round's exact tool envelope — same schemas, same server-web flag — so the provider prefix stays a cache hit; "tool-less" describes the instruction and the host, which executes no call, so a reply that still asks for a tool, with or without prose, is incomplete on every rail and degrades through the host fallback path instead of publishing as the model's final. Every forced rail uses the common terminal recorder; an eligible task with no panel run records an eligible bypass with zero runs and the rail's trigger. No forced rail can take another model round, so a dangling revision_requested terminalizes as finalized_unaccepted (revision_unavailable_on_forced_rail) on EVERY forced rail through one shared helper, naming the prior reason; accepted and finalized_unaccepted decisions are never overwritten, no bypass reason is stamped over a panel that ran, and the pair stays outside the blocked-terminal set so the objective remains best_effort. The forced children_unabsorbed rail still runs the panel for an acceptance-eligible root with a quiescent subtree, with the undispositioned-children debt in its evidence, keeping a forced delivery distinct from both clean acceptance and no-panel-warranted. A superseded panel remains an audit row with its run_count intact; pending_delivery_acceptance is a transient eligibility, never a terminal state. The root's post-task phases use the minimal root_phase_checkpoint: startup replays only a durable pending_once phase, and an indeterminate running phase is disclosed as degraded rather than replaying paid work. A deadline or late unread message retains unaccepted work rather than claiming a new subject was reviewed.

Headless finalization and workspace patch capture

A workspace task's completion compares against the captured preflight base — task-local commits stay in the delta, not git diff HEAD — and the patch is bound to task_constraint.base_sha; a moved HEAD fails closed only for self_worktree (a shared tree relies on reverse-patch verification), and an unborn repo diffs against the canonical empty tree. workspace_patch_capture.py streams the tracked binary diff plus admitted untracked files under the pure rules of workspace_patch_rules.py (5 MiB per untracked file; untracked_capture_veto_reason is the one composite the patch and the delegated-run execution snapshot both ask), excluding each vetoed entry with a per-file reason; otherwise eligible oversized or binary untracked outputs ride complete manifest+zip file artifacts, and tracked files whose old or current size exceeds 50 MiB stay in the same file-reference manifest instead of a giant Git patch. Generated output (dist/, build/) is governed by the project's own .gitignore, honoured through --exclude-standard, not by a host name rule — git-ignored files are outside the capture universe and are not listed as exclusions, because a project whose deliverable IS its build output must not have it silently dropped; a sensitive-looking untracked credential is excluded per-file and disclosed as sensitive_blocked. workspace_patch.json is written for EVERY workspace finalization, no-change and failed included, and is the truth source for CLI strict-patch (it distinguishes omitted, no-op and failed); workspace.patch exists only for ready_with_changes. A forked or empty child drive under data/state/headless_tasks/<task_id>/data is execution state: the result copies back to the canonical root, declared artifacts rebase to data/task_results/artifacts/<task_id>/, and once the canonical result is terminal a late copy-back cannot overwrite the parent-owned terminal marker or the cost/round/token fields. The startup prune (headless.prune_headless_task_drives, after prior-process custody) removes a child drive only when the canonical parent is terminal, artifact finalization is terminal, retention has elapsed, the recorded child path matches the expected directory and no child-ref promotion is pending — everything needed after deletion must cross the canonical handoff before a task is presented as settled. The capture manifest explains its acting/admission/empty-tree/capture base and current branch/upstream observations; these do not establish task authorship or change application-patch bytes. An auxiliary comparison uses explicit vcs_diff(base=..., head=...) inputs rather than guessing a target. The CLI contract stays in §1 CLI / Headless Boundary.

Owner routing verbs

promote_chat_to_task, route_to_project, steer_task and ensure_project_scope ride one receipt rail: an act succeeds only once its token-matched supervisor facts are durable in the task result, queue snapshot, annotation or mailbox authority; among several possible tasks the LLM chooses, code auto-delivers only the unambiguous one-target case, and an unconfirmed or stale receipt fails visibly instead of launching a second root. Receipts are retained per (owner message, routing token): an earlier act's receipt stays readable by its token (chat_annotation_receipt) while the message's latest row is the UI projection and the picker's liveness test. A KNOWN rejection returns rejected with its reason, never a timeout, so UNCONFIRMED keeps its one meaning: no matching receipt exists. A receipt proves admission, not completion; an unread indicator proves a visible revision, not memory isolation.

WHO is speaking is ONE host-minted fact on the event (control_routing._routing_issuer; the model has no argument): an OWNER TURN (a direct turn the owner door stamped: is_direct_chat and run_origin.owner_ingress) or a TASK speaking for itself (a promoted root inherits the stamp as ancestry, not as issuer; a root relaying an owner message it just drained; a consciousness wake-up, a Presence event or the auto-resume template on the direct lane, which nobody typed — a wake's promotes mint consciousness roots inheriting its origin, ledger category and autonomy level). An owner turn's steer travels as owner text: [Message from my human], the owner corpus, the generation bump that supersedes a reviewed answer, the room veto from the registry lane of the issuing chat (a Project room reaches its own roots, Main every host-listed root) and the owner acknowledgement. A task's own words NEVER travel as owner text: they go through the one task-message writer forward_to_worker also uses, as independent_task provenance, to any host-listed active independent root (hidden roots included, no room veto; a Presence sender only to its own binding's work, chapter 12), render as [Message from independent task <id>], enter no owner corpus (owner_source_sha256 and the acceptance premises stay the owner's), carry no attachments, and are confirmed WRITTEN or refused with the host's reason. A relay keys and publishes its acknowledgement on the owner message it drained; other task acts use their own synthetic receipt id without a chat acknowledgement. Neither receipt grants owner authority; author and target ride the receipt and one task_message_routed Logs row, and the receiver's task_message_injected row names the sender. Independent roots learn the roster from a [INDEPENDENT_ROOTS] TAIL note (peer_roster.py, ROSTER_NOTE_CAP = 40 rows shown, the cut disclosed), appended only when the roster changed and never merged into a sent row. Rows identify live direct conversations without claiming an owner initiator; messaging one is the model's call. A root may publish one bounded update_focus(text, source_ref) record, exposed in grouped notes and paginated live_roots and retained by recent_tasks on dormant results. A source_ref names a reader; update_focus answers it through that reader under the caller's registry admission (disabled_tools included) and retains the exact answer (≤256 KiB) write-once on the canonical root as focus.source_handle, a native task_source ref peers read from any drive via get_task_result(include_focus_source=True, focus_source_sha256=…) against the PHYSICAL author's record: the digest selects the immutable historical file, so no later focus or retry can substitute its evidence (the roster's retained_source); a refusing reader or changed snapshot refuses the focus (FOCUS_SOURCE_UNRESOLVED), and the roster drops a settled root's focus. Focus has authored time separate from host observation and carries no TTL or owner authority. Explicit authorized project journal/workpad reads return exactly the requested source; foreign scoped writes refuse; children/Presence keep their capability ceiling.

steer_task relays the owner's exact ingress bytes only on the turn's FIRST routing act, while it still acts on the message that started it; the window ends with the latest owner message the turn DRAINED (ToolContext.last_owner_delivery, stamped at the loop's mailbox drain; a message only written to the mailbox ends nothing) or with a landed promote/route/steer receipt already on the origin message (a refused or unconfirmed act carried nothing, so the next act still relays). Past either, the turn RELAYS its own words and its receipt is keyed on the message relayed (the drained delivery's own client id, else the synthetic agent-steer:<routing token> id), never again on an origin this turn already routed. An agent-authored steer belonging to no owner message earns its receipt under that synthetic id, confirmable through the same routing_wait poll; no chat row carries the id and the owner message's own receipt (what a later decision turn reads) stands. Each steer's mailbox entry is keyed by its routing token, so several instructions under one origin are several deliveries while a retried emit of one steer stays one.

Every promote/route refusal names its typed reason; producers holding a fact beyond the code (the workspace family, source and attachment staging, persistence failures, the project fence) add detail, the cause with its repair (self-explanatory codes such as duplicate_task_id or worker_pool_unavailable get no prose), and _persist_promote_rejection is the single durable writer. workspace_admission.workspace_repair_hint composes the workspace repair from the typed source of the refused folder; an empty Project whose genesis provisioning fails is typed workspace_provisioning_failed on the promotion path, never a fallback onto the system repo. The exact Ouroboros repository root named as workspace_root is the documented default, mapped to the workspace="none" sentinel and disclosed; a subfolder of the repository, the data drive and every /api/tasks, Presence and source caller stay refused, typed. Who is told follows who can narrate: a tool-issued act gets its receipt, error row and detail, and the model narrates; a host-issued act (host_initiated) gets ONE typed System row (task_not_started, or task_start_unconfirmed for an unconfirmed admission) in the chat the owner wrote in, from the one publication boundary _handle_promote_chat_to_task wraps around every outcome, never a bubble in Ouroboros's voice. A steering refusal whose owner act carries no owner-message receipt uses steer_not_delivered through supervisor/steering.py::_refusal_row, in the issuing chat. The Project start row is announced only once the task is enqueued. An admitted promote reports the destination the admission receipt returns, not the requested project_id/project_name (an implicit scope is re-resolved under the origin claim lock, §6 Project binding by task and by origin); a promote landing in a project other than the one the request's work already has says so: promote stays a free choice, disclosed. A promote by a root whose Swarm planning obligation is still UNMET (force_plan with no plan wave recorded, owner_hurry.unmet_force_plan_obligation) carries the obligation onto the new root and releases the promoter in the admission transaction (supervisor/plan_obligation.py), disclosed in the result; a met obligation moves nothing, and ensure_project_scope keeps it.

A routing/promote decision turn receives host-built ground truth (each project's registry working_dir plus bounded typed projections of the lane's recent ROOT results and live roots, never raw result text) through the registry row's durable last_task_result_id pointer, stamped by ROOTS only, with a bounded scan fallback; only the absent-pointer case writes back (a failing non-empty pointer is usually a split-drive result mid copy-back). A swarm child is not offered (reachable through its root) and counts as its own omission instead of evicting a root. The list is a HINT sized by runtime_limits and scoped to the lane that receives it (a pooled task holds none and names any id it read through recent_tasks/get_task_result); the door is the predicate in §10 on the result itself, never on the caller's room or the landing project, and a helper predecessor or a landing outside the predecessor's project is disclosed in the receipt (_predecessor_notes), a free choice like the second-project note. list_projects, route_to_project and the promote receipt read the project registry through the canonical data root, because a task on a forked execution drive never carries state/projects.json. A routing receipt on the SAME owner message is a fact, never a ban: a second root stays the model's judgment.

Tool capability and execution

tool_capabilities.py is the SSOT for the tool classes (core, meta, cognitive-memory, parallel-safe, stateful-browser, untruncated, capped-result, reviewed-mutative) and the child allowlists, which remain deliberately narrower principals. tool_policy.py chooses the initial capability set; ToolRegistry remains the execution authority; loop_tool_execution.py owns timeouts, concurrency, live evidence, result handling and mutative ceilings. Ordinary top-level presets share one built-in name surface: project focus changes the default target, while root policy, runtime mode, task-contract disables, credentials, resources, selected-resource bindings and delegated-child profiles narrow independently — a registered or discoverable tool is not thereby callable for a particular target. Nothing disappears silently: lazy discovery returns an explicit capability omission or CAPABILITY_UNAVAILABLE; a policy-filtered REGISTERED tool (a contract-disabled extension/MCP name included) is "hidden by policy: " (ToolRegistry.policy_hidden_reason), never "Not found". A name dispatch cannot run gets one answer (ToolRegistry._name_miss_result): its typed reason plus the callable names of the ONE namespace it addresses (tool_policy.tool_namespace), or above 12 their count and list_available_tools(namespace=…), never another namespace's or a hidden name. list_available_tools (tools/tool_discovery.py, bound by the loop) reads schemas() in every mode; loaded is residency, not availability.

Outcome classification keeps policy refusal separate from execution failure, so an expected authority boundary never becomes the task's headline failure: user_files_path_blocked, cwd_blocked and artifact_output_undeclared are typed non-failure surfaces, while a declared output that cannot be registered is the genuine artifact_output_error. The identifier register (tool_result._EXACT_IDENTIFIER_CODES) applies the same rule: a typed refusal is recorded as a refusal, never as ok. It covers the whole routing family (STEER_REJECTED/STEER_UNCONFIRMED, PROMOTE_REJECTED/PROMOTE_UNCONFIRMED, ROUTE_REJECTED/ROUTE_UNCONFIRMED, ROUTING_UNCONFIRMED, NEEDS_MANUAL_TARGET), so a promote or route that scheduled nothing is never read as a SUCCESSFUL call; a _REJECTED result carries its (reason: detail), and the refusal stays visible to the agent (is_error) without degrading the execution axis — the host refused, the agent did not fail. File discovery distinguishes a confinement refusal (an error), a failed listing (a failure) and a path holding nothing — a MISS (LIST_FILES_NOT_FOUND, a warning beside DATA_NOT_YET_CREATED) that leaves the rest of the outcome uncoloured, because the recovery scan credits only a later success on the same target. The delegate tools (delegate_start, delegate_wait, delegate_cancel, delegate_answer) reach the same register from the producer side: delegate_shared._fail writes the ok: false + host_code envelope BESIDE the domain payload (whose reason keeps its own vocabulary), so every substrate refusal (daemon, ownership, containment, cancel, subscription window) is RECORDED as a refusal, never counted as a successful call — mapped to TOOL_REPORTED_FAILURE (never degrading), a malformed call to TOOL_ARG_ERROR (degrades, feeds reflection), never to a timeout or a generic tool error — and a terminal run in state failed or cancelled is a SUCCESSFUL observation of an unsuccessful leaf.

Tool API v2 exposes neutral canonical names directly (read_file, list_files, search_code, write_file, edit_text, edit_batch, apply_patch, run_command, run_script, verify_and_record, service tools, commit_reviewed, vcs_*, schedule_subagent, manage_schedules, schedule_followup, wait_task, wait_tasks); legacy public names are neither exposed nor translated. search_code uses one basename-glob selector, including comma brace alternatives, before ripgrep or the Python fallback reads allowed files; selected-file count and content matches remain separate facts. The file tools share a path-based public ABI; payload-borne paths (edit_batch entries, apply_patch targets) miss the dispatch seam that rewrites a path ARG, so both ends canonicalize through tool_access.canonical_repo_relative_path — one normalization contract keeps a guard from judging repo/BIBLE.md while the write lands on BIBLE.md — and _ROOT_ARG_REPO_WRITE_TOOLS is the single set every repo-write fence keys on. verify_and_record (tools/verify.py) runs the declared check through the SAME pre-execution guards as the process tools — a post-execution check cannot gate a receipt already written — and appends a host-attested receipt to <drive_root>/task_results/artifacts/<task_id>/verification_receipts.jsonl (expected_match: substring by default, exact, exact_line, json_equals, bytes_equal); it reads only public task info, reports an owner-settings change as a typed note and never auto-reverts it (a revert cannot prove causation), and delegation_zero_run writes only incomplete/unknown, after the custody scan proves no open run.

Resource roots and physical file identity

Filesystem tool output is self-locating: results use canonical root:path labels and run_command/run_script echo the resolved cwd. A direct Project room selects one active physical folder for reads, writes, edits, process cwd, VCS and delegation; governance stays at system_repo, and a missing folder keeps its selected address with an availability note, never a fallback to Ouroboros source. Plain folders support ordinary file/process work and directory delegation (delegate_directory.py); Git-only snapshot and integration paths fail typed when their target lacks Git. Physical file resolution preserves the requested address: an absolute path inside any selected base normalizes to that base, an absolute path with no root binds to the permitted root holding it, an outside absolute path is refused before safe_relpath can turn it into a similarly named file, and the repo basename-prefix, lineage and delegated-artifact read redirects share the resolver guards and handlers use (ToolContext.repo_path/drive_path). A Docker workspace's mapped backend absolute address (workspace_executor.map_backend_path) is accepted only inside active_workspace; other roots and unmapped absolute paths stay confined.

user_files is the first-class root for user-visible files under the owner's home (the Ouroboros repo and runtime control plane are refused); task_drive is task-scoped scratch; artifact_store is task-scoped under data/task_results/artifacts/<task_id>/, created lazily on a write or output registration, where deliverables written through user_files or declared process outputs are copied for audit, and a rewritten user-visible file keeps its prior copy under task_results/artifact_versions/<task_id>/ (§1 tree). subagent_projects and deliverables grant read/list/search only; a read-only child reads deliverables and its lineage's (parent/root, each also on its own non-symlinked headless drive) task_drive/artifact_store, never a sibling's; a continuation also reads the task_drive/artifact_store of the one predecessor its contract names (predecessor_authority.source.task_id, one hop, read-only; a child carrying the envelope reads it too) (issue #1232); a top-level task writes the Deliverables container through user_files. Process admission keeps the original argv and the prepared physical target, and that one binding serves every other authorized destination; command words, script examples and unknown interpreter effects establish no write intent — the selected Safety Supervisor receives the full task source (§6 Safety and runtime mode) — and the post-execution shell audit is observational: no replay, rollback, interpreter or attribution proof. Invalid path-like prose yields no finding. A failed audit adds a bounded diagnostic (typed results record its exception class in output_audit_unavailable) while keeping stdout/stderr and completed exit, signal, timeout and runtime facts; it cannot invent a spawn failure or undeclared output. Computed targets stay a disclosed parser limit, not grounds for a new semantic scanner.

Credential fence and byte masking

Ordinary-mode protection of credentials, Bible and identity keys on PHYSICAL owner locations, source operands and actual VCS metadata, never on names: ordinary project names such as tokens.json, credentials.json, .nojekyll or .well-known do not identify host credential stores, and an owner's input or output is never refused on a SUFFIX or a WORD in a file name — on attachment ingest, on export/Deliverables, on a user_files mutation or in the git lanes — beside the exact-name and retained dotenv-tail rules. The enumerated physical owner stores (credential_shapes.owner_credential_locations) are the mutation-fence authority; path-selected attachment ingest checks exact credential leaves and credential/control directory components; ordinary .config, Library, settings.json and exact ~/.ssh/config mutations need no special authority; ordinary git capture adds content evidence (workspace_patch_capture.pem_private_key_reason, shared by the patch, the coop checkpoint and the attached-folder snapshot). This is not blanket protection for all credential stores: unlisted locations such as .cargo/credentials.toml, .terraform.d/credentials.tfrc.json and .kaggle/kaggle.json keep ordinary root configuration access, and children keep source visibility for ordinary auth//tokens/ paths and public PEM certificates. core_secret_paths.restricted_data_roots anchors owner/control checks on the child, canonical parent and configured admission roots, shared with vision/media; runtime secret/control entries keep their predicate even inside the project. Reads and searches share ONE byte masker in every scope: restricted repository views mask known token formats and complete PEM private-key blocks while ordinary long identifiers, hashes and source bodies pass, so search and read deliver the same bytes; owner-home output does not mask opaque runs either, and known credential bytes stay masked on delivered views.

Web access mechanisms (three distinct paths — do not conflate)

The three web paths differ in who chooses the query, which model reasons, which authority performs the fetch, and where evidence is recorded:

path reasoning model fetch authority evidence
Main-loop native search (provider server tool, attached only when the main-loop setting allows) the same solve model decides whether to search; no second model enters the scaffold provider-side citations and request counts fold into llm_usage and the task's host-attested retrieval fact (context for acceptance, never a criterion by itself); absence means only that this path recorded no search; the provider-side query is not available to the host and is not claimed as logged
web_search function tool a provider-backed call can introduce a second model and its own cost the configured web-search route or a keyless retrieval backend, executed by ToolRegistry arguments and bounded results in tools.jsonl; never native-search usage on the answering call
Browser tools (browse_page, browser_action) the main loop a local stateful Playwright session that can fetch or act on arbitrary pages arguments and result previews as tool evidence; browser state and local action semantics are not equivalent to either retrieval path

For ordinary delegated profiles the URL, route, private-range and control-plane guards apply to local-readonly and acting children alike (target policy: browser_policy.py, §6 MCP and browser-facing external tools); a Cyber-effective acting child inherits the selected access while explicit task restrictions remain. A concrete private origin is reachable only through host-established resource_policy.allowed_origins: ordinary Main/root scheduling names it in schedule_subagent, delegated and external Presence callers can only inherit or narrow it (tools/control_subagent_spec.py), and the origin normalizer rejects paths, credentials and wildcards. Of the delegated profiles only a valid acting child may run JavaScript evaluate, on its current page; acting children keep the pre-existing shell-to-loopback /ws route without WebSocket authentication. This separation is methodological authority: an evaluation or acceptance claim must name the path used.

Context fitting, retry, and compaction

config.get_context_mode() is the effective Main sizing/rendering source and config.get_owner_context_mode() the persistent owner-intent/P3 source; they differ only during the auto-Low compatibility window (context_mode_compat.py). prompts/SYSTEM.md and BIBLE.md are tier-0 and full in every projection (context_layout.TIER0_ALWAYS_FULL). In Max, ARCHITECTURE.md is full-resident for every Main task class because it is Ouroboros's capability/tools/access map (economy comes from dropping the handbook for external work, never from hiding the map), and DEVELOPMENT.md follows the active repository binding: full when the work targets Ouroboros's own body, a visible on-demand pointer for a bound external workspace or an API/CLI/scheduled external surface (explicit context_requires_development/context_requires_self_body_docs flags override the binding). In Low and Nano BOTH books are replaced by their book navigation (context_layout.book_navigation: the authored chapter introductions plus each chapter's heading index with line ranges into the PHYSICAL chapter file, because the composed book exists only in the prompt and an offset into it would address the wrong bytes), and a subagent child receives that navigation of both books in every mode, against BIBLE P1's wording (issue #1026). Disclosed residual: Low and Nano do not consult the explicit per-task context_requires_self_body_docs flag, although the per-task requirement should win in either direction (issue #1019).

The one path that trims tier-0 is the local-model overflow compactor (llm_local.py, re-exported by llm.py): when a local window cannot hold the prompt it replaces every ## section outside its keep list with an omission notice — SYSTEM.md sections (hence the prompt's load-bearing floor lives in its preamble) and the knowledge index (the one disclosed exception to the always-loaded index) — keeping Identity, Shared understanding and the dynamic memory sections (Scratchpad included), and raises LocalContextTooLargeError rather than truncate silently. Disclosed defect, not design: BIBLE.md's own ## Principle headings are not on that keep list, so the compactor also cuts the principle bodies (issue #1018).

context_fit.py renders Max, Low and Nano projections from one immutable core and measures each ordinary Main candidate on one labelled density basis (context_fit.estimate_context_prompt_tokens) against the selected route capacity, while budget reservation keeps RAW chars because over-counting money is the safe direction. Owner Nano uses its compact source view and bounded output target; Low adds an elastic 200K total-context economy target (context_budget.py); crossing either target is never a synthetic failure. With unknown capacity, owner Max gets one honest Max call while owner Low or Nano may reclaim toward its known target. Predicted pressure retains the selected document projection and may request one deficit-sized mutable-history pass; only an actual provider overflow authorizes task-local Low, and one same-route semantic recovery is permitted only when the final candidate has the same route, round and response reserve and strictly fewer context-bearing bytes. Owner mode and P3 applicability never change. context_health.build_health_invariants runs ONCE per task attempt, so the Health Invariants block is a task-start snapshot; its delegated-run lines are visible to every task but ownership-aware, so a non-owner is never handed a call that structurally refuses (§6 Delegated subagents).

context_compaction.py is a requested materializer, not a second threshold, timer or retry authority. The helper-summarized path selects a positive-reclaim prefix of completed assistant-call plus contiguous matching-result units; owner turns and malformed, interrupted or visually opaque units stay verbatim, and an unfinished Anthropic native unit is ineligible. A non-empty selection first writes an exact private checkpoint, then summarizes complete, gap-free hashed map/fold input; independent covered units may apply while a failed unit stays raw, and opaque custody never enters summarizer text. No eligible positive reclaim means no checkpoint, no summarizer call and no transcript mutation; the one route+round latch prevents repeating the same automatic pass.

Main also uses this materializer for an actor-authored working view. At the existing successful-response boundary, after retry/reprepare/fallback and before returned tools execute, compact_context.record_context_view retains the canonical source and active schemas; the acceptance observer still sees the separate send projection. compact_context can retain exact inspected unit IDs, restore checkpoint sources, or pair an authored note with an explicit keep_last_n count. IDs take precedence; note-only keeps raw units, and count-only uses the helper path. Original source bytes remain retrievable, newer owner/tool messages stay untouched, and no helper model rewrites the authored note. The completed tool boundary measures the full proposed messages, schemas and receipt through Main's existing fit arithmetic; estimates are disclosed, not new capacity or economy-target admission gates. Actual overflow retains its existing recovery. Publication updates the view and schemas together; no-op observation leaves cache identity and the live turn unchanged.

Ouroboros stores one provider-neutral, function-shaped conversation. OpenRouter omits the neutral tool_choice="auto" while building the physical request, before request-wire identity is bound, so an implicit API default cannot become an unsupported parameter under strict routing. Direct OpenAI agent/tool traffic stays on Chat Completions: openai_chat_custom.py translates only the physical copy into Chat custom tools and normalizes returned custom calls back into canonical function calls; the dialect ladder is fixed at custom/function/none (ordinals 1/2/3), only an exact dialect rejection advances a rung, the custom→function control-flow fallback is never learned as a durable dialect action, and the compatibility layer never migrates the conversation to Responses or changes the model/provider/API surface. Direct Anthropic tool turns keep a private route-bound receipt with the complete native assistant content, replayed byte-for-byte on the immediately matching same-route continuation and scrubbed on any provider, endpoint, API-surface or model change; opaque values never reach summarizer or public projections. Owner-requested none is sent as thinking.type=disabled; no guessed legacy budget_tokens mapping is invented.

request_wire_recovery.py is the single provider-neutral adaptation driver for both exception and HTTP-200 body-error shapes, keyed on one credential-free exact route/profile: only bounded set_value, drop_field and registered replace_dialect actions, at most eight per call. A learnable reactive action becomes durable — the 14-day cross-process store state/request_wire_compatibility.json — only when a semantically valid normalized response is bound to the exact settled physical-attempt capture; provider prose cannot switch routes. Disclosure needs only the identity half: the candidate a RETURNED response carried is disclosed as usage.request_wire once its capture passes the exact-identity check (request_wire_attempt.validate_wire_attempt_identity), because a ledger write that fails after the provider answered cannot unmake which request shape was sent; the finalizer's place on the successful-return path is what proves a response arrived, and money keeps its own authority in the attempt ledger, so an unsettled attempt discloses but never teaches. Disclosures aggregate in order as bounded request_wire_history — not the physical-attempt ledger; state/usage_attempts.jsonl remains the complete monetary authority. Private custom-argument receipts never enter stored history or public observability.

Retry budgets are failure-class specific: empty/incomplete responses and transient 429/5xx failures may retry the same model with deadline-bounded backoff; auth, quota, permanent bad requests and confirmed oversize fail fast; exhausted compatibility returns to the configured model fallback chain rather than becoming a second router. LLMClient keeps leading system messages authoritative and demotes later notices to visibly marked user notices.

Transport failures are classified by physical-attempt custody; each class has its own owner and rails. A REMOTE pre-dispatch failure is transport_unavailable (loop_llm_call.classify_llm_exception: released custody, $0, on a non-local provider). It takes one physical attempt per call; pacing belongs to the round-level wait episode (ouroboros/loop_transport.py), which tries the fallback chain once only when USE_LOCAL_FALLBACK makes it local (remote candidates never dial over a proven dead egress), waits (durable network_wait events, progress that keeps the idle rail alive, owner-interruptible backoff), then redials the SAME round for free. A managed task waits as long as its existing rails allow — owner deadline minus the dispatch-admission reserve, budget, Stop, absolute ceiling; never a new setting, and OUROBOROS_TRANSIENT_RETRY_MAX bounds only transient PROVIDER failures — and its exhaustion is its terminal result. The interactive class (every turn stamped direct-chat: owner chat, Presence turns, consciousness wake-ups) waits the same way but has no queue rails and ordinarily no owner deadline, so its episode is bounded by the task idle timeout (OUROBOROS_TASK_IDLE_TIMEOUT_SEC) measured from each episode's entry: it limits idle WAITING only and never cancels in-flight work, the shorter of it and an explicit deadline binds, the durable ended detail names the expired rail (interactive_wait_window_exhausted, or the deadline's own), and exhaustion emits an owner note alongside the provider-outage terminal. Closing network_wait rows exist only for cooperative exits — an external kill, Panic, crash or direct-turn hard stop leaves none — so correlate a confirmed task_done or durable terminal result; silence never proves completion.

The mid-flight class is provider_outcome_unknown: a dispatched request with no terminal provider fact keeps its unresolved monetary bound and is never resent. A NEW logical request afterwards needs a unique host-attested input, granted per caller class. Ordinary managed tasks and native API children use the upstream-observation continuation (§6 Review delivery), whose wait admits a marked new physical attempt on observed connectivity. A configured-session supervisor with exactly one live delegated leaf latches a durable unknown-provider hold (ouroboros/delegate_hold.py, closing any transport episode as network_wait ended / hold_latched) and parks in the ordinary supervised_wait, resuming in a NEW round only on a meaningful leaf wake whose receipt is that input (control wakes and an ineligible hold grant nothing; budget admission stays fail-closed on the unresolved bound); without a hold it uses the managed continuation. Interactive primary rounds get the bounded transport-death repeat below: no-effect by construction (no Host-executed tool or correspondent action can stem from the failed attempt; priced read-only provider-owned retrieval may rerun; §12), not an event retry. Every other surface — forced-final, fallback-chain candidates, review actors, safety, probes, web search, consolidation/summary/reflection, external-harness delegated runs — keeps no-resend (budget 0, or never enters call_llm_with_retry); an unknown outcome not typed as a transport death is never resent.

The transport-death repeat rail: a typed death (transport_custody.is_retryable_transport_death — httpx ReadError/WriteError/RemoteProtocolError through the explicit __cause__ chain, or a requests wrapper carrying ProtocolError/RemoteDisconnected; never a timeout, a provider status/body error, a pre-dispatch failure, a local provider or a loopback route) lets the PRIMARY main-loop round dispatch alone (_dispatch_round_model, transport_death_retries) repeat the SAME logical request at most twice per round. Each repeat is a NEW physical attempt with its own ledger row, and the earlier rows stay unresolved at their upper bound (a fully dead round reserves up to three: honest accounting over a cheaper rail). A round record (execution_id:round:round_idx) counts the repeats and names the class the terminal reports; llm_non_retryable_same_request marks only the exhaustion. A current typed finalize_now control or the deadline can refuse a not-yet-sent repeat (finalize_control_pending names the control case); the round then exits through the unknown no-resend terminal. A round holding a repeat record sends nothing but those typed-death repeats: a repeat failing with any other class ends the round on the same terminal — no transient burst, empty-response retry, compaction retry, forced-final dial, fallback chain or local-only pass — because the earlier request may still be live; only a repeat that never left the host (released custody) stays the free wait episode's to redial, and when that window closes the terminal is still provider_outcome_unknown_no_resend.

Retry rails nest, each with its own owner and bound; the table exists so the multiplication is visible in one place (a physical send is always its own execute_physical_attempt lifecycle, whichever rail asked for it):

Rail Owner Bound What it repeats
SDK max_retries client construction (net_transport.py, llm.py) 0 everywhere nothing: every physical send is visible to the ledger
request-wire recovery ladder request_wire_recovery.py inside one LLMClient.chat at most eight typed compatibility actions per call a rejected request in a corrected wire shape
transport-death repeat call_llm_with_retry, primary round dispatch only ≤2 per round, 4 s then 8 s deadline-aware backoff a dispatched request that died with a typed transport death, as a new attempt
transient burst call_llm_with_retry OUROBOROS_TRANSIENT_RETRY_MAX attempts per call (the outer ceiling of the same loop) typed 408/429/5xx and empty/incomplete responses
round redial loop_transport.py wait episode a managed task: owner deadline, budget, Stop, absolute ceiling; an interactive turn: the task idle timeout measured from episode entry (OUROBOROS_TASK_IDLE_TIMEOUT_SEC), or an explicit deadline when that is shorter a released ($0) pre-dispatch failure of the same round
served-model redo llm_substitution.py inside one chat_claudexor OUROBOROS_SERVED_MODEL_REDOS per call, and none under a pin, an admitted candidate, a spent send budget or a spent owner window a round the engine says another model answered, as a new operation. It nests inside the rails above, so a repeat granted by one of them starts a call whose redo budget begins again
review physical rail review_substrate.py, review_native_episode.py 2 sends per packet or session actor on the P3/acceptance surfaces; a native-retrieval slot carries no send count (its bounds are the episode's own, enumerated below); no rail elsewhere a released review send, never an unknown one

The cached OpenAI-compatible clients, the no-proxy per-call clients and the web-search clients share one transport factory (net_transport.py) that sets platform-guarded TCP keepalive socket options — on Linux and Darwin the idle threshold, probe interval and probe count from config.py (Darwin: TCP_KEEPALIVE, not XNU's 75 s × 8 default), on every other platform (Windows included) SO_KEEPALIVE alone — so a NAT/VPN mapping silently dropped during a long silent reasoning stretch is detected by kernel probes instead of hanging until the read timeout. When any proxy httpx would honor is configured, the cached and web-search clients skip the explicit transport (httpx env-proxy mounts require it absent). Disclosed residual: proxy-routed installs, the Anthropic-native requests lane, every non-Linux/non-Darwin platform and a handful of library clients run without keepalive tuning.

Main-loop model cognition is a separate typed in-flight fact: immediately before each exact provider-call seam the worker sends a direct supervisor started event bound to the exact attempt, execution, round, call and retry, and every terminal sends the matching fact, so a stale terminal from an earlier retry or attempt cannot clear the current row. The supervisor keeps only that process-local active row, with no elapsed-time expiry, consulted only by the idle predicate. OUROBOROS_LLM_TRANSPORT_READ_TIMEOUT_SEC (default 2700 s) is a configurable dead-socket bound, not a cognition deadline; deadline, budget, cancellation and the absolute ceiling remain independent hard axes.

OpenAI-compatible response choices keep their outer finish_reason as the bounded observational usage fact response_finish_reason; it never enters canonical history and changes neither the empty-response classifier nor retry policy. The trusted provider canary accepts a schema-valid native tool call even when the assistant also returns text (length/hash recorded as warning telemetry); a malformed native argument, invalid schema or missing call remains a red contract failure — no prose parser, salvage, provider hop or unbounded retry.

Prompt caching is stable-first: governance and task-stable contracts precede mutable evidence, review builders disclose the stable/dynamic boundary, untrusted payloads stay outside the governance cache block, and context_fit.seal_task_transcript owns the single message-side breakpoint (the task-contract boundary until a rolling tool-result seal qualifies, migrated in the same call so the four-breakpoint cap can never drop it). OpenRouter derives its sticky session from stable policy/model and a copied first-user content projection without cache or host-only metadata, so moving a seal in either direction does not change the session. Explicit review affinity and provider-reroute rotation retain their precedence. Provider-specific cache hints are sent only where supported and receive one exact retry without the rejected hint; rejection evidence is durable and route-specific. Gateway response-cache recovery is narrower and reactive: only a main-loop provider_incomplete_response arms a fresh-response request for later attempts, and only the generic openai-compatible route renders LiteLLM's extra_body.cache.no-cache control — direct providers and OpenRouter never receive it, and no URL/model heuristic exists.

Vision and local image evidence

analyze_screenshot and vlm_query are bounded secondary-model calls through LLMClient.vision_query; view_image attaches a local image natively. Send-time image routing (vision_routing.py: inline vision, captions or placeholders) works on a per-send copy and never mutates canonical history. Model-wait reprepare starts from that canonical source, preserves original images and prepares the new route's send projection; a typed invalid native continuation retains its existing reset without replacing the source with captions. Image payloads are validated, capped/downscaled and confined to readable roots derived from the Tool API policy matrix plus the protected-artifact rule. vision.attach_local_image_to_context is the single attachment seam for explicit view_image and the typed auto_attach_image opt-in — same durable copy, trust boundary and live-image eviction budget; auto-attachment failure is non-fatal. The loader's own refusals are typed TOOL_ARG_ERROR results (a policy owner's refusal keeps that owner's typed marker), so a known failure is recorded as one. Vision/local-media tools are not web tools and may be withheld by the task contract. The uploads and skill-output roots resolve through the task's canonical data owner and count as independent image admission before user-home confinement, per-path secret/owner-state/project-store/protected-artifact checks still applied; an admitted image is copied into the task's uploads/views before same-round attachment, so the original and its retained copy both stay available to explicit readers.

Background consciousness and Evolution

Background Consciousness is Ouroboros taking a turn in its own Main chat when nobody wrote to it. There is no second mode: a wake-up is an ORDINARY Main direct turn (_is_direct_chat) with the same system prompt, context builder, tool loop, chat log, activity block and usage ledger, so consciousness structurally cannot count its context, its money or its tools differently from Main. Only three things are new: an alarm clock, a wake-up MESSAGE that stands where the owner's words would be (prompts/CONSCIOUSNESS.md rendered by consciousness_wake.py from existing readers, never silently cut), and the three owner settings beside the existing interval bounds. The alarm (consciousness.py) owns no thread: the supervisor pass calls tick(now), whose ordered typed outcomes are disabled → a live wake → a live owner direct turn → not yet due → the rolling-24h allowance (an unreadable ledger is the disclosed skip allowance_unknown, never a silent block) → an owner chat must be bound → launch through supervisor.workers.handle_wake_direct. The next wake is last finish + interval (set_next_wakeup or the default, clamped into the owner's [min, max]; a failed or unadmitted wake doubles it up to max), notify(reason) pulls it forward to one shared floor, an owner message never wakes it, and a wake's own finish never re-arms it. Version awareness stays ambient: every task's and wake-up's Runtime context carries official_update (running version versus the official target at the last fetching check, plus the update letter, update_letter.py); no check forces a wake, and mentioning an update to the owner is the mind's judgment.

The rendered wake message puts the fresh cause and settled facts before outstanding cards. Failed/cancelled or host-authored Presence terminals stay visible even on the direct-turn lane, with their stored outcome, task pointer and deferred work reference; ordinary owner direct turns and the wake's own turn stay excluded. This exposes an external failure without starting another wake or prescribing recovery: consciousness inspects the task and actual delivery through existing tools. Cards in open and expired_terminal retain late-answer semantics, explicit state, age and bounded preview. consciousness_last_wake_at preserves the factual boundary across restart; the best-effort in-memory trigger may fall back to heartbeat.

Observe retains a narrow positive path. consciousness_authority.OBSERVE_ARGUMENT_NARROWED is the set of mutating-table names Observe still SEES, because each has a genuine read-only or own-work use that withholding the name would destroy along with the mutation: schedule_subagent and delegate_start for read-only research, write_file/edit_text/journal_write/workpad_write for its own notes, and cancel_task for the children it started. What such a name may be ASKED to do is decided by one predicate, observe_argument_refusal, which the dispatch guard consults directly — no fail-open wrapper, because a policy that cannot be consulted must not read as permission. It refuses a write root outside task_drive/artifact_store, a mutative write_surface/may_mutate, and a delegated session asked for workspace_write or a payload selector; an OMITTED access is the read-only shape (_derive_authority pins it) and stays allowed. cancel_task is the one name whose narrowing needs durable lineage, not arguments, so it rides the tool's EXISTING own-child custody check (join_ledger._cancel_task), the same seam a delegated task meets: Observe may stop the children it started, never the owner's running work. A read-only child may delegate further read-only descendants under the inherited depth cap but never passes a mutating delegation budget onward (delegation_may_mutate). manage_schedules (tools/followup.py, beside schedule_followup — one module for the two tools over one table) is argument-aware, not name-wide: the root turn may list and change, a delegated task may only list, and a Presence conversation gets neither — Presence authority is not expanded here.

Three facts distinguish a wake-up from an owner's turn, and all three ride its metadata. Its ORIGIN is metadata.initiator == "consciousness", inherited by everything the wake starts: it labels every surface; the wake speaks to other tasks as a task ([Message from independent task <id>]), never as the owner, because no owner door stamped it. Its AUTHORITY is the owner's autonomy level (metadata.consciousness_autonomy, owned by consciousness_authority.py): Observe researches, keeps internal memory/project notes, stops the children it started, inspects and governs existing schedules through the audited narrow surface, and may delegate read-only research, while the dispatch guard refuses owner-file, source, skill/settings and publication writes; Act (the default) is everything the runtime mode allows except editing Ouroboros's own code and prompts, evolution, restart and settings; Full adds evolution. A level has two structural consequences — an exception list in task_contract.disabled_tools and a per-task runtime_mode_cap of light below Full — plus, at Observe, the argument-level narrowing of the names it keeps (observe_argument_refusal); for a consciousness-origin task that list binds at DISPATCH ONLY, so the wake's tool schemas and cached prompt prefix stay byte-identical to an owner turn's while a withheld call is refused when it is made: the typed dispatch refusal is the mechanism, and the wake message names the level. The stricter of install mode and cap binds at the repo-mutation, protected-path, start_service and shell-write gates, at the acting-child write surface and subagent patch integration, and in a started task's Runtime block, so the owner's level holds even on a Cyber Pro install; a capped tree is refused a self_worktree child and a system-repo patch integration in every install mode, the mutative-subagent toggle included. Its MONEY is one rolling 24-hour allowance covering the wake-ups and every root tree they started, read off the usage ledger (consciousness_allowance.py) by the alarm before it wakes and by the ONE admission door (supervisor/queue.py::enqueue_task) before any work it asks for starts — no separate "consciousness task" type, only the existing tools through the existing door; a Presence cycle a wake starts (Act and above) is an independent root outside this allowance. What is left of the allowance (capped by the owner's per-task cap) is the wake tree's graceful ceiling (root_cost_ceiling_usd, inherited by its members): the in-task stop lands a planning margin before it while the ledger fence stays at the owner's per-task cap, so a wake never dies before its first call; a remainder at or below that margin counts as exhausted.

An active campaign owns an explicit objective, campaign id, transaction and task claim (supervisor/evolution_lifecycle.py); evolution_mode_enabled is only its scheduling projection. Dispatch, review, commit, publication and restart revalidate that exact authority: a restored row without a live uncommitted claim is cancelled, a reviewed commit binds to the claim by exact SHA before publication, and if authority changes after commit the commit moves to a private inspection ref and any attempt-created tag leaves the normal namespace, without resetting concurrent index/worktree edits; exact terminal replay resumes only incomplete effects, never double-counting. The restart-verify claim (state/pending_restart_verify.json; one writer helper serves the supervisor's evolution restart and the agent's restart tool) is written whether the supervisor restarts automatically — OUROBOROS_EVOLUTION_AUTO_RESTART off skips only the restart — so the owner's manual restart verifies the cycle by exact claim rather than by the weaker markerless reconcile.

Campaign cleanup is deterministic and custody-aware: a no-op or abandoned cycle may restore the transaction base while preserving dirty/ahead work in recorded stash or local refs, and cleanup is skipped when another task, a live test or the operator kill-switch makes reset unsafe. A byte-identical diff whose last terminal is a review-verdict block is refused for free from the first block; changing the diff, or a rebuttal whose content hash is new to the streak, creates a new paid reviewable case, and the shared cycle cap bounds paid triad+scope cycles per root task. Checkpoint and outcome rows (evolution_checkpoints.py, append-only) preserve git/memory identity, cost, rounds and explicit omissions so the promotion loop learns from failed and absorbed cycles. An agent-requested restart first drains heartbeat-fresh running tasks up to its configured bound, then fails closed rather than cut another task silently.

Post-task evolution is owner-gated and default-off (post_task_evolution.py): the worker may recommend one backlog item by writing a durable request — never enqueuing or enabling evolution itself, and never from an evolution, deep-self-review or subagent task — and only the supervisor's idle tick converts it into one normal campaign cycle. Its CLOSED-objective projection distinguishes absorbed/abandoned objectives from this attempt's typed review-cap or unavailable-review stop: a temporary review block does not permanently prohibit reconsideration. Retained task sources remain discoverable, without scheduling a retry or resetting campaign counters. An owner stop closes the campaign, clears queued requests, persists the evolution_owner_stopped flag, and cannot be autonomously reversed. Evolution is hard-blocked in light runtime mode at the owner and post-task start entry points, again at idle campaign enqueue, and again at assignment (§5); it uses the normal task/review path in advanced or pro.

Loop self-checkpoints remain plain user-message reminders: a second structured, tool-less reflection protocol in the hot loop produced unusable records and destroyed cache continuity. The reminder's recent-call list carries each call's recorded outcome and, for a failed call, the first line of what the tool answered — host facts only, no threshold, no repetition verdict and no stop. It is built from llm_trace.tool_calls, so name, arguments and outcome come from ONE record: joining the visible transcript to the trace by call id cannot be made sound, because providers reuse ids (GigaChat answers call_0 every round) while compaction removes whole units from anywhere, so a surviving call was printed carrying an evicted namesake's error — by id, and still by (id, tool) when both calls were the same tool. Without a trace the transcript's calls render with no outcome, not a borrowed one. Durable learning belongs to the post-task reflection flow (§6 Post-task reflection).

Safety and runtime mode

Tool dispatch uses the registered capability, the explicit task/resource contract and one prepared physical target. Explicit task resources and the actual Git target are host facts, distinct from any semantic judgment about the call; this is not an OS sandbox.

Runtime mode is effective Access and a self-modification ladder. Light refuses mutation of the Ouroboros repository and control plane (LIGHT_MODE_BLOCKED; skill payloads under their own root stay editable) while user deliverables outside them remain buildable; Advanced evolves the application layer but not the protected set owned by runtime_mode_policy.py — SAFETY_CRITICAL_PATHS, frozen contracts, release/managed invariants, one policy shared by the registry, git tools and gateway guards, refusing typed CORE_PROTECTION_BLOCKED; Pro may write that set, announced by CORE_PATCH_NOTICE, and its commit still passes the normal triad + scope gate. runtime_mode_policy.mode_has_unrestricted_agency derives Cyber authority from effective Access; saved next-boot settings never rewrite a running task or a prior request. In Cyber Pro, internal review and Safety findings are advice under BIBLE P0/P3: configured enforcement, independent assessments and the decision to proceed remain distinct facts. Real unavailable resources, malformed executable forms and external errors remain failures, not invented success, and an explicit read-only assignment keeps its meaning.

Process admission preserves the original argv and the resolved cwd. Command words, quoted source examples and unknown interpreter effects are not proof of a forbidden action: shell_parse.py and tools/write_shape.py supply observed targets and syntax facts, never permission judgments. The Safety Supervisor (safety.py) judges each call from the complete retained initial and later owner directives, the task contract, the task constraint and the prepared resource target; the bounded newest-first recent-conversation view (_SAFETY_CONTEXT_CHAR_BUDGET, omission marker reserved inside the budget) supplements that source and never replaces it. Coverage is OUROBOROS_SAFETY_MODE Full/Light/Off over each tool's TOOL_POLICY entry (skip, check, check_conditional; an unknown tool falls to DEFAULT_POLICY = check): policy-skip tools and established safe conditional subjects make no call in any mode, Light additionally skips conditional checks, Off skips every Supervisor call, and every waved-through check leaves a durable safety_mode_skip event. full → light → off is a decreasing coverage ladder, so outside Cyber Pro a downward step is owner-only (POST /api/owner/safety-mode). A selected assessment runs on the light-model route. No second consent store, appeal pass or repeated execution exists.

Foreground commands, scripts and run-kind verification accept env_from_settings through the same Settings-selection authority as service startup. tools/process_facts.py resolves references once after admission and masks diagnostics before result and receipt persistence. Foreground execution keeps its inherited environment; Docker transports target values through inert host aliases, so target PATH or DOCKER_HOST cannot reconfigure the host client. Verification matches raw output before masking. Python selection uses the effective child environment without a version probe; wrapper and non-local backend resolution remain unknown when not established. The existing Node health check runs after gates with the selected PATH.

Safety Supervisor outcomes

The first line of a Supervisor result is the typed outcome. SAFE passes silently; SUSPICIOUS allows the call with SAFETY_WARNING; DANGEROUS, an unrecognised status, a failed check and an answer still unparseable after one repair retry (safety_parse_retry, safety_parse_failed) block with SAFETY_VIOLATION. Outside Cyber that warning/refusal stands; Cyber records a negative or unavailable assessment as advice (SAFETY_ADVICE, durable safety_advisory event) without manufacturing SAFE. A provider 429 is an infrastructure fact about the Supervisor, not a verdict about the call: after one deadline-capped retry the call is blocked with the typed non-verdict SAFETY_UNAVAILABLE, which tells the agent to retry the same call rather than reword it, and a confirmed storm arms a process-local latch (_SAFETY_STORM_COOLDOWN_SEC) that answers in-window checks without a provider call, because re-probing from the highest-frequency light-model consumer would amplify the storm it reports; each is a durable safety_check_rate_limited event, while a structured insufficient-quota refusal stays PERMANENT and blocks as a verdict. The serialized subject has its own 250,000-character budget (_SAFETY_SUBJECT_CHAR_BUDGET, rendered ensure_ascii=False): an over-budget subject is refused fail-closed with SAFETY_SUBJECT_TOO_LARGE_BLOCKED plus a durable safety_subject_too_large event — never truncated, because anything past a cut would run unreviewed. Only the no-backend cases fail open with SAFETY_WARNING (no provider key and no local lane; a remote key that does not cover the light model's provider, with no local lane; a local runtime chosen as FALLBACK that fails or is rate-limited), so a misconfigured runtime does not hard-block every unknown tool; an explicitly primary local runtime (USE_LOCAL_LIGHT) that fails is a real SAFETY_VIOLATION.

Registry dispatch passes its prepared binding into start_service; a direct handler consults the same Supervisor only when no prepared registry binding was supplied. Each admitted process effect executes once; post-execution settings/repository observations annotate the original result and never repeat or roll back the operation.

Outside Cyber Pro, read-only shell git is allowed everywhere; mutating shell git only when its resolved target lies outside the Ouroboros system repository and runtime data drives (git_shell_policy.py; network-disabled tasks still fence network git). Acting self_worktree children remain read-only because patch capture requires an unmoved HEAD. git init/git clone are judged by destination, relative destinations and path-valued retargeting flags included; resolve_shell_cwd supplies the same canonical cwd to target checks and execution, and command text is not an independent runtime-file permission classifier. The generic Tool API VCS family (vcs_status, vcs_diff, vcs_pull_ff, vcs_restore, vcs_revert) defaults to root=active_workspace and accepts explicit root=system_repo; protected Ouroboros path names constrain generic restore/revert only on the explicit system target, so a project's own BIBLE.md or contracts/ remains ordinary project content, while preflight_review, commit_reviewed/vcs_commit_reviewed, vcs_rollback and promotion remain system-repository lifecycles. vcs_diff without refs retains unstaged/staged behavior; base compares a resolved local tree to the worktree or index, while base plus head compares two local trees. Results name requested refs, exact tree IDs and comparison kind, without fetching or implying a merge-base comparison.

Outside Cyber Pro, a task contract may declare resource_policy.protected_artifacts[] as execute-only black boxes (protected_artifacts.py): registry guards allow the declared execution but refuse reads, copies, hashes, static inspection and trace/debug wrappers over those paths. Light-mode cognitive writes are redirected to update_identity, update_scratchpad or knowledge_write instead of raw memory-file edits; a corrected cognitive redirect is advisory, while an ignored user-file root correction remains a blocking deliverable failure.

Review delivery

review_execution_projection.review_actor_progress_text says in one sentence per row who the reviewer is (the frozen route display name), what it was asked for at start, and at settlement whether it answered, how it ran (the observed harness, model and account, or that this was not reported), or why it did not (the engine's reported sentence, quoted; a spent window's reset instant); a $0 refusal reads as not sent. Requested effort/model is not provider-observed truth; duplicate model slots keep distinct identities in the record. No global last-execution record supplies task progress authority; no code, id or custody state reaches the row.

Review delivery has two closed route kinds in review_execution.py: api_chat and agent_session; vendor and harness names are route targets, not new kinds. A slot is bound to one immutable route before its first send and never falls back to another transport — a route that cannot deliver raises on its own slot (ReviewRouteUnavailable). The API executor memoizes the assembled review messages, so its durable prompt record and its at-most-two physical sends share the same bytes; one logical interaction may spend one bounded second send on transport or empty-output recovery, never while a dispatched outcome is unknown (task acceptance may also spend it on malformed-format repair). That two-send rail belongs to the PACKET api row: a native tool-round episode carries no send count — its bounds are the ones below — and neither retrieving delivery is sent a format-repair resend, because their executors canonicalize their own answers. The hosted-agent executor starts one read-only delegated session through the shared Claudexor nanny (route-owned instructions and retrieval pointers, its own tools, no assembled API pack); extraction canonicalizes the collected transcript and never launches a second session, and custody, cancellation, full-artifact recovery and settlement stay on the delegated transport contract (§6 Delegated subagents). Mutating external coding work uses the delegated subagent path. The Claude Agent SDK gateway is retired; two residues stay on purpose: reviewer_slot_config.py migrates legacy Claude-SDK advisory targets at parse time (_migrate_sdk_advisory_target; an unmapped target is typed legacy_claude_sdk_target_unmapped), and the legacy filename tools/claude_advisory_review.py hosts the live generic advisory gate.

A session reads with its harness's own tools. review_session_reads.session_read_facts derives diagnostic harness_observed ranges from completed calls in attempts/*/events.jsonl, using the same manifest and range vocabulary as native reads with weaker provenance: an inferred shell range does not attest the exact bytes the model saw. Simple standalone cat/nl, sed -n, head/tail and bounded file reads have modeled extents; compound commands, pipelines, redirections and searches do not. File reads without enough range information, including ordinary Cursor read/view_file, stay unobserved, as do missing or unparsable journals. These limits never invalidate the answer, remove a response from quorum or trigger another review. Changed candidate bytes remain a source-gap diagnostic; publication binding is independently checked by the commit gate.

Native tool-round episode

review_native_episode.py (NativeToolRoundReviewExecutor) delivers every retrieving API row: configured-subagent reviewers on any surface, plus every scope, deep-review and advisory API row. Scope declares ReviewSlot.native_retrieval_override through review_substrate.scope_delivery_rows; deep review and advisory bind the episode directly. One episode is one logical review attempt: chat(tools=…) rounds against a fresh instance-local inspection registry (read_file, list_files, search_code, query_code, vcs_status, vcs_diff, compact_context; the local_readonly_subagent constraint, network off) until the reviewer answers. There is no round cap (P13; the retired round-cap key: §11.4). The bounds are four. The SEND bound review_native_transcript_bound is min(owner ceiling OUROBOROS_REVIEW_NATIVE_MAX_TRANSCRIPT_CHARS, the reviewer window's density-calibrated capacity in chars) measured as the wire size of the next send (serialized messages plus tool schemas, recomputed after every append) and enforced before every provider call. The owner deadline_at (passed by commit triad, scope, advisory and task acceptance) binds together with the slot's independent logical window: each send's transport timeout is clamped to the window remainder and a spent window refuses before dispatch. Each inspection call adds the loop's per-tool timeout, narrowed by the inherited dispatch deadline; past it the call is abandoned as native_tool_abandoned on an error receipt. The paid ledger prices every physical send as its own attempt. Mandatory reading never raises the send bound: a surface's policy["native_mandatory_read_chars"] declares the whole required surface, which the reviewer reads through successive working views (compact_context authors a new view; the data plane is the host's empty scratch unless the surface opts into policy["native_data_root"], and only the scratch is removed at episode end); when one view cannot hold the declaration the shortfall is disclosed typed as native_mandatory_read_disclosure = native_multiple_windows_required beside the declared chars (native_mandatory_read_bound is a separate advisory estimate — never a floor, a grant or a refusal). A surface may also declare policy["native_required_sources"]; exact host-observed intervals and byte-identical inline delivery fold into native_read_coverage. Missing or unobserved reading remains diagnostic evidence beside the verdict, never an episode stop, lost quorum seat, commit block or automatic paid repeat (§6 Scope review by retrieval).

Typed ends: a bound that leaves no room below the first send is a pre-send refusal (native_bound_below_first_send), as is a registry without projectable inspection schemas (native_inspection_unavailable); both are not_dispatched ($0) in review custody and leave the receipt keys (resolved model/provider) empty, so the public execution wire never mints a native run that never sent. The host posts one [EPISODE_BUDGET] landing notice per working view at NATIVE_LANDING_FRACTION (80%) of the bound, and every tool result is clamped to the room left below it so no single read can jump past the bound; exhaustion is a typed refusal (native_transcript_cap_exceeded) for verdict shapes and a disclosed native_incomplete product for the report shape. A round carrying tool calls but no well-formed one and no prose is the malformed end native_round_without_progress (a progress floor, never a ceiling); an empty answer stays the episode's honest end on the empty-response rail. An episode that ran paid rounds is dispatched (settled) in custody whatever exception ended it — deadline, transport or the paid ledger (budget_exhausted). Every end — delivered, refused or errored — leaves one review_native_episode custody row and its facts on the actor usage (failure_custody for the error actor). The default output contract of a retrieving row follows the surface's output shape (triad_review.default_output_contract) unless the surface hands over its own policy["output_contract"].

Native read evidence rides the actor usage as native_tool_receipts (outcome = executed/refused/error/withheld, host_file_read_attestation: host_observed); an executed read_file receipt carries start_line, end_line, total_lines, eof, opened_path and opened_root from the reader's own stamp (ctx.last_read_view in tools/core_file_tools.py; tests/test_native_tool_round_executor.py pins the writers), so a bounded read is distinguished from full coverage without trusting the model's path spelling — a missing stamp proves no extent. Every native end publishes native_rounds, native_tool_calls, native_transcript_chars (last physical send), native_transcript_bound, native_transcript_refused_chars when a next send did not fit, native_landing_notified (posted) versus native_landing_sent (physically sent), native_end_reason, native_custody_row, native_view_changes and native_history_source; a non-delivering or incomplete end after any assistant round also keeps native_terminal_round, a bounded, redacted, structurally valid JSON copy of the envelope and tool results that ended the episode, because receipts alone cannot reconstruct it. These facts survive on failed actors through failure_custody, deadline, ledger and transport ends included.

Late completion and typed refusals

Late review completion is retained in the operation-addressed prompt/response CAS (review_substrate.py, review_custody.py): the response manifest stamps the complete producer outcome with the original task/root, attempt, slot/route, operation, subject, contract and roster/epoch binding, so exact reconciliation verifies that source and runs the existing surface reducer without a second paid attempt; pending, unreadable, mismatched and unknown outcomes keep custody and their full source references. Plan waves carry their dispatched set, health epoch and operation ids through the same path; free replay spends no new cycle. Skill Review reserves every chunk digest/retry key and slot operation id in review_job.review_wave before paid dispatch and keeps that binding in terminal history; the next authorized call may record review_resume_of only for the immediately preceding unsuperseded wave with identical task/root/attempt, skill/state roots, group, content, contract, rebuttal and chunks, and then reaggregates complete CAS at the original wave id without another paid stamp, even at the cycle cap. Missing or partial producers stay pending, unstarted chunks acquire no verdict authority from their reserved ids, and a new task/root, changed material/contract/rebuttal, superseding lifecycle or explicit cancellation cannot inherit the old wave. Late completion alone never revives a job or enables a skill: review-state persistence, grants/dependencies and enablement keep their lifecycle guards.

Terminal task delivery and later inspection derive a bounded host notice from the delegated-custody audit after cleanup — confirmed terminal cancellations, live runs, unresolved invocation ids and undisposed patches stay distinct — beside unchanged model text and continuation narrative, so a later cancellation never rewrites historical authorship or the answer hash. A valid acceptance PASS with partial, missing or rejected criteria remains a parseable contributing verdict; only the clean/applied host decision can authorize objective completion.

Malformed structured slot configuration is refused at save and is a typed loud review-time failure on every surface — commit triad, scope, advisory, plan, skill review, deep self-review (deep_review_slot() raises; run_deep_self_review returns the typed deep_self_review_unavailable result instead of a report) and task acceptance (a typed DEGRADED panel, reviewer_slot_config_invalid); no surface silently chooses the opposite route or a default panel. Technical failure and commit permission are separate facts: producers keep the failure origin (context, delivery, format or window authority) beside the original status, source and findings, the shared physical exception projection keeps budget/admission and authority refusals distinct, and what each enforcement mode permits after such a failure is the protocol's rule (DEVELOPMENT "Review & Commit Protocol"); neither path changes quorum, and custody, Stop/deadline, ownership and candidate binding stay independent. Contributor slot evidence retains configured-subagent identity while its stored wire uses the mutually exclusive reference-or-route form, preserving native retrieval and its API budget admission.

Usage ledger substrate vs. accounting policy

The startup seal audit selects its bounded manifest set before reading live attempt rows; missing live ids share one lazily validated archive-id set for the whole audit (an UNKNOWN result included) while forward checks continue independently, and the next audit refreshes that observation through the unchanged archive reader and its integrity/cache checks — no per-manifest archive traversal, no change to monetary authority or archive retention.

usage_ledger.py owns the durable append-only physical-attempt ledger (cross-process locking, sequence and transition validation, append+fsync, replay, loud tail quarantine); usage_accounting.py is the one-way policy layer above it (pricing, reservations, settlement, scopes, budget fences, imports, projections, admission), and the substrate never imports policy. The boundary means a pricing or budget-policy change cannot redefine valid ledger storage, and a locking or repair change cannot silently change what an attempt costs; compatibility events, state mirrors, task fields and UI projections may carry attempt ids and derived totals but never become a second charge source. _usage_response.py is the one NORMALIZER of a provider's usage block for accounting (its importers are usage_accounting.py and loop_llm_call.py), where a zero-usage body error settles at a confirmed $0 to release the reservation, so a provider storm cannot manufacture phantom budget exhaustion. Not the only READER of that block: every provider adapter reads the raw usage dict for its own response envelope, so a "consolidation" there would centralize a read that was never centralized.

Caller-owned subscription model calls

claudexor::<source>=<model> selects a model operation in the Ouroboros-owned engine, not an Agent Run: LLMClient dispatches sync and async calls through llm_claudexor.py before the OpenAI-compatible filtering/retries, and Ouroboros keeps its SYSTEM, BIBLE, canonical messages, tool selection and execution. The raw-model adapter is Codex; connected Claude/Cursor and other harnesses keep their Agent capabilities, and direct API keys keep their existing routes. Every main-loop execution of one install sends the same Codex cache key (llm_claudexor.cache_key_for_model: per data root and model, projected to prompt_cache_key and the session_id header), so a new task, child or consciousness cycle is served the governance prefix its predecessors already cached; per-execution turn states ride nativeContinuation unchanged under that shared session.

ClaudexorGateway uses purpose-bound byte uploads, one idempotent operation ID, status/result reads, cancellation and explicit result acknowledgement. Claudexor performs at most one generation per operation: a lost local HTTP reply rejoins the same operation, and an unknown provider outcome never authorizes another generation. It shares its managed profile with Agent sessions — the adapter reads current access credentials while the official CLI keeps refresh ownership — and no shared external engine, relay process, copied operator auth store or public OpenAI-compatible endpoint. Metadata helpers use read_owned_gateway (owned-only discovery and handshake; they never install or start the engine); actual model calls and explicit lifecycle actions use ensure_owned_gateway.

Failure evidence: before creating an operation the host discovers the catalog's captureFailureEvidence=true query and freezes that choice for the invocation, same-key rejoin after control loss included (older catalogs keep the legacy request; the serving runtime pin remains the independent adoption boundary). A supporting engine puts every received failed-response byte and processing exception — incomplete responses and HTTP refusal bodies too — in the private result, which the observability CAS retains before ACK; usage and events carry only compact problem context, and a received prefix never claims the missing provider tail. A coherent provider terminal whose message cannot be normalized keeps its outcome, route and usage: the operation fails response_rejected, the host settles the reported usage and raises the non-repeatable stream_rejected class — no unknown-outcome recovery, no healthy account marked failed. ClaudexorModelError.display_message adds sanitized typed vendorCode/parameter details to the error event and terminal preview only after classification; unknown outcomes expose no provider details.

Money and bytes: physical attempts stay in the usage ledger, where exact zero cash, known charges, estimates and unknown cash remain distinct — a subscription is not evidence of zero cost. Claudexor removes terminal request bytes and acknowledged result bytes, an unacknowledged ready result expires after 30 days, and compact operation receipts survive body deletion, so an expired read cannot start a new generation. Model-purpose resources are not valid Agent attachments. Private continuation envelopes stay in canonical history but out of public/summarizer projections; their reuse is bound to the actual source/model/profile/account identity.

Image capability comes from the selected role/account's model catalog, not the global model-id overlay; main send preparation, browser image attachment, captions and registered VLM tools share that reader. Missing metadata stays unknown — image input is retained for the real call rather than silently omitted on a cold engine — while a confirmed text-only model keeps the caption/capability-gap behavior, and text-only messages and image-off mode never read the image catalog.

The send copy omits foreign provider envelope metadata such as stop_reason (canonical history retains it), preserves a nonempty refusal as assistant text even when the provider supplied no ordinary content, and normalizes all-text tool results so moving the message-side cache boundary never rewrites bytes already sent (non-text blocks stay intact). For the same reason the transcript is append-only between the sends of one execution: compaction is the one sanctioned rewrite, stamped at its seams, and ouroboros/transcript_prefix.py records every other break (a context-fit reprojection after a real overflow included) as a prompt_prefix_break checkpoint instead of blocking the send. Only the exact model_request_invalid create refusal proves validation failed before command admission; a generic HTTP error or failed status read cannot prove non-dispatch.

The live turn slot

Beside assistant-level history the engine keeps ONE live turn per model client session, and its opaque token is transport, not content: it belongs to the caller that is running, not to any stored message. So the caller owns a single mutable slot (llm_claudexor.ModelTurnState) that rides the ordinary LLMClient.chat parameters to the engine boundary, the only writer — a dispatched durable result replaces the value, a released invalid_continuation repair clears it together with message envelopes; other non-dispatched, unknown or legacy-shaped outcomes preserve it. Holding state never licenses another generation, and silence about a turn never ends one. WHY a slot rather than the transcript: the live turn survives compaction that drops the messages and outlives a body that fails after its headers, so the last stored assistant envelope would rebuild the wrong lifetime. One run_llm_loop invocation is one turn (rounds, owner steering, acceptance follow-up, forced finalization, quota waits and reprepares continue it; a next loop and a cold restart start empty), and a consciousness wake-up is one such ordinary turn keyed by its own task id. A dispatch that leaves this transport for a direct API or local route ends the turn at the caller and never revives it on return; the engine alone compares route identity. Opting in requires a serving engine at config.CLAUDEXOR_MODEL_TURN_STATE_MIN_VERSION, read from the last SUCCESSFUL handshake — a failed probe never un-proves it, so concurrent status polling of that singleton cannot change the shape a running caller sends — and an older or not-yet-observed engine keeps the legacy request shape instead of risking a schema refusal. The token is never usage, an event, a progress note or a task card; mechanism in the llm_claudexor.py docstring.

Roles, accounts and context

Model roles own account pins and context assertions through OUROBOROS_MODEL_ACCOUNTS and OUROBOROS_MODEL_CONTEXT_WINDOWS; ordered fallback entries keep their original ordinal, reviewers and configured actors reuse their route credential fields, and identical Main/Light or reviewer model strings never imply identical roles or accounts. Auto uses the largest advertised window of the exact account route, not CLI compaction thresholds; a manual value is a sizing assertion, not a provider unlock or reviewer authority; unknown capacity stays unknown, and input size, response reservation and capacity are separate. Each operation records its submitted options beside the engine's applied options (an absent applied report stays explicitly unknown), and the first changed reasoning effort of each model in a task produces one typed owner-visible notice. An account change rebinds preparation before another physical send; ordinary sends, prospective wrap-up payloads and forced final replies share one acting-role/account binding, and prospective subscription accounting uses the dispatch serializer, not an OpenAI-compatible approximation. The final header identifies the responding model/account; a paid review that ends incomplete retains its full report and custody with that classification, never switches to another delivery or inherits the previous route's window.

Host-authored sampling defaults stay separate from explicit parameters: review requests carry default_temperature as a hint each effective API route resolves at dispatch, raw model operations use their provider default, and an explicit temperature (zero included) is preserved and may be refused. Routes without structured output use the existing safety text-JSON parser, repair and refusal path — no new safety bypass. Ordinary confirmed provider errors keep each helper's previous handling, while quota, owner interruption and unknown physical custody propagate without becoming a helper result or verdict.

Quota and auth waits

model_wait.py holds confirmed quota/auth waits inside the live call. Typed owner actions use the existing mailbox and the /api/decisions family model_wait:<task>:<wait>; task-result rows are projections, not restartable stack checkpoints or a second attempt ledger, and metadata polling makes no generation. Auto may prefer the last successful same-route account unless it got a status-null or typed per-subject refusal this execution. An HTTP vendor quota refusal of a named account is re-asked with the same payload, no preference; a pre-dispatch verdict never is; only its poolCause proves the pool spent; an account chosen twice stops rotation, claiming nothing. Pin never rotates. A route may also answer with ANOTHER model than the one requested, with no error and nothing in the account's quota: the engine states that as a typed fact on its completed result, and the transport never accepts such an answer as the round. It is retained and acknowledged like any other answer, so the horizon it would have run under stays reconstructible, but it adopts no turn state, runs no tool call, and is replaced by a NEW operation that names no account — the engine, which deprioritises that account for a while, alone chooses where every redo lands. OUROBOROS_SERVED_MODEL_REDOS bounds those redos; a pinned account names one account, an already admitted candidate is bound to exact bytes, and an actor with no send left on a bounded budget would otherwise read that rail's own budget error instead of this refusal, and a spent owner window buys no further send, so none of those is asked again. An outcome nobody knows is never a substitution: the transport raises it before the rail sees a result, so the existing unknown-custody fence, not this budget, governs what happens next. Each discarded generation's retention and acknowledgement state travels with its durable row, so a failed one is visible rather than assumed. A discarded generation teaches the density witness nothing, because a witness belongs to the model that produced it. Once no redo is available the typed model_substituted refusal reaches the caller, which is an ordinary cross-model fallback trigger for a root and a typed failure for an exact-route child; it never cools the requested model, because one substituting account says nothing about the model. The first substitution of each model in a task also becomes one timeline row of that task's card. The rule covers this transport only: an Agent-session delegation is a separate engine capability with its own model evidence. Fallback keeps one attempt per model and a 120-second cooldown; on quota each configured route runs before the owner wait, opened from the retained refusal (no generation first); an intermediate fallback's own refusal remains the wait/retry route if a later fallback fails differently; nothing sleeps to a reset; inline Presence never waits (typed refusal; safe retry proof: §12). A temporary switch affects only the wait; Settings persistence, acceptance and task application remain distinct. Fallback Local remains a group-wide setting: a one-role persistence request may change model/account only when Local is unchanged, a different Local choice is task-only, and permanent group changes remain in Models. The current task-attempt identity selects actionable wait rows; a failed mailbox delivery retries the same accepted request without authoring a new one; terminal mailbox cleanup waits for post-task synthesis and pending accepted attachment promotion to settle.

ouroboros/gateway/task_model_wait.py owns model-wait decision effects and the bounded settings-writer receipts; task_decision.answer_decision returns (status, payload) to both the Web and authenticated Host Service transports (they supply the installation root — no synthetic HTTP request or parallel owner registry); supervisor/task_model_wait.py projects worker events and quota clocks through the existing supervisor, and ouroboros/gateway/decision_contracts.py owns the decision-family vocabulary. Quota pauses use union duration across simultaneous waits and do not consume internal execution time, while explicit calendar deadlines remain fixed; the worker and project-write ownership stay held (a one-worker queue waits), completed tools and reviewer results remain on the same live stack, closing a browser does not stop the wait, and stopping the Ouroboros process ends this continuation guarantee. Image tools retain the tracked VLM child: vision_process.py carries typed errors, physical capture, cancellation and operation checkpoints across that process boundary, and a lost child result remains unknown, never free.

A consciousness wake-up is an ordinary direct turn (handle_wake_direct): the same direct-turn registration, model wait, Stop and bounded interactive transport-death repeats as an owner's turn, an ordinary task result, and no background owner, locked memory, pause or stop event of its own. The alarm clock only decides when the next turn starts.

Streams and transport waits

Main remote completions stream (llm_stream.py), assembled inside the physical send closure before accounting settles: a complete stream yields the same normalized response shape as JSON; a partial stream never yields a usable tool call or answer, and its exact wire bytes and partial assembly stay in private CAS through the physical_stream manifest and attempt id. The assembler is strict about completeness, tolerant about form: on the Chat path identity scalars and metadata keep their first value, index gaps are forgiven and every forgiven fact is disclosed in the stream receipt (anomalies); the native Messages assembler rejects post-terminal shape (non-contiguous blocks, an unsigned thinking block, an incomplete tool block). A complete-but-unusable stream settles with its usage as a provider error (stream_rejected); unknown-outcome continuation is reserved for a stream that never reached its terminal frame (a lost socket, a malformed mid-stream native frame), and a complete response with absent final usage still has unknown money. Comments/pings do not define cognitive deadlines; every recovery candidate checks inherited calendar and quota-adjusted execution bounds before reservation, after preparation and at dispatch, and HTTP phase bounds stay distinct from the overall logical wait.

An ordinary managed task whose provider outcome becomes unknown stays in the transport-wait episode (loop_transport.py); the old attempt and its unreported cost remain unknown. After non-generating upstream observation, a user-role [SYSTEM NOTICE] supplies explicit recovery input for one new physical attempt under the existing budget, cancellation, owner deadline and absolute ceiling. An attempt that fails again, is unknown after dispatch or released before it — and a free redial that crosses dispatch and dies unknown — returns to the same episode: the backoff keeps growing (4→60 s), one redial is counted per wait iteration, an unknown repeat re-arms the probe's freshness bound at that latest failure and refreshes custody, one continued row names the transition (continuation_outcome_unknown, continuation_transport_unavailable, redial_outcome_unknown), and no cap on continuations — budget, deadline and Stop are the only stops. Finished tools are retained. Configured-session supervisors prefer their live-leaf hold (delegate_hold.py), then use this same recovery for their own model; reading a result or disposing a patch is not a condition for cognition. Direct turns keep their separate contract; manual Restart/Panic gains no resume authority. Recovery is proved by provider provenance, never by a local handshake: direct remote endpoints are observed through HEAD under the same no-proxy policy, reusing the ordinary connection allowance for every socket phase narrowed by the owner remainder and holding no cognitive in-flight lease; subscription metadata qualifies only when the catalog reports generic provenance="provider_http" with an original observedAt after wait entry for the selected source/model and effective profile/account fingerprint — cached reuse never advances that timestamp, and a static or pre-outage catalog cannot prove recovery. Capability discovery remains authoritative, not a new core provider table. Loss of the Claudexor control connection keeps reading the same accepted model operation across endpoint rediscovery, without creating another operation.

Delegated subagents (Claudexor transport + the nanny)

Ordinary delegation requests no extra engine review panel; new ordinary runs on Claudexor 3.9.8+ default to no panel — a version boundary of the behavior, not a second release pin. The start receipt's engine_version is the handshaken serving version, distinct from the release pin; engine review outcome, execution success, the parent's integration decision and Ouroboros review gates remain separate facts. Delegated snapshot capture stays relative to its recorded baseline: committed bytes can be captured with a disclosed head_moved, the instruction still forbids committing, and the distinct self_worktree unchanged-HEAD check is preserved.

Children coordinate through tree_note and tree_read; read-only children can read/list the project-scoped knowledge store without writing its index; only the parent may use override_delegation_constraint, and a review_requested note carries an exact evidence reference/hash and wakes the parent without starting a paid cycle. Read-only and acting children alike hold the descendant-scoped forward_to_worker, peek_task, cancel_task and discard_child_result controls, and recursive delegation never widens filesystem, budget, depth, deadline, commit or owner authority. forward_to_worker also reaches the child's parent or a sibling (same parent and root) as peer_task provenance, its prefix naming the stamped relation (sibling|parent): context-only, relay refused, 8000-char bound. delegation_budget governs descendants (may_delegate, may_fan_out, additive depth provenance; a free-form intent note is never authority); persisted admission facts outrank later Settings changes, and a lower permitted depth is reported capability_reduced, never a silent flat tree.

Registry. configured_subagents.py owns OUROBOROS_SUBAGENTS: strict {enabled, items}, at most ten rows with hidden subagent_id, route-handle name (§3 Available subagents), any-language recommended_use, normalized API/session route, optional effort, account pin, access and row enabled. Session access defaults to full or the owner's workspace_write; snapshots lacking access retain workspace_write. Legacy env inputs fail closed. Row enabled defaults true, rejects nonbooleans and serializes only false, preserving fingerprints/receipts. Off rows retain configuration but leave catalog, alternatives, legacy matching, local autostart and reviewer resolution. Explicit selection returns subagent_disabled, distinct from global subagents_disabled and live unavailability. Tools take handle or stored id; roster edits refuse twins; records name their own engine. Captured snapshots never consult live settings. Descriptions guide the LLM, never host ranking/keyword routing; exact selection starts that route or refuses without substitution.

Scheduling. schedule_subagent requires subagent_id, a focused objective and expected_output; the remaining public fields describe child-local context, constraints, capability needs, write surface, narrower deadline, delegation budget and acceptance claims. access=inherit (default or omitted) preserves the configured session access; readonly|workspace_write can only lower it. API-model rows ignore this field with a result disclosure; write_surface stays the read/write authority. There is no model-visible lane/executor axis and no public effort override: the selected row is the complete execution choice, lineage/bounds/route/budget are host-derived, and omission never inherits the parent's acceptance claims. subagent_runtime.select_subagent_snapshot copies an immutable snapshot of the exact enabled row into the child task: an api_model row becomes an ordinary recursive API child on that exact model/effort, an agent_session row an ordinary recursive Ouroboros nanny on that exact external route. Burst cash (each sibling launched before the first sibling's first response pays its own prefix write on cache-write-priced routes) is disclosed in the tool description as an affordance, never scheduled by the host; so is an exchange of addressed turns, stated in objective/constraints (never a contract field or host section; no stage, topology or participant count fixed). The old lane/executor resolver serves only old durable records: for historical lane envelopes schedule_subagent reports the requested lane only (effective facts remain on the dispatched child record), and a task carrying configured_subagent goes straight to subagent_runtime, so legacy policy cannot reinterpret an active selection.

Waiting on children. wait_task/get_task_result return the full single-child handoff (verification receipts red/masked-first, exact omitted count); wait_tasks stays batch-compact: task_id, status, child_result_sha256, outcome_axes, result, terminal_host_notice when present, trace_summary, capability_delta when the child has something to disclose, duplicate_of, plus the nullable cost-finality pair accounted_upper_bound_usd/cost_final and, when the child's envelope carries one, execution_evidence (§11.1); the retired cost_usd spelling is tolerated only when reading stored rows, never emitted. Both use task_status.SETTLED_STATUSES, and a pending cancellation is the typed cancel_state: "pending" projection, never completion. A batch wait that expires with children still running returns the typed wait_expired_with_live_children block (live child ids, requested window, clamp ceiling): facts only, no advisory text, no host floor on the next window (how long to wait is the mind's call, BIBLE P13); an id this tree never minted stays unknown_task_id, never a live child. An optional known_result_sha256 / known_result_sha256_by_task compares the join-ledger semantic result identity (join_ledger.py; CHILD_RESULT_STALE): an exact match returns result_unchanged without the repeated text while status, costs, outcome and custody facts stay; no persistent seen-state is inferred. await_messages(timeout_sec) holds the worker slot until an unread mailbox entry or its bounded window, delivers nothing and takes no lease: the in-flight tool lease spares the idle rail and its close stamps progress.

What a delegated run costs. Claudexor reports the amount in summary.spendUsd and its exactness in summary.spendEstimated; delegate_custody.disclosed_spend is the single reader of the pair, so the ledger row and the payload the nanny relays cannot tell different stories. Runs ask authPreference: subscription explicitly, because the engine default falls back to a paid key invisibly. Four cases:

disclosure ledger finality
disclosed settled zero 0.0 (subscription_session row) cost_final=true — the proven free session
disclosed settled charge the amount final
estimated amount the amount cost_final=false — an estimated zero is not a proven free session
undisclosed cost_usd: null, increments unknown_unmetered drops cost_final for the projection

An undisclosed spend contributes 0.0 to accounted_usd — inventing a conservative bound would fabricate a number the harness never gave (BIBLE P1) — so a TOTAL_BUDGET fence cannot stop spend it was never told about; the honest consequence is loss of finality, not a guessed charge.

What a delegated run READ. Harnesses count input tokens with incompatible conventions, so settle_run carries the engine's own normalized split, summary.inputTokenUsage, onto the subscription_session row as the optional input_token_usage object: total_tokens, cache_read_tokens, cache_write_tokens, each a nonnegative integer or null for unknown. record_subscription_session is the one validator and writer: all three keys are required together, and a partial, extra-keyed, negative or fractional object is unknown as a WHOLE, not repaired field by field, because a repaired counter reads exactly like a measured one. An engine that reports nothing leaves the row unchanged, and the object stays OUT of the idempotent row identity, so an engine that starts reporting never rewrites or duplicates a session already settled without it; the legacy prompt_tokens/cached_tokens axes keep their own meanings, compaction never folds these idempotency-bearing rows, and direct physical attempts take no copy.

The nanny model. An agent_session subagent is an ordinary recursive task-tree child acting as a nanny supervising at most one active bounded external leaf. The task node keeps lineage, authority, deadline, budget, acceptance, cancellation and descendants; the harness process stays a non-recursive tool leaf: session rows never flatten the task tree into harness processes. The nanny is the host: verification receipts stay host-authored, and harness output is a claim to check, never proof. A nanny-to-nanny chain through schedule_subagent is the host-attested realization of a nested subscription swarm, each level one metered supervisor task plus one free harness run; schedule_subagent.requested_depth is the typed ABSOLUTE request counted from the root, telemetry that never narrows the configured caps, while the legacy depth_remaining envelope keeps its narrowing semantics (both disclosed on the contract), and the root's swarm_efficiency.depth block reports requested, permitted and achieved. A nanny terminal names BOTH routes as separate facts: the host's own model_execution carries its model, provider and last typed error (last_llm_error_kind), each terminal_runs row the leaf's model, profile and selected actor, so a nanny that died on its own lane is never read as a verdict about the leaf, and a leaf half whose reconciliation was never persisted is omitted, not guessed. Harness-agnostic by construction: the row holds an opaque Claudexor target, Ouroboros asks for an access profile derived from task authority and lets Claudexor choose the mechanism; no harness-name branch selects a capability or fallback in core dispatch, and login-wire asymmetries stay presentation adapters in gateway/claudexor_accounts.py.

Transport. gateways/claudexor.py is pure transport (descriptor read, /v2 handshake, the config.CLAUDEXOR_MIN_VERSION 3.2.0 floor). The daemon bearer token grants the entire /v2 surface, so it never leaves this module, and the HTTP client runs trust_env=False so an ambient proxy cannot intercept the loopback control plane. Production starts obtain a handshaken owned gateway from claudexor_daemon.ensure_owned_gateway — exact reviewed engine/Node pins (claudexor_runtime.py; the reviewed pin IS the next-spawn selection — no mutable current pointer, no background updater — and OUROBOROS_CLAUDEXOR_BIN is the explicit operator override), never a PATH install. Keeping lifecycle above transport keeps account status side-effect-free and harness mechanisms out of Ouroboros (claudexor_daemon.py/claudexor_runtime.py docstrings; stop path, spawn latch and typed start failures: §9). Native child processes are contained by env token (process_containment.py, OURO_PROC_CONTAINER_*), because a surviving descendant can become invisible to parent-child traversal once its controller exits; an alive-or-undeterminable member is an honest hard-block answer, never a kill guarantee.

Custody is durable, because the run is not ours to kill. A delegated run lives inside the daemon and survives our worker, so the AUTHORITY is the durable delegate_run_* custody rows (delegate_run_started and friends) on the canonical/budget root (ouroboros/delegate_custody.py), read by every replay-class consumer through custody_rows — the process-local memo in delegate_custody_memo.py (the _usage_rows_memo shape: an ordered (st_dev, st_ino, consumed, st_mtime_ns) fingerprint of the folded chain prefix, only appended bytes folded later, a refold on any doubt, a bypass and never a cache while the chain is unreadable; the rows stay the authority, each process pays one cold fold, a durable compact projection is the next step); one compact incident projection, <drive_root>/logs/containment_faults.jsonl, exists because the event log grows without bound and a tail-bounded scan can bury an unresolved fault. A lookup answers OWNED, FOREIGN or UNKNOWN — collapsing UNKNOWN into "not yours" made a restarted owner indistinguishable from an intruder. An ABSENT custody log is a positively established clean state; an EXISTING-but-unreadable one audits as delegated_run_state_unknown:custody_log_unreadable, never cleanly reconciled. Every INTENDED start mints a fresh per-intention invocation UUID as the wire Idempotency-Key; the content hash is only the LOOKUP identity — a content-stable wire key would hand a deliberate re-run the finished old run — and reuse happens only by explicit token (pending_invocation_id/retry_of, replaying the STORED canonical body under the SAME key). The complete replay envelope lives unredacted in the existing private observability CAS, written before delegate_run_start_requested; the event carries request_ref and prompt_chars, keeping a large reviewer packet out of every custody scan (the memo swaps a legacy inline body for a locator delegate_pending.request_body re-reads). delegate_pending.request_body resolves legacy inline first, then verifies the CAS reference for both invocation readers. JSON values and the canonical request digest stay unchanged, including the thread fields needed to reproduce its wire projection; a missing/corrupt blob leaves the request unknown while preserving pending custody identity for start blockers, terminal audits and snapshot retention; recovery retains that invocation without POSTing. Known pending review tokens take the same missing-request path, while absent or definitely refused skill-review records keep their existing fresh-retry behavior. Pending scans resolve only surviving invocations, while legacy event/archive bytes remain untouched. reconcile_orphaned_runs visits every open run whose owning task left the live set and settles the terminal ones, but it CANCELS only behind a deliberate verdict: a durable owner result that is readable, truly terminal and finished by the task's own decision. A custody row carries its owner's kind, and task association confers no lifecycle authority: a run a review surface registered (RunCustody.review_owned, durable source under the review substrate) belongs to its panel, bounded by the slot window and its own maxSeconds, so the sweep spares it unless the owner task was itself cancelled — a left_live row names the panel, and a task that consciously finished under a running acceptance panel keeps its reviewer alive; a pending review invocation is likewise retained, never re-posted by the generic recovery, because the review substrate owns its rejoin. The owner-cancel kill boundary STATES the verdict it is about to write, because it audits custody before that write, so an owner cancellation stops the paid run at the boundary, not at the next sweep. A provider or transport death, a worker crash, a missing or unreadable result all SPARE the run, left live with a durable left_live reconciliation row: an undignified nanny death must not kill a healthy paid run, and "unknown" is exactly that case. maxSeconds is the damage limitation for a spared orphan, never custody. SETTLED is published before registration retirement, after the ledger row lands; settlement and registration retirement are separate durable duties, a failed retirement stays replayable on project_owned for the later sweep, and writing settled over a suppressed ledger failure would turn a lock timeout into a permanent leak. A start whose row did not land reports started_uncustodied: no supervision, no replacement, until the original run is proven absent or terminal.

No terminal or cancel claim without a verified receipt. delegate_cancel returns confirmed (read back terminal), requested, failed or containment_fault_run_may_still_be_live; the last two hold a durable CRITICAL containment fault until a receipt or settlement clears it — an overpowered run that may still be alive is an incident, not a reassuring string — and the state read decides, so a refused control is never a verdict about the RUN. One daemon_says_absent predicate decides everywhere that a 404 is the daemon ANSWERING that the resource is gone (scoped to the daemon that answered), never a failure to find out; custody closes such a run delegate_run_closed_absent (unreachable, not settled), inventing no terminal detail, usage or spend. Results are delivered, not severed: delegate_wait stages the whole terminal detail atomically under task_drive/delegated_runs/<run>.json with a typed output_delivery block, and cut fields are renamed *_preview so a partial read of head-truncated JSON fails loudly instead of looking like an answer.

Four nanny verbs (tools/delegate.py): delegate_start, delegate_wait, delegate_cancel, delegate_answer. There is no fake hurry: the transport supports cancel and answers, not in-place steering. delegate_shared._fail adds ok: false and host_code without replacing the payload; _owned_run retains its durable OWNED/FOREIGN/UNKNOWN boundary for wait/cancel/answer. Start selects an exact session subagent_id (API ids refused), or recovery-only retry_of. subagent_bootstrap starts the configured nanny's snapshotted leaf through delegate_start(prompt="") before its first model round, without waiting (configured_session_started); the model chooses delegate_wait. Recovery adoption precedes zero-run/unknown fences because a prior run may be live; a fence-wake outranks terminalization.

A definite refusal needs a typed cause, no custody handle and either producer definitely_unrun or _DEFINITE_UNRUN_REASONS; it ends UNRUN at $0. Ambiguity wakes the model: uncertain live work cannot become a zero-spend terminal. Unsupported readonly/payload geometry (directory_execution_unavailable) and unregistered roots refuse before start with durable start-blocked evidence; the nanny registers first and retires owned registrations at settlement (delegate_registration_policy). The payload selector names authority through fresh ResolvedResourceBinding(skill_payload.write), never grants it. Actor replacement/zero-run fences exclude review_owned runs: those belong to review custody, while settlement, money, containment and registration still see all runs. New delegation readers join the consumer matrix in tests/test_custody_owner_kinds.py.

Finite leaf continuation (delegate_continuation.py): engine maxSeconds cancellation settles as outcome_reason=wall_clock_exceeded, replayed as terminal_reason. Explicit continue_from=<run_id> admits a NEW run/cap/key only for the caller's settled run with that confirmed cause, full output read to EOF, captured patch applied/rejected without ambiguity, and unchanged actor/route/access/mode/isolation/authority target. Stop/Panic, owner deadline, user cancel, failure and unknown endings refuse. continuation_of and the host block retain predecessor/cause/disposition and the prior run's recorded access, so a mutating run that captured no patch is described as having written in place (its effects already on the target), never as read-only; the model authors remaining work in prompt. No session state transfers; delegate_recovery.NO_RESUME_CAUSES stays unchanged.

Execution evidence. delegate_evidence.task_execution_evidence projects the custody rows read-side; delegate_start_attempted counts blocked and uncustodied attempts too, so a refused-but-obedient nanny is never disclosed as nudge-ignoring (nanny_nudge_recorded). applied_access_profiles is read off SETTLED rows only (empty = no receipt disclosed it, never "no access"); acceptance_patch_dispositions is the bounded section over delegate_run_patch_verdict rows (cap 20 with the exact omitted count, unreviewed_delegated_apply headline) whose absence means NO disposition was recorded, never "reviewed clean"; an unreadable custody log is the typed evidence_read_failed marker, never an empty-therefore-clean section.

Work orders. The compiler (subagent_work_order.py) sends the entire chosen assignment, preserving context and instruction roles without an arbitrary host-size cutoff: direct starts carry the normalized host contract once in instructions (delegate_start_instructions.py; the separately hashed coordination appendix is absent from the host pre-start) and the chosen assignment separately in prompt, and coordination context stays complete. Real native/HTTP limits return their actual failure with the original input and any pending invocation retained. Exact-source readers stay optional capabilities; incomplete source coverage never authorizes a terminal PASS or apply (delegate_source_coverage.py). Legacy partial starts keep their exact renderer digest, source-interval validation and stored-body retry; removing partial-start production certifies no incomplete old run and launches no duplicate after ambiguous dispatch.

Supervision. delegate_wait is model-visible as an event-only sleep, not a caller-sized poll: delegate_supervision.supervised_wait renews bounded transport windows in host code at zero LLM calls, and only a meaningful event (settlement, interaction, fault, addressed message, child signal, control, recovery judgment) becomes a coalesced durable wake, replayed across worker interruption until acknowledged; deadline, ceiling, budget and cancellation stay outer bounds. Every receipt and wake carries one host-rendered coordination_context (intent, time remaining, root-tree spend, active descendants, remaining paid acceptance capacity) (facts for LLM judgment, never thresholds), observed READ-ONLY, so a metadata-poor task reports time.state = "not_set" instead of latching an anchor from a poll. Polling writes nothing of its own beyond the canonical usage-ledger reader's bounded maintenance, identical for every reader: the torn-tail quarantine after a SINGLE crash mid-append (one verbatim row in state/usage_attempts.quarantine.jsonl, one usage_ledger_tail_quarantined event), the empty state/ lock directory on a never-initialized root, and owner-aware usage_attempts.lock recovery (§1 Platform substrate); an absent ledger is that reader's known-zero; a crash inside the quarantine repair can leave the sink torn (a disclosed residual). A requested future inspection (checkpoint_after_sec) wakes once and is consumed by any earlier real event: no cadence, stall classifier or hidden polling. An observation the transport could not complete is a quiet renewal carrying its typed reason, never a refusal that spends a model round: observation_read_timeout is our own read bound expiring against a live daemon (quiet, nothing more), daemon_unreachable a socket that carried no answer, the only half worth an owner line, delivered once per episode with one recovery line on the supervising task's own progress surface. The beat stays three seconds with no backoff, durable counter or outage latch, and deadline, ceiling, budget and cancellation still cut a long unobserved stretch. On a wake the nanny holds its full tool surface and the parent's captured model/effort. Nanny economics are structurally quiet: only BASELINE_RESET_TOOLS (delegate_start/schedule_subagent) reset the burn baseline, and coordination verbs never buy metered silence (nanny_pacing.py).

A run's question is the nanny's to answer. Supervision wakes immediately with typed status="waiting_on_user" on a NEW pendingInteractions entry instead of burning the engine's answer timeout in dead polling; answer keys are echoed verbatim into delegate_answer (custody-gated like cancel, relaying the engine's typed outcomes: subscription_window_exhausted carries reset_at; a transport death or 5xx is delivery_unknown), and delivered interaction ids are acknowledged only after transcript injection, so a question neither re-triggers a round nor disappears across recovery; every waiting payload states continuation: same_session (an answer resumes THIS session; each turn paid). A question above the nanny's authority escalates to the nearest live ancestor. The codex lane has no mid-run channel: a terminal with outcome_facts.reason=input_required (continuation: new_physical_run) is answered by a plain new delegate_start(subagent_id=..., prompt=...), never the engine's rerun verb, which would start a run outside this task's custody trail.

Recovery is exact and cause-specific, not generic task resurrection. Only a proven non-signal worker crash (or a planned self-restart's typed handoff) lets a successor adopt the exact run or pending invocation, before any LLM call or new start; anything ambiguous returns typed recovery-required, never a duplicate mutator; owner restart, panic, signals, deadlines and cancellation are explicit no-resume causes (delegate_recovery.py; delegate_pending.py replays a pending invocation under its original key and body). When no physical run exists and none can be started, the model may finish only with a typed zero-run receipt (verify_and_record) after custody proves nothing is open; a session actor's terminal is CLEAN only through its own physical leaf or that receipt ("completed direct child ⇒ clean" does not exist), and physical_leaf_not_started rides the terminal projection as an incomplete/unknown execution axis. An unknown-provider hold (delegate_hold.py) parks the nanny in the supervised wait until the leaf wakes and never resends.

Configured dispatch is exact (subagent_runtime.resolve_configured_actor_dispatch, resolved at dispatch, the last moment route availability is current): a bad choice returns a typed reason, current alternatives and any reset time with host_fallback: false; the host never ranks alternatives, converts session work to API work or picks the first healthy row. The WHY is recorded so no redesign undoes it: selecting an agent_session row IS the parent's typed LLM decision that this work executes on the harness, the FLOOR the host hardcodes (truth, money, authorship), while topology, decomposition and supervision judgment remain the model's CEILING (BIBLE P13/P5). At completion actual_substrate derives only from custody rows; a typed startup refusal never authorizes work on a different substrate: another session route or API fallback is an explicit LLM-selected start.

Route health. subagents.route_health is the one health reader for dispatcher, nanny and review slots, and it does not guess admission: the doctor status describes only the DEFAULT credential store while real accounts live in the engine's credential-profile pool, so admission is the engine's, whose start POST answers an empty pool with its own typed refusal (credential_pool_exhausted + earliest reset) at $0; the gateway types it as a timer-healing window, scheduled on its reset like a spent subscription window, only when resetsAt is dated; an undated (structural) pool stays a plain ClaudexorUnavailable. enabled is honored as route_disabled (the owner's switch, not an observation); the other structural refusals are route_not_in_capability_catalog, an access-profile mismatch, engine_rejects_delegated_marker and positively proven quota exhaustion. A fully-used ratio needs a valid future reset before it can refuse a route; an incomplete reading delegates admission to the engine and never certifies available quota; explicit active cooldowns bind. Review slots inherit "the engine decides", never a silent fallback onto metered API spend: on the auto lane a dead daemon keeps the native fallback with its visible marker, while "daemon alive, pool empty" is discovered at the engine and disclosed. The model sees a semi-stable facts-only catalog (model_visible_subagent_catalog); invalid, list-disabled or all-rows-off configuration projects nothing; dispatch is the live check.

Historical helper observations use state/subagent_last_delegation.json, owned by subagent_history.py. Task finalization joins configured API actors to their own attempts, preserving requested pins and observed routes/accounts; fallback success never certifies the failed route, nor does code/test failure condemn a provider. Sparse pre-response failures stay in usage/history/raw events; served-call trace references retain received response-backed calls, including incomplete responses. Foreground/recovered session terminals and typed start failures carry original effort/processing through RunCustody/STARTED/START_FAILED facts, without history-only archive reads; missing options stay unknown, distinct from captured defaults. Pre-invocation bootstrap refusals keep their source; reviewers keep separate receipts. Known occurrence and first observation differ: replay neither refreshes age nor replaces newer dated evidence; definite evidence may resolve an unknown start. MAX_CONFIGURED_SUBAGENTS bounds latest actor rows, with old single receipts readable. Dynamic Runtime context and owner UI consume history without changing catalog, admission, quota, dispatch or ranking.

Read-only and mutating session rows share one nanny transport. The only difference is the run shape, whose ONE owner is subagents.delegated_run_shape, asked one question — is this an acting child? — by tools/delegate._derive_authority (live ToolContext) and by resolve_configured_actor_dispatch (configured snapshot); a shape re-derived in two places drifts unsafely in exactly one branch.

task authority access mode isolation execution.delegated
acting subagent (valid write surface) captured full or workspace_write agent live true
Ordinary root with a validated external workspace or a selected Project room captured full or workspace_write agent live true
top-level task selecting an exact skill payload (root="skill_payload") captured full or workspace_write agent live true
anything else, including a fail-closed subagent or an invocation lowered to readonly readonly ask envelope (default) not sent

WHERE a mutating run's changes are destined is the second, separate record — the host-derived mutation authority (tools/delegate_integration._mutation_authority / _payload_mutation_authority, re-exported by tools/delegate), never model-supplied:

source target_root derivation capture_mode
acting_constraint the child's own task_constraint.write_root, required to equal the genuinely ACTIVE workspace root delegated_snapshot
external_workspace_root the root task's validated external workspace or selected Project room delegated_snapshot
skill_payload the exact payload the fresh skill_payload.write binding resolved delegated_snapshot
readonly ordinary active root (nothing to write) none

A payload target gets a standalone private Git snapshot (subagent_worktrees.provision_payload_snapshot): the live payload is never initialized as Git, and capture trusts nothing under the child-writable snapshot's .git. Disposition (integrate_payload_patch) applies a live, index-free git apply in no-repository mode under a whole-payload content-hash CAS (drift = typed conflict; identical content = idempotent applied); GIT_CEILING_DIRECTORIES is pinned at the payload's resolved PARENT (git still searches the ceiling entry itself, and an ancestor Git worktree above the runtime data root could otherwise make git skip every hunk at rc=0), and reserved paths refuse the WHOLE apply as blocked_reserved_paths with the candidate preserved. The post-apply outcome set is complete: a live loader hash equal to the recorded RESULT hash is the success; a hash equal to the recorded BASELINE hash with a non-empty touched set is a provable non-mutation that RESOLVES the apply intent (typed INTEGRATE_APPLY_NO_OP, the apply_no_op arm: no success, no disposition, no reconcile queued, retry lane open); anything else is the ambiguous mismatch, whose intent stays PENDING and whose reconcile marker IS queued (the payload did mutate). A successful apply queues the extension reconcile (request_extension_reconcile) and the skill's review goes STALE pending fresh skill_preflight/skill_review; a run whose ONLY change is a mode flip is already refused at CAPTURE time as unreviewable_metadata_change, so a live hash still equal to the baseline means nothing was written.

A mutating run normally executes in a PRIVATE EXECUTION SNAPSHOT (metered children keep sharing the tree; their patches integrate through integrate_subagent_patch: sha256-bound, 3-way --index, protected-path gated, genesis refused, coop_already_in_tree a no-op). At delegate_start the host snapshots tracked, staged and eligible untracked state, deciding sensitive/credential vetoes BEFORE hashing (git add -A would put secrets such as .env in the shared object database). Git's binary verdict for the whole untracked inventory comes from one index-versus-worktree git diff --numstat over a scratch index (two when empty files need their attribute verdict; workspace_patch_capture.untracked_binary_verdicts, shared with patch capture; a failed batch falls back to the per-file verdict with a warning), never one process per file by design. The machine-wide worktree ops lock (subagent_worktrees._ops_lock) guards SHARED metadata only — the registry file, a target's .git/worktrees and its baseline pin — held twice, briefly (row FIRST, then the ref, then worktree add --no-checkout, so a crash after the row is GC-nameable); the acting self_worktree lane is split the same way (worktree add --no-checkout -b + branch under the lock, reset --hard --no-recurse-submodules and deletion outside it — the body's own post-checkout hook no longer fires at provision either), and only the boot-time prune_orphans sweep and the millisecond genesis git init still do their work under it; listing, classifying, hashing, populating (reset --hard --no-recurse-submodules: what worktree add runs internally, minus the target's post-checkout hook), copying the source's exact bytes over the checkout and one update-index re-recording their stat (a CRLF-converting checkout otherwise leaves every such file "modified" in the child's git status) run outside it (issue #1241: one 67k-file provision held the lock 40 minutes and every other mutating start timed out). A held lock refuses typed (cause: lock_busy + the holder's pid/task/op), a SIGKILLed holder is evicted by the owner-aware stale check, every refused snapshot provision is definitely_unrun with a durable START_FAILED row, and the start receipt discloses snapshot size/time. Disclosed residuals: the registry read-modify-write under the lock is O(registered rows, their per-file baseline maps included — moving those maps out of the row is a follow-up), and every eligible untracked text file is hashed into the TARGET's own object database, where it stays unreachable after the snapshot is removed until that repository's own gc. The baseline is pinned by refs/ouroboros/delegated/ and checked out as a detached worktree; scope.root stays the authority target and execution.workspaceRoot names the snapshot. The host appends a separate typed binding after the immutable work order: the snapshot is writable and the authority is read-only until integration. Directory-copy runs use the engine-created copy and never the source folder. Full native access has no filesystem sandbox; the binding names the writable root but does not enforce it. Terminal capture records authority drift as evidence: ready-no-changes stays no-change with unknown authorship, while ready-with-changes keeps its private artifact and the locked baseline proof decides integration. Nested Git directories are excluded and disclosed; skill payloads use content-hash CAS. The binding is durable before POST and retry replays it; a GC-collected snapshot is a typed execution_snapshot_missing refusal. Worktrees live in state/subagent_worktrees.json; removal is explicit or custody-cross-checked startup GC, fail-closed on unreadable custody. The run still uses live from the engine view, so the scoped-HOME/delegated marker below applies.

At terminal, delegate_wait captures the run's diff against the baseline durably into the task's artifact store; NOTHING reaches the target automatically: the nanny explicitly applies or rejects through integrate_delegated_patch. Git and skill captures are whole-result operations: omitted paths and an exact empty list select the same captured result, disclosed when explicit; nonempty selectors stay directory-only. Engine-directory empty/subset semantics are unchanged. The staging substrate differs: a GIT workspace target applies under the repo git lock after PROVING no touched path drifted from baseline_sha (a plain git apply relocates hunks by offset; the touched-path set is read from git apply --numstat in BOTH directions, each naming only the paths it writes), then applies and STAGES, never commits. A SKILL-PAYLOAD target captures through the payload adapter over a parent-owned trusted index and applies LIVE into the non-Git payload: nothing is staged into any active root and no .git or index is created in the payload. The protected-path gate applies only when the target IS the Ouroboros body; a conflict (proven drift) is owned by the still-running nanny, with snapshot and patch persisting until explicit resolution or discard. Mutation rides an apply-intent protocol: a durable delegate_run_patch_apply_started row lands before any tree mutation, so on replay a pending intent without a disposition answers typed INTEGRATE_DELEGATED_APPLY_AMBIGUOUS, resolved only by explicit acknowledge_ambiguous=true, while the provably non-mutating outcomes (a lock error, proven baseline drift, a failed apply, a verified revert, a baseline-equal payload hash) RESOLVE the intent. patch_verdict.py is the ONE verdict writer for both pipelines: subjects are minted run_<rid> by the writer, never prefix-matched by readers, and each decision lands twice (artifact plus typed delegate_run_patch_verdict custody row), a failed artifact write disclosed on the row. artifacts.delegated_capture_read_target narrowly rebinds artifact_store READS for the owning task's own delegated_runs/ prefix, and delegate_shared.orphan_capture_read_target extends the same read-only, one-directory READ to the terminal-owner ORPHAN the disposition rule authorizes (confirmed by orphan_disposition_status), so the actor that may dispose a patch can inspect it without wider write authority. A read-only child stays in Claudexor's default envelope: one transport with one derived difference, not a second pipeline.

Terminal reconciliation captures only over PROVEN terminality, and patch_captured means "a usable artifact exists". When the owner task is gone, a settled mutating run's diff is captured through delegate_integration.capture_terminal_patch_for_drive, capture only, never an apply: the decision stays with an owner, and the obligation stays visible (delegate_custody.undisposed_patches) until the durable PATCH_DISPOSED row clears it. The obligation is held by the THING (the run's custody rows and the target itself), not by the task that created it, so a terminal owner's run keeps its debt visible and disposable instead of locking the target forever. While the owner task is LIVE, only that identity may dispose its captured patch; once it is terminal, any live TOP-LEVEL task may. Apply requires the caller's active Git root or fresh payload binding to equal the run's recorded target, or, for an ORPHAN only, to CONTAIN it while both lie under the host-minted subagent-projects root (the swarm aggregator shape: <project>/contributions/<track> clones the host already checkpoint-commits), every apply-path guard unchanged. Reject requires only the owner's proven terminality: it releases a dead task's locks and snapshot without fresh target authority. The PATCH_DISPOSED disposition row records who did it as disposed_by_task_id, and the delegated_runs_unreconciled projection stays the evidence trail the sweep heals. An owner whose terminality cannot be proven (task result missing, unreadable, or still an unreaped running row) keeps the lock, with no time-based release. A run closed absent or left unreadable captures NOTHING (across the provisioning boundary the child may still be alive and writing; an eager capture would freeze an incomplete patch and serve it forever); the snapshot stays preserved for capture-on-demand, and patch_captured is minted only over a ready manifest so a failed manifest leaves every retry point open.

The stored delegated_runs_unreconciled projection is healed only from the write side, at four seams. delegated_custody_unreconciled is the disclosure a task carries when it wrote its terminal result while one of its OWN delegated runs was neither applied nor rejected; a review-owned run is never one of them, so a task whose only open runs are its reviewers audits clean and carries no custody notice (the review projection's late_result_pending discloses a still-running panel). The overlay is ADDITIVE (outcomes.custody_debt_axes): it merges an objective WARNING and sets the top-level reason code but never rewrites the derived execution, review, objective or artifacts axes (the paid verdicts the task earned survive); a rail truncation code from BEST_EFFORT_REASON_CODES keeps the single Reason line (a round-limit or budget-exhausted victim is more usefully labelled by its truncation). While the debt is open the task reads Done with warnings, not Failed (the raw code stays typed on the row); the debt list lives on delegated_runs_unreconciled plus the delegate_terminal_reconciliation envelope. The stored code stays historical; the owner-facing Reason line is resolved at render time by the Task lifecycle rule: host rows in project_dialogue._custody_debt_reason, the browser card in log_events.taskReasonDetail, both from the stored list the live task_done event copies from the row (agent_task_pipeline._custody_debt_event_fields), so no surface infers a debt it was never given. Disclosed residuals: a provider-death terminal carrying custody debt shows the custody code on its one Reason line while the debt is open and remains Failed (the devtools-only infra codes are not in the runtime best-effort set); a CLEAN derived outcome plus custody debt does not increment evolution_consecutive_failures (a warning is not a failure); and project_dialogue derives failed/degraded from axis STATUSES only, never objective.warning, so a custody-debt child reads as a plain success there. Readers serve the stored projection (projection-over-replay, no live custody join), so a run settled after its task's terminal write stays stale until one of four seams: the periodic sweep refreshes tasks named in its own reconcile outcomes; the boot backfill (delegate_terminal.backfill_terminal_reconciliations, once per generation) re-audits stored terminal rows still carrying a disclosure; the cursor pass (delegate_terminal.refresh_recently_settled_terminals, a durable byte-offset cursor state/delegate_terminal_refresh_cursor.json over the append-only custody log, 5 MB per tick) catches terminal-boundary settlements no outcome-driven refresh can reach; or a kill path clears a stale list. Every refresh is audit-only in both directions (unreadable custody proves nothing) and never rewrites reason_code or recomputes the frozen delegated_runs_* counters: counters stay a historical snapshot while actual_substrate and the envelope mirror are rewritten from live custody, current liveness lives in the delegate_terminal_reconciliation envelope, and patch debt survives every refresh as patch:<run_id>, never a blind clear.

Disclosed delegated-isolation residuals (deliberate): a live top-level task with a different active root may reject and release another dead task's snapshot; the orphan-apply widening by containment means the disposer's blast radius grows to exactly the host-minted descendants the checkpoint-commit already writes; an orphan disposed by a non-owner writes its verdict artifact and delegate_run_patch_verdict row under the DISPOSER's task while the capture directory and the PATCH_DISPOSED row keep the OWNER's task id (readers key verdicts by run_<rid>, but the owner's acceptance packet will not list that verdict); a GC-lost snapshot or permanently failing capture discloses in a typed refusal that its obligation can never be satisfied; the baseline is worktree-primary; the git lock is task-drive scoped, so two nannies integrating into the SAME external tree can interleave apply+stage (the drift check makes the loser's apply a typed conflict); and a credential-shaped file the CHILD creates fails the whole patch instead of shipping a partial diff.

A mutating run requests its captured native profile, reads back what it got, and DISCLOSES the gap instead of refusing the work. Full requests no filesystem sandbox; the private snapshot owns delivery, not OS confinement. The owned gateway creates a scoped full-access grant only when no trust record exists; an existing denial is preserved and addressable by lowering the invocation or capping the row. Grants persist per scope like stable project registrations; no automatic trust cleanup is implied. Workspace-write keeps its boundary disclosure. In place because Claudexor otherwise hands the harness the operator's real $HOME, which holds the daemon token (a compromised child could start its own runs at any access level). Four facts, one mechanism. (1) The execution.delegated: true marker rides in the same record as isolation: live, built from delegated_run_shape in one place, so one cannot be sent without the other. (2) The version floor is about the SCHEMA: config.CLAUDEXOR_DELEGATED_MARKER_MIN_VERSION (3.3.0) is the oldest engine whose strict RunExecution accepts delegated; below it the start is a 400 and no run exists (route_health blocks dispatcher and nanny identically before a token is spent). Read-only delegation sends no marker and keeps the 3.2.0 transport floor; an engine between the two floors serves read-only and refuses mutating. The floor is a schema fact, never a containment fact (the OS boundary is platform-dependent, a build declares one version everywhere); threat model and measured bands: docs/DELEGATED_ADMISSION.md. (3) What was APPLIED is asked of the attempt, never of the OS: delegate_wait reads harness_home_isolated, confinement_mechanism and the proven denied path from the attempt record (gateways.claudexor.attempt_containment, delegate_containment.py); sys.platform appears nowhere in the decision. (4) A missing boundary is disclosed (durable delegate_run_unconfined event, the child's instructions, the parent's terminal payload), not refused (the child already holds a shell in this worktree; cutting the lane on every boundary-less host costs more than it prevents); a home nested under the operator's own is disclosed-unconfined (home_nested_under_operator_home), never relabelled as isolation. A recorded FALSE is still a fault: harness_home_isolated: false, or a scoped home that IS the operator's own, cancels as a typed containment fault exactly like a widened access profile (those two exact facts are the WHOLE breach rule); a MISSING home fact is neither breach nor proof, so unproven is REPORTED.

Delegated authority cannot widen. Start exposes prompt, session subagent_id, max_seconds, lowering-only access (readonly/workspace_write), recovery retry_of, continuation continue_from and the payload selector; continuation refuses beside retry/payload selection. No mode/isolation/scope argument exists. Retry preserves route/root/access/permissions; host instructions remain unforgeable. Native process access never overrides task constraints or the assigned edit target. Reviewers remain readonly/ask even if an actor row allows writes. Every fetched detail passes _containment_breach, checking both access and harness HOME because Claudexor derives effective access rather than echoing the request. Wider access cancels as access_profile_widened; narrower access is valid.

Read provenance on the accounts surface. GET /api/claudexor/status carries a reads block (ClaudexorStatusReads: catalog/accounts/quota, each ok|not_read|failed) (the owned daemon starts lazily): an idle daemon serves empty collections under a 200, and "no account connected" must not be inferred from a collection that was never read: ok makes the matching collection authoritative (empty means empty); one parity-tested client reader (facetReadState) applies the same rule, and the aggregate daemon.state is never the negative answer. Login jobs: the daemon stays the sole process/fence authority and reconcile is an explicit POST, never passive polling (gateway/claudexor_accounts.py; the route inventory is its §1 row and §4).

Git and commit review

tools/git.py owns repository writes, staging, reviewed commit, rollback or restore, tags, push and CI follow-up; its leaves are git_plumbing.py, git_repo_edit.py, git_vcs_ops.py, git_review_cycle.py (staging plus the advisory/triad/scope review and the reviewed-material binding the commit gate consumes) and git_evolution.py (campaign authority at the reviewed-commit and publication boundaries). File-edit tools validate their own atomic write shape. mutation_attribution.py captures the root-task baseline and projects the clean-at-baseline system-repository delta plus an explicitly adopted predecessor's exact retained changes — a changed pre-existing dirty path, a stale or missing baseline, or a failed scan blocks automatic staging; commit_reviewed(paths=None) stages only that attributed candidate, explicit paths must be a subset, an empty candidate returns GIT_NO_ATTRIBUTED_CHANGES, and managed update transactions keep their separate typed whole-tree authority. At startup a later independently initiated task may adopt exact retained candidates from its host-validated predecessor source in the same repository. predecessor_adoption preserves the observed dirty baseline and original source; terminal quiescence, unchanged path content and unchanged per-path base are required, while unrelated dirty work stays excluded. Missing fingerprints or size-only observations establish no transfer, and adoption grants no review approval or automatic new task. Commit preparation verifies the exact local working-branch ref before an unambiguous checkout: a missing ref refuses without changing the current branch, index or files (remote guessing and implicit branch creation are disabled), detached work retains the checkout -B <branch> HEAD recovery, and managed assisted merges retain transaction-owned precommit verification.

A reviewed commit is bound to one staged fingerprint. A cheap LLM-first advisory pass may run before the expensive gates; it is advisory, and skipping it never skips independently applicable tests, triad, applicable scope review, aggregation or exact-SHA binding. The hermetic preflight runs the candidate in a disposable worktree and data root; triad and scope inspect the same staged snapshot, aggregation preserves actor evidence and obligations, and any mutation stales the binding. Managed exception: a managed-update resolution commit reviews the declared M0→S subject (tools/review_subject.py) and the commit gate binds S to the exact index write-tree the fingerprint pins. External review wrappers report readiness but do not grant commit authority. The exact binding includes the git write-tree SHA, ordered HEAD and MERGE_HEAD parents, indexed VERSION, expected v{VERSION} tag, any existing tag target, and the binary staged-diff hash; after commit, tree, parents, VERSION and tag target are re-read before success or push is recorded, and an existing release tag is never silently accepted or retargeted. release_sync.sync_release_metadata projects version carriers during ordinary commit preparation; VERSION_CARRIER_SPANS/substitute_carrier_spans is the ONE span primitive the managed-update resolver and the commit-triad pack cut share, and carrier_only_change names a carrier changed only inside its declared version spans. Durable review state keeps attempts, obligations, readiness debt, raw actor evidence, and the final commit or tag binding; raw advisory output lives in state/advisory_review.json (selected by snapshot_hash/ts), not in review_status(include_raw=true), which exposes raw triad and scope attempt evidence; commit_gate.py classifies review blocks, refuses an identical verdict, counts paid cycles against the ceiling and fingerprints the review contract. BIBLE supplies review authority, CHECKLISTS supplies criteria, and Development supplies the procedure; snapshot identity, advisory coverage or audited-skip evidence, deterministic results, actor evidence, and final Git identity must all describe the same material.

Material ordinary Advisory commit review returns the complete outcome and exact review_reference before commit, tag or push. The author may correct, accept unchanged bytes, request another permitted review or stop. Explicit commit_reviewed/vcs_commit_reviewed continuation with that reference and author_disposition binds the current attributed candidate, reruns independent required preflight/tests and exact Git checks, and dispatches no critic. The author record references the original attempt instead of rewriting its subject or paid facts. Settled handback is reviewed/review_only; actual pending custody remains reviewing/late_wait and collectible. Clean supported review keeps its one-call path; Blocking gets no author override. Evolution receipts record actual triad_scope_status, never infer PASS from a successful Git commit.

Commit review evidence

review_evidence.capture_commit_review_evidence freezes selected browser/vision calls and same-round automatic image-attachment observations, explicit unavailable-image gaps included, after cheap/free admission and before preflight/triad/scope. The borrowed loop trace and original call refs retain exact redacted arguments/results and the immediately following visible response in the same execution, excluding provider thinking; the original context binding is restored through the attempt's _LoopExitContext cleanup even if tools._ctx changes while the loop runs. Adjacency never attests visual inspection. One canonical task source handle holds the selected UTF-8 view, its initial exhibit within _ACCEPT_NOTES_CAP with counts, completeness and source identity: native reviewers read artifact-store ranges under the real canonical root, sessions receive byte-identical bytes in ignored .review-drive/<task>/<sha>.txt, and packet-only reviewers receive a bounded, explicitly partial view when needed; existing request evidence/refs and the preflight execution evidence_source_ref bind it. A pending rejoin restores the same source, a recorded empty selection included, and carries it from preflight into triad/scope without selecting new trace. Commit cleanup waits for physical custody; standalone or uncertain views remain retained without a new cleanup registry.

Hermetic preflight proof

The hermetic runner (preflight_runner.py) alone mints ctx._preflight_test_proof, after actual green lanes and containment; it binds the assembled checkout tree and installed index, effective pass specs, isolated environment (test_environment.py owns writable roots), and the invoked/resolved interpreter and active Node identity (a Python symlink's invocation path selects its venv). The candidate is ONE hardened raw-bytes git diff --binary … HEAD capture, because a staged+unstaged diff pair cannot faithfully materialize unmerged entries; a capture/apply failure is the typed PREFLIGHT_CANDIDATE_ASSEMBLY block, never a test verdict. The node lane runs first (preflight_node.py: bundled signed node then PATH, floor 20.11; PREFLIGHT_NODE_MISSING/PREFLIGHT_NODE_TOO_OLD hard blocks and NODE_TESTS_FAILED on a red suite, never a silent skip), then the two-pass parallel/serial pytest split under one budget (LANE_EXCLUSION_EXPR is the marker-lane SSOT); a dead xdist worker or a missing required plugin is a distinct named block, never a retry or silent serial fallback. Equivalent ordinary and managed checks reuse that process-held proof only after their distinct baseline checks; a phase label alone cannot establish equivalence. Every workload binds HEAD, because even an unmarked default-lane test may inspect committed files or history: repeated preflights can reuse an unchanged subject, creating a commit changes HEAD and requires a new run, a skip or no-suite None is not a green proof, and a restart loses the proof and reruns the suite. Creation and reuse emit preflight_test_proof observations through the existing event log, naming the task, HEAD, tree, index and workload fingerprint; an unavailable log falls back to an explicit diagnostic, never a test failure or a new authority source. Managed tests_evidence is forensics, not reuse authority. Unproven lane teardown retains the temp tree and Git registration (_release_hermetic_tree), including raised lanes; deleting inputs beneath surviving processes would destroy custody.

Commit advisory cycle

preflight_review(commit_message="...", deterministic_only=True, source="worktree" | "index") returns a release-only diagnostic report before sync, staging, custody/state access, provider checks, tests or fingerprints. commit_admission.release_metadata_diagnostics owns source acquisition and reports all independent applicable findings plus unavailable sources; release_sync.release_metadata_findings reuses the carrier validators, release grammar and P9 history counters. The report identifies its source and status (clean, blocked, unavailable or not_applicable), gives no review freshness, and writes no review history. It checks the selected VERSION, not a future release. VERSION/README absence and failed reads are unavailable evidence; an absent optional older carrier stays optional, while a readable malformed carrier is a finding. Coverage is release metadata only, not syntax, structural size, tests or critic review; size policy stays warning-only locally.

Standalone advisory retains automatic carrier sync and worktree reads. Prepared advisory follows the commit gate's whole-index applicability, regardless of a narrower paths hint, including partial staging and the version-neutral carve; standalone documentation-only scope stays exempt. A present malformed VERSION is now a finding, not a silent skip. Optional[str] wrappers format the complete release findings, separating unavailable evidence from defects. Author continuation uses the shared name-status formatter and unavailable classification. Changelog prose, history trimming and version allocation remain deliberate.

Advisory availability is evaluated from the current configured slot and route, never inferred from a stale stored verdict: a disabled advisory slot is an audited bypass, an api_chat row requires provider credentials for its RESOLVED model, an agent_session row a resolvable session route. claude_advisory_review.preflight_review (callable advisory_review alias) owns that admission policy: an api_chat row rides the native tool-round episode (§6 Review stack), an agent_session row the session review executor; a native episode that ends on its own transcript bound — keyed on the structured native_transcript_cap_exceeded code, never on message text — reaches the caller as the typed non-blocking ADVISORY_SKIPPED with reason native_transcript_bound_exceeded, carrying the bound, refused chars and paid rounds, and every episode exception keeps failure_custody() as the advisory meta's usage, never an empty {}. The retrieving advisory brief carries a touched-path manifest instead of duplicated file bodies and applies the shared span-only release-carrier cut over HEAD→working-tree, with the same PACK EXCLUSION NOTE; omitted bodies remain readable through its own tools. Governance comes from the shared tiers (§6 Governance delivery). If the commit advisory is unavailable, the commit gate runs its compensating hermetic preflight only when tests remain independently applicable (not explicitly skipped, diff not documentation-only). Readiness projections receive only an exact repo/hash-matched advisory record and keep its failed status and freshness.

The reviewed-commit cycle (git_review_cycle) checks authorization and unresolved prior work before mechanical preparation and staging, binds the exact candidate before free-cycle/budget admission, runs a needed preflight inline with the full rebuttal and applicable tests, and revalidates the index and worktree snapshot before triad/scope; it owns entry-point resets and pending/blocked finalization, and its custody check joins current reviewer facts with the strict durable advisory record before index cleanup, an interruption before local metadata was updated included. AdvisoryRunRecord.execution_pending preserves physical custody for history retention and the external wrapper; blocks_preflight separately tracks logical admission. An explicit audited bypass releases that admission while retaining the original task, invocation, source and unknown physical outcome; late results update their original history row without superseding the newer bypass. Free advisory replay still checks freshness and buys no automatic preflight; stale coverage requires an explicit audited skip with applicable compensating tests preserved.

AdvisoryRunRecord.execution binds delegated preflight intent and candidate to the existing invocation token before POST and retains terminal usage/receipt evidence on failure; delegate_custody.invocation_record owns the immutable request, full prompt included — no second prompt store. An exact pending rejoin restores that request through the shared executor: a delivery-kind change cannot replace an unresolved session, same-session model/profile changes still rejoin the canonical request, and changed, foreign or lost identity is refused before preparation, naming the mismatched fields; review_status projects the recorded rejoin intent and execution identifiers through the public secret redactor without exposing the private reviewer prompt. Failed checkpoint writes cannot dispatch; unresolved rows survive history trimming. Only a corroborated failed_definite invocation discharges a stranded logical checkpoint automatically — age, missing identifiers and unreadable state cannot prove never-started work — and an audited skip never posts a replacement preflight. Released history stays eligible for an exact delegated rejoin but does not lock a later explicit standalone request for different intent or evidence; an unchanged audited bypass can satisfy the freshness shortcut without a new model call, naming the actual bypassed status. Native preflight uses its executor's end-event and monetary custody, with no separate advisory operation checkpoint or native resume protocol; a returned unknown physical outcome remains failed evidence on the existing refusal and audited-skip paths, and a hard process death has no native exact-rejoin handle. The external wrapper runs the same cycle, exports the real advisory record to advisory.txt, and preserves a pending checkout without restaging.

Review stack

review_cycles.py owns the shared paid-cycle ceiling OUROBOROS_REVIEW_MAX_CYCLES (a positive integer or unlimited; anything else fails closed to the shipped default). One number, four gate meanings: paid reviewer-panel cycles per task for plan review; paid panel runs per task for acceptance, independently of ordinary author responses; paid triad-plus-scope cycles per ROOT task for the commit gate, counted from the attempt ledger at DISPATCH so the ceiling counts money and only undispatched attempts stay outside it; paid panel dispatches for skill review. Byte-identical material replays each gate's recorded verdict for free under that gate's identity. Physically dispatched technical failures consume capacity; undispatched refusal and collection do not. The last allowed panel still permits author reaction. An explicit task-local max_improvement_passes=p retains its p author responses and p+1 paid ceiling; without that override N creates no N−1 response limit. unlimited removes only the paid count, not ordinary task rails.

For skill review, convergence guidance to the reviewer and retry coaching to the author use the group's review_round, so revising payload bytes does not reset the coaching. snapshot_attempt remains the separate ordinal shown in history and headers; free replay keeps its exact material and contract identity. A paid-ceiling refusal names author finish only where the existing enforcement predicate permits it.

Review waiting has six independent axes:

axis bound
transport dead-socket read bound
operation typed active-operation lease
logical slot/task deadline
policy budget / cancellation
absolute task ceiling
aftermath late-result custody

Physical custody

review_custody.py is the small worker-lifecycle seam under review_substrate.py: it schedules no tasks and keeps no second timing ledger. Parallel slots hold independent operation ids; the active-operation map only keeps a live physical call from reading as idle — deadline, budget, cancellation or ceiling still wins. Delegated-session expiry uses the verified cancel path; API/thread calls disclose in_flight and reconcile a late answer before the same retry identity can dispatch again, and a settled terminal API error stays in the cycle's replayable actor roster, so a sibling can neither erase its terminal fact nor buy a duplicate physical call. Physical custody is proved by the capture/operation state, never by the synthetic operation id, in usage_accounting's vocabulary: only reserved/released proves pre-dispatch, a positive capture outranks a contradictory synthetic label, and dispatched/unresolved without a typed terminal HTTP status stays custody-lost/no-resend. Custody never infers pre-dispatch provenance from Python's implicit __context__ (a fallback raised inside a prior provider handler inherits that earlier attempt): only an explicit __cause__ or typed transport metadata can release a paid row. A retry token without a durable invocation is custody-lost on every delegated review surface, and a durable token is valid only for its recorded surface, slot and operation. Retry identity is the explicit material/cycle identity when supplied — mutable prompt or history is deliberately not part of it, while a changed snapshot, owner intent, reviewer route or admitted cycle mints a new one.

Timeouts compose, never override: send-time VLM captioning keeps its direct 90-second provider cap and Anthropic's direct route its 120-second provider default — neither is a generic review-reasoning cutoff — and the whole hierarchy is narrowed inside the owner deadline and finalization reserve before dispatch. A returned provider response or typed terminal error is settled even with an empty body, so bounded repair/retry may apply; a dead socket after dispatch is provider_outcome_unknown and cannot trigger another paid route. A spent owner window yields a typed $0 not_dispatched row before fan-out; under blocking enforcement an in-flight triad row stays pending rather than becoming a final quorum verdict, and a primary call at the deadline boundary enters the local finalization rail (deadline_local), not a provider-outage relabel. A reviewed commit has no independent outer tool cutoff: the foreground caller keeps custody until settlement.

Paid stamp and owner custody

The paid fact behind the cycle ceiling is recorded WRITE-AHEAD at physical dispatch (review_dispatch.py): a gate installs one once-only ReviewPaidStamp for its wave, and each delivery stamps at its own boundary — a packet api row at the usage ledger's physical-attempt transition, a native episode on each paid send, a session before its replayable START_REQUESTED row. Assembly-only refusals exit before that seam, so a $0 attempt stays outside every ceiling, and neither a worker outliving its logical caller nor a crash after dispatch can lose the durable fact. The stamp is idempotent across the concurrent triad and scope callers, so no side's transport starts before the paid fact is durable; commit review and task acceptance treat a failed write as fail-CLOSED, skill review only when resuming a wave, other callers keep fail-open accounting. Task acceptance binds a strict exact-hash tree-wallet claim to that same stamp on every delivery its rows run — one idempotent claim per panel, because the paid identity is material, not route; a wallet or cancellation veto at the claim releases the reservation and no reviewer transport proceeds on any row (the launch floor is evaluated once, at loop admission: §6 Task lifecycle). Disclosed residual: a compatibility transport that raises with positive physical capture without having entered the canonical marker is stamped after the send; if the tree's last paid cycle is consumed concurrently at that late stamp, the wallet refusal replaces the captured exception and the substrate may resend.

Before either parallel surface starts, one locked write records paid=True plus both complete slot rosters and operation ids in the commit-attempt row; a delegated slot patches its reserved row with pending_invocation_id before the provider POST. Exact resume preserves rows and tokens; a missing or mismatched operation is custody_lost under every enforcement mode. The stamp records the process-custody server session and pid owning the reviewer threads (review_owner_custody.py): owner-loss is proven by pid death, never elapsed time — a TTL would convert waiting into resend authority, and starting another Agent is not evidence this owner died. Tokenless rows settle as typed infrastructure failure only after a death seam confirms that exact pid gone, and a legacy row without owner identity stays fail-closed. Plan review re-enters a recorded in-flight cycle only through exact live custody or complete operation-addressed CAS for its recorded dispatched rows; a delegated poll that loses transport after a run id exists preserves the exact durable invocation for retry custody rather than cancelling a healthy unknown run.

Per-row delivery is configured only through the structured reviewer-slot SSOT (reviewer_slot_config.py); the older per-row route envs are retired and ignored. The one shared session target is OUROBOROS_REVIEW_SESSION_ROUTE (review_execution.REVIEW_SESSION_ROUTE_ENV), an opaque harness[=model][:effort] spec that falls back to OUROBOROS_SUBAGENT_HARNESS when unset, so one delegated route is configured once; a set but unparseable value turns review sessions OFF with a warning rather than silently re-routing them onto the subagent route.

Surfaces and money admission

The review surfaces:

  • Advisory pre-review (claude_advisory_review.py) — a cheap, staleness-aware error-finding pass; the audited skip covers only advisory admission, never authoritative review, test policy or snapshot binding.
  • Triad diff review (tools/review.py) — configured reviewer slots cover the Repo Commit Checklist with JSON findings under config.adaptive_quorum (the same SSOT as scope/plan/skill/acceptance review).
  • Scope review (tools/scope_review.py) — intent/scope/coupling across the whole repository by retrieval on every row, in every context mode; the exact candidate, required-source manifest and diagnostic reading coverage stay distinct from substantive findings and configured enforcement (§6 Scope review by retrieval).
  • Parallel orchestration (tools/parallel_review.py) — two-phase admission: the triad packet is assembled and fit-checked and every scope brief is prepared before any reviewer is dispatched, so a deterministic assembly block anywhere dispatches nothing and spends $0 everywhere (a lane paid while its sibling was doomed is real money for a half-verdict). Money admission (tools/review_admission.py: commit_gate_paid_seats / admit_commit_gate_wave) is all-or-nothing and scope-first: every PAID seat of the wave — the blocking scope seats FIRST, then the triad — is priced with the reservation math reserve_attempt will apply (its exact first send, including native work-order and tool schemas, plus its own output reservation; later native rounds reserve themselves), under the usage scope its substrate sends under (review_usage_category + slot, so a warm split of the caller's transcript never stands in for a seat's cold prefix), against every fence reserve_attempt enforces — the global TOTAL_BUDGET remainder (unbounded when the owner set no finite global budget) and the task's CURRENT root fence — and admitted as ONE wave through review_wave_budget_gate(surface="commit_gate"). A wave that does not fit is a typed $0 not_dispatched record on every seat plus one review_wave_budget_insufficient event carrying binding_axis (global or root) and both remainders, and the block names the binding fence, the shortfall and the knob that moves it — never a half-dispatched panel whose non-blocking seats hold the money the blocking reviewer needs. Admission is a read-only pre-check without a wave-level hold: the per-seat reservation stays the enforcement, so money a concurrent task takes on the global axis between admission and a seat's reservation can still refuse that seat (disclosed residual). A paid seat is any api row (packet or native episode); an agent-session row rides the owner's subscription — its ledger row is written at settlement, never reserved — and is neither priced nor waited for. Every executor transition on the way to a seat runs under contextvars.copy_context(), so the usage scope the wave was admitted with — its bound root fence included — is the scope the reviewer rows reserve under; a settings reload that changes the per-task limit mid-turn changes neither side. An admission that raises is fail-open like the gate's own unknowns, but loud and typed (review_wave_admission_unavailable; the wave dispatches unadmitted). An admitted wave submits the scope seat first and starts the triad only once the scope seat's OWN reservation is on the ledger (nothing else releases the hold), bounded by NESTED_SETTLEMENT_MARGIN_SEC and by the scope seat itself; a hold that ends without observing it emits a typed review_scope_lead_unobserved event. Admission touches no transport or logical timeout contract.
  • Shared helpers (review_helpers.py, triad_review.py) — packet/brief building, checklist loading, JSON extraction, usage events, obligations scaffolding and reviewer actor records.

A free refusal never wears the form of a verdict: a refusal that spent nothing is recorded as a typed not_dispatched fact plus its reason, never as a DEGRADED panel, a synthetic actor or a verdict. That is the one shape across every $0 exit — a plan-review locator the evidence policy cannot attach is a named omission row the panel is dispatched with, an acceptance packet that overflows takes the ladder rather than reporting a verdict, a truncated or self-pageable row names its cut, and a request never sent is one seat with operation_state='not_dispatched' that stays in the denominator.

Rationale: diff reviewers catch line-level mistakes; the scope reviewer catches cross-module contracts; running both on the same staged snapshot prevents one result from hiding the other. The managed-update exception reviews the same declared M0→S subject in both lanes. Task acceptance is a separate root-owned post-delivery system specified once under Task lifecycle; it shares adaptive_quorum and these transports, has no acceptance scope actor, and stores its verdict separately from the terminal result.

Session identity and advisory parsing

Session reviewer identity comes from gateways.claudexor.final_attempt_facts: the unique final_attempt_id row in the engine-owned final/telemetry.yaml, bound to the requested run id, and model, harness and credential profile are read from that SAME attempt. summary.model/harnesses echo requests, and summary route/auth projections may borrow earlier-attempt facts; none supplies a missing observation. Reviewer usage, last-execution views, delegated terminal payloads and settlement preserve known facts or explicit absence; the original custody/billing route remains the chosen authority, separate from the observed actor. No quorum rule or model-name mapping follows from this disclosure.

Advisory row parsing owns its PASS|FAIL and critical|advisory values at preflight_review_run._is_checklist_array, canonicalizing case and surrounding whitespace once for every consumer. An unknown verdict, or an unknown or missing FAIL severity, rejects the whole array into the bounded extraction rail; unresolved output stays parse_failure with its full source retained. PASS without a severity remains compatible and the separate genuine-empty-clean predicate is unchanged; an optional array validator lets the shared canonicalizer honor this surface contract without changing triad, object-verdict or report semantics. Canonicalization never changes the reviewer's judgment by searching the repository for words or identifiers.

Structural gates

Structural smoke gates are a deterministic BIBLE P3 codebase-size component. ouroboros/review.py::iter_gated_modules is the one source inventory for smoke, codebase_health, census and the UTF-8 byte gate (Python everywhere plus first-party web/**/*.js, vendored/minified excluded). ouroboros/size_ratchet_manifest.py is a generated, data-only debt register consumed through AST literals: exact module debt above 1600 lines, exact function debt above 300 lines, the 1001–1500 band with rationale authority, and exact byte debt above 200,000 UTF-8 bytes. validate_size_ratchet proves the live and staged manifests exact against their trees and shrink-only against the merge-aware committed authority — no first-parent history replay: the previous manifest resolves from HEAD or any of its parents, and a checkout with no committed manifest anywhere bootstraps from its own tree, so a fork whose local line predates the manifest is never condemned by inherited topology. Official-repository CI runs the blocking size_ratchet pytest lane while every local surface reports the same findings as warnings; two disclosed residuals — pairwise validation covers only the base→HEAD interval, and the official block presupposes branch protection. Within a validated pair, debt can shrink but cannot be swapped, re-entered, grow on the byte axis or survive stale. MAX_TOTAL_FUNCTIONS (ouroboros/review.py) remains the coarse runtime ceiling. codebase_health and review readiness also report informational headroom from that same inventory: runtime function count, LF-normalized module lines/UTF-8 bytes and function size against current limits. The overview prioritizes ordinary near-limit modules before registered debt; readiness selects touched paths. Debt and omitted rows are labelled. Positive capacity stays separate from warning findings and adds no gate or setting.

A deterministic hot-store growth invariant sits beside these gates: agent_startup_checks.py::hot_store_growth_notes (surfaced by context_health.py::build_health_invariants and once per worker boot) stats eight hot stores plus the archive/chat_*.jsonl aggregate — state/usage_attempts.jsonl, logs/events.jsonl, logs/tools.jsonl, logs/supervisor.jsonl, logs/task_reflections.jsonl, logs/progress.jsonl, state/scheduled_tasks.json and state/skill_review_root_tasks.jsonl — against justified byte thresholds in ouroboros/context_budget.py and emits a WARNING with a remediation pointer.

Three gen/verify inventories ride the same discipline as the size manifest (generator scripts/regenerate_inventories.py, verify tests/test_generated_inventories.py, staleness = red): the frozen-contracts inventory (docs/inventories/FROZEN_CONTRACTS_INVENTORY.md, a machine extraction of §11.1 with every owner/anchor path resolved against the tree and the ouroboros/contracts/ package-coverage gap pinned), the data-layout inventory (docs/inventories/DATA_LAYOUT_INVENTORY.md, every entry of the §1 "Data layout" tree probed as a tracked repo path or a runtime-source literal, so a durable file renamed in code while its tree row survives turns red), and the facade inventory (docs/inventories/FACADE_INVENTORY.md, the AST-derived noqa: F401 re-export surface with per-leaf domains from ouroboros/domains.toml).

Prompt size, density and windows

review_helpers.REVIEW_PROMPT_TOKEN_BUDGET bounds assembled inputs: triad packets, per-slot plan packets, task-acceptance packets and skill-review chunks with their own headroom. Scope and deep review assemble briefs, not repository packs; native sends use review_native_transcript_bound and reach other sources through tools. Every surface scales its output reserve to the route's usable window (reviewer_window.window_scaled_reserves); a reserve sized for a much larger window can otherwise consume all input room before a request reaches the provider.

Sizing is density-calibrated, not keyed to model names. usage_accounting.execute_physical_attempt records settled (prompt_chars, prompt_tokens, route_fp) witnesses in capability_evidence.json. Main uses the newest fresh exact-route witness, then exact-model witness, then neutral density; reviews use the densest fresh exact-model witness, which can undercut COLD_START_TOKEN_DENSITY. Stale, absent or cross-model evidence keeps the conservative floor. The witness TTL and bounded retention preserve both its densest support and recent observations. calibrated_input_token_limit combines the shared cap, measured density and absolute margin per call; triad, plan, acceptance and native transcript sizing consume it rather than freezing an import-time result.

Only the triad packet has a cold-start probe rung: review_admission.density_probe_before_size_refusal invokes capability_evidence.cold_start_density_probe when an irreducible packet exceeds a cold route's cap. One bounded slice of the actual prompt is sent on that exact model under physical_attempt_limit(1); its witness buys one resize and fit pass. It never probes a warm route, retries the probe or runs when the packet fits. Progress and review_density_probe retain the attempt; a paid-ledger refusal is budget_refused and leaves the size refusal unchanged. This breaks the loop in which a request rejected before dispatch can never teach its tokenizer density.

reviewer_window.resolve_reviewer_window is the route-specific sizing source, with metadata probes serialized per route and limited by evidence TTL. There is no model-window table or window-authority floor. ReviewerWindow.sizing_window has a full-window sizing default for an unknown API window; raw subscription routes keep no numeric unknown-window assumption. scope_window adds its designated-default or conservative fallback and distinguishes confirmed, asserted, stale-unverifiable, designated-default and unknown provenance. POST /api/owner/capability-ack records an asserted window bound to one route fingerprint, invalidated by a route change; it sizes sends and grants no reviewer authority. A failed send supplies no verdict, while a small window alone does not remove a responding reviewer.

Scope review by retrieval

Every scope row retrieves in every context mode. A native API row preserves its configured identity, model, route, account pin, effort, processing preference and local-route flag while the surface selects inspection delivery; a session row uses its configured delegated session. The first review where a stored bare API row acquires episode semantics emits review_scope_delivery_migrated and writes state/scope_delivery_migration.json: changed cost and duration are disclosed before becoming an unexplained bill. Failure to write the marker repeats the notice later rather than failing review.

review_admission.prepare_scope_review builds each brief through scope_review_session.build_scope_session_task, without dispatch. It carries intent, rebuttal, history and open obligations; the touched-path dispositions and candidate sizes; the repository index; shared governance tiers and scope checklist; calibration/output contract; task evidence; and the required-source manifest. Stable governance precedes change-relative context for cache reuse. The full staged diff stays inline while the first send fits, measured by native_first_send_chars including schemas and wrapper. Otherwise it becomes one exact paged source in the task artifact store, with a byte-exact .review-drive view for a delegated session. An unavailable store is disclosed and keeps the diff inline. Native inline sizing uses the landing threshold within scope_first_send_bound; sessions use SESSION_INLINE_DIFF_CEILING_CHARS because the harness owns their context selection. A managed-update subject is the authoritative M0→S resolution delta, never a substitute git diff --cached over the two-parent candidate. Invalid roots, missing subjects and failed assembly remain typed pre-dispatch failures; only an irreducible first send reaches native_bound_below_first_send.

When context construction fails before dispatch, preparation best-effort appends one scope_review_preparation_failed row to the existing events log, whose normal sink forwards it live. It names the task, slot, final model and original context_unavailable reason. This is row-local preparation evidence, not a panel verdict; logging failure preserves the same typed result, and model-control exceptions still propagate before logging.

scope_required_sources.py declares the change-relative reading minimum: touched protected runtime, prompts and frozen contracts; protection-owned families (GIT_OPS_FAMILY_PATHS and tool-dispatch leaves); and declared Python/browser contract twins, with both sides included. An ordinary touched file needs no whole-body obligation because the diff already carries its complete change. Each readable row binds raw-byte source_revision, normalized-text complete_sha256/complete_chars, range basis and candidate tree. Deleted or renamed sources retain their exact baseline preimages; paged diffs join the same manifest; unavailable sources remain explicit rows. Byte-identical governance already inline is satisfied by delivery without another read. The final rows are shared by brief and policy, and required_sources_ref plus SCOPE_REQUIRED_SOURCES_POLICY bind the actual manifest and review-contract fingerprint, so old replay authority cannot survive a changed contract. This minimum never restricts the reviewer's further reading.

Coverage records what could be observed: native receipts match source identity, opened root/path and delivered character intervals; sessions fold weaker journal evidence over the same rows. States are complete, incomplete with missing ranges, declared_empty for an explicitly empty manifest, and unobserved when measurement is unavailable. A changed source is a source_gap; availability or inline delivery never proves understanding. review_context_atlas.repository_index supplies orientation only: compact tracked-path dispositions, collapsed excluded classes, touched-file and direct-importer facts (size, digest, language, symbols and imports, with bounded totals disclosed). It renders no bodies, spends nothing, writes nothing and never limits what the reviewer may open.

The scope reducer keeps received findings and counts responding reviewers through config.adaptive_quorum independently of coverage. Reading gaps remain actor and panel diagnostics: they cannot change status, block a commit or trigger another paid review. The agent judges whether a particular gap needs more reading. Substantive critical findings still follow OUROBOROS_REVIEW_ENFORCEMENT; an unanswered reviewer does not count, and an all-not_dispatched panel retains its assembly/admission failure. Oversize requests, route failures and candidate-binding failures never become PASS; budget, custody, deadline and ownership keep their separate contracts. Retrieval keeps the reading obligation proportional to the change while exposing the actual evidence, instead of letting total repository size decide whether review can occur. If the change cannot fit a reviewer, split the change rather than weaken the reviewer.

Governance delivery

tools/governance_context.py owns the shared tiers for triad, scope, advisory and deep review. Tier 1 carries the applicable CHECKLISTS section, BIBLE.md and CHECKLISTS_ARCHIVE.md inline in a stable prefix. Tier 2 selects the review-protocol chapter, DESIGN for web/ changes and DEVELOPMENT chapters naming touched files, within runtime_limits.REVIEW_GOVERNANCE_INLINE_SHARE of the usable window. Tier 3 gives physical-source book navigation for ARCHITECTURE; packet triad rows can additionally inline relevant sections within the same share. Each non-inlined document has a manifest disposition and navigation pointer. Retrieving rows can follow those pointers; tool-less packet rows receive an index of material not delivered, never an instruction to call unavailable tools. A whole document with the same source identity is credited as delivered_inline without a duplicate read.

Guaranteed-fit ladder

The triad packet retains its one-pass fit sequence in review_admission.fit_triad_prompt: shorten the optional evidence projection, replace duplicated post-change snapshots with a touched-path note, then re-render the same pinned subject at -U0 while preserving all file/hunk identities and added/removed lines. The cold-density rung above can precede this sequence once. Advisory briefs and triad packets share the disclosed span-only release-carrier cut; triad additionally omits governance bodies byte-identical to its inline prefix, while managed subjects retain full texts. Every omission is named, and an irreducible oversized packet is a typed refusal. Scope uses its paged exact subject instead; plan review sizes its typed specification and declared evidence per slot, retaining the constitutional material self-modification plans require.

Planning, deep review, reflection, memory

Plan review, task acceptance, commit review, and deep self-review answer different questions and never inherit one another's authority: planning judges a proposed approach before implementation, task acceptance the delivered objective, commit review a staged self-change, deep self-review the whole system. Post-task reflection and memory persistence learn from execution but approve none of those boundaries.

Plan construction and review

plan_task reviews an INTENTION before the work starts — the same organ for code, research, deliverables, and actions (BIBLE P3). The obligation is constitutional and the finalization gate is structural (owner_hurry.force_plan_decision); the order in which a task asks, explores and plans is the mind's judgment, and no prompt choreographs it. The envelope carries the goal, the plan prose, and a typed domain-neutral SPEC: in_scope, non_goals, acceptance_claims, invariants, decisions, deferred, affected_paths (REQUIRED — the files the work will CHANGE, [] when none), affected_resources (the same question in words: systems, services, projects, people, never resolved as paths) and evidence. A spec without affected_paths is refused before any dispatch with the typed PLAN_RESOURCE_FORM_REQUIRED, which records nothing. ouroboros/tools/plan_spec.py validates and normalizes the full operative content without shortening or dropping anything, hashes it, and mints the only valid breaks targets (goal, claim_N, invariant_N, decision_N, deferred_N) positionally, because a caller-chosen id could shadow another target and corrupt what a blocking finding breaks. Governance documents always come from the system repository; declared targets and evidence resolve against active_repo_dir_for(ctx), and a path escaping the active subject or an unreadable root is a named omission, never a silent gap.

ONE structural fact tiers the governance pack: constitutional is true iff a declared affected_paths locator resolves under the Ouroboros system repository, whether the file exists yet (creating ouroboros/new_module.py IS self-modification). Nothing else buys the pack: affected_resources is prose, and an evidence locator is something to LOOK AT — reading a repository file is not changing it — so system-repo reads are only named in the disclosure note. A constitutional plan carries BIBLE.md and ARCHITECTURE.md in full (inline on an api_chat row; mandatory full reads on a retrieving agent_session row); that packet is not tiered, so assembling it without either is a typed failure, never a disclosure. Every other plan carries the runtime heading-derived navigation maps (context_layout.generate_doc_nav_map, never a copy) plus resolvable pointers. There is no plan-kind taxonomy, no agent-declared plan_class, no planning scouts, and no assembled repository pack for a plan.

The ONE caller-facing strength axis is the envelope's optional reviewer_effort: the panel's effort for THIS order, placed on the default rung of each row's ladder (reviewer_slot_config.row_effort: an explicit per-row effort, then a compound Cursor/Agy route slug, then the declaration, then the owner's OUROBOROS_EFFORT_REVIEW) and passed as an ARGUMENT of plan_review_slots only, so every other review surface keeps reading the untouched rows. Effort is roster identity: a different declaration re-dispatches a paid panel within OUROBOROS_REVIEW_MAX_CYCLES, the same one replays free, and it is recorded per wave (reviewer_effort; declared_effort in the last-execution projection, apart from the row's saved effort). Against BIBLE P1 this is the review panel's strength for one order — the owner's setting is the default, Ouroboros may order stronger or weaker per envelope — never the core's own model, effort or horizon. Disclosed residual: a cheap panel's closed GREEN is earned authority for that envelope, and a later stronger order does not reopen it.

Declared evidence is resolved by ouroboros/tools/plan_evidence.py against exactly two roots — the active workspace and the system repository — with the shared sensitive-name policy on locator and target and the disclosed bounds EVIDENCE_PER_ITEM_BYTES / EVIDENCE_TOTAL_BYTES: every refused, missing, truncated, oversized, binary or URL locator becomes a typed omission row in the manifest (the host never fetches a URL); a locator may carry an exact range (::lines=A-B, ::bytes=A-B, ::tail=N, ::symbol=Name for .py sources), and an oversized source is attached head-first with the cut named. A need_evidence locator a reviewer names is attached by the host on the next cycle through the same policy; what it cannot attach becomes a [reviewer-requested] omission row the panel is dispatched with, never a reason to run no reviewer. The manifest hash joins the spec hash and constitutional in the wave fingerprint, so changing what the reviewers can see changes the identity of the review — an attached locator makes the next envelope a new fingerprint, never an idempotent replay — while the root exploration log stays outside that identity. The evidence continuation uses a fresh full-packet dispatch only when no exact artifact reference exists, disclosed per slot as a capability_delta; an unreadable referenced artifact instead fails closed with plan_review_exact_artifact_unavailable and never mints replacement authority.

Own-room dialogue is automatic evidence: ouroboros/dialogue_evidence.py reads the complete retained room through Memory.read_chat_generations (rotated and consolidated history included) under the shared project_dialogue.room_membership predicate, plus the progress stream and addressed owner_mailbox entries — both speakers, options, quiz recommendations and answers, attachment names (the canonical manifest label), typed outcomes and sender/relay provenance. The winning quiz answer is written to canonical chat history as its structured block (after record_answered; an answered quiz_closed reply heals it), so eviction and mailbox cleanup never erase it and History shows it in the original card without a second bubble. Missing history, room binding and unstable capture remain explicit coverage gaps; acceptance freshness stays with the acceptance directives ledger, and planning consumes no separate directive section.

ouroboros/tools/plan_dialogue.py captures the redacted full room as an immutable task_source in existing artifact custody, bound to the author-request identity (goal, prose, spec, declared evidence, the constitutional fact). Identical author inputs and free collection reuse the recorded snapshot — progress, panel-settled messages and other room growth cannot buy a new panel — while a genuinely changed author request (a changed plan or an explicit evidence request) captures the current room. The wave fingerprint includes the snapshot digest; dialogue_source_ref and the author identity survive hot-index compaction and child promotion beside spec_source_ref; every result names the original capture and excludes later messages. chat:<id> and chat:<id>@<sha256> locators address the room and its recorded snapshots through plan_evidence.resolve_evidence with the same selectors, redaction and omission accounting; runtime files do not become general path evidence, and related Main/lineage rooms (bound siblings included) remain labelled pointers, never unsolicited inline neighbours.

ouroboros/tools/plan_packet.py keeps the current objective, spec and prose complete, and own dialogue bypasses the ordinary per-item/total evidence bounds. Delivery follows the per-delivery source contract: Packet/API slots receive all of it while their calibrated route window, governance and output reserves permit, otherwise the newest text plus the exact immutable omitted byte range; native retrieving slots fit under review_native_transcript_bound and the mandatory-reading declaration — native sizing measures the complete first request (native_first_send_chars, schemas and wrapper included) — with full artifact access through native_data_root; delegated sessions receive a mandatory full-read instruction and the complete redacted artifact outside active_root, their harness owns context selection (no host-invented numerical window or 1M claim), and an actual overflow is disclosed as exact coverage/omissions. File availability never attests that a reviewer read it all: plan review declares no required-source manifest to fold, so coverage stays declared/unobserved and a retrieving reviewer is recorded host_file_read_attestation: unobserved.

Slots come from reviewer_slot_config over the review_execution._review_route_executor seam. Reviewers return ONLY a typed findings array — blocking with a breaks id, note, or need_evidence with a locator the host attaches (or a spec id in breaks as a question to the author, who answers, escalates or defers it openly in the disposition); the HOST validates membership, demotes invalid findings with disclosure, keeps failed slots in the quorum denominator, and computes the aggregate — GREEN, REVIEW_REQUIRED, REVISE_PLAN, or the honest DEGRADED — through config.adaptive_quorum. No reviewer emits GREEN as authority. Premise criticism and simpler alternatives are optional brainstorming notes whose adoption is Ouroboros's; no mandatory competing plan, and repetition alone cannot promote advice to a blocker.

plan_review_state v2 inside the root task result is the bounded durable index: each wave records a hash-bound task_source reference to the full operative spec, the hashes, constitutional, validated findings, aggregate, dispositions and paid state (full waves, specs and manifests: plan_review_artifacts.py; review evidence stays redacted). State reads share the task-result writer's lock without rewriting JSON; source resolution starts after lock release. Current-authority readers restore the full spec before comparing plans or binding its acceptance_claims (contracts/task_contract.effective_acceptance_claims); compaction and child promotion retain both references, and an unavailable recorded source is a typed infrastructure failure, never empty claims or a new unpaid cycle. A v1 record is read-only (legacy_v1_projection; an OPEN v1 wave projects legacy_open_requires_resubmission, never auto-closed). A corrected full goal/plan/spec can be selected by an explicit review_disposition.author_action with author disposition and the referenced review fingerprint. current_attempt.author_subject holds its exact source and separate critic reference; plan_review_artifacts.current_author_plan restores it without a new wave, GREEN or paid cycle. Advisory finish can select its claims as acceptance_claims_source=author_plan, after ingress claims and through the existing receipt-support checks. Blocking can save and stop but cannot use the plan as approved; closed_plan_review_wave remains literal critic-closed authority. The gate and terminal disclosure follow the author's explicit review_fingerprint for historical critic outcome and pending custody as EVIDENCE INTO A GAP — a revised plan with no wave of its own stops reading unavailable — without granting that wave's closure to the revised plan: the enrichment fills only an ABSENT outcome and never moves the attempt's own lifecycle, because overwriting the status with open took the release back off an exhausted cap and off a degraded rail (an author permitted only to finish honestly was pushed back into a cycle it cannot buy), forcing closed=False contradicted closed_plan_review_wave, which still bound that wave, and replacing a current wave's own aggregate would let a historical verdict stand in for the one this plan earned. The projection therefore carries historical_critic, and both of owner_hurry.py's renderers read it: the terminal disclosure labels that outcome as the earlier plan's instead of showing it as this plan's own, and the blocking reminder answers it with its own branch, because its outcome-keyed advice would otherwise send the agent to dispose a wave that cannot approve the revised bytes. A rail that forced finalization keeps its own sentence with the author clause appended, since returning the author decision alone hid the cap is spent; the task ends blocked; an author STOP carries allow=True under every enforcement, so it is excluded from the advisory-continuation sentence rather than described as work that proceeded — and because the projection's stop REPLACES the status, a rail that fired under a stop is not separately named, the stop being the fact that ended the task. The open-wave status reads cycles_exhausted from the attempt and the wave, because an author finish at a spent cap stamps only the attempt (_apply_author_subject, the one writer with no in-flight guard) and the gate otherwise held a task that may only finish honestly; that read is guarded by custody_pending, since a paid slot still working could yet close the wave and releasing then would trade an honest hold for a premature blocked terminal. Missing criticism stays unavailable, and author finish/stop stays separate. Disclosed gap: web/modules/review_presentation.js::planReviewGroupFromTaskDetail still selects the critic by the author's review_fingerprint and renders that aggregate as the group's own verdict, so the CARD can show an earlier plan's GREEN where the terminal disclosure says the revised plan has no verdict; the same projector also reads rail_degraded from the attempt but cycles_exhausted only from the wave, so an author finish at a spent cap on its own wave now releases the gate and reads blocked in the terminal while the card still shows the wave's aggregate. Both are one presentation gap needing its own change with browser verification, not a backend fact. Old-wave collection never replaces the selected author source. The first exact wave also stores the request policy and per-slot prepared inputs/fit sizes, which collection reuses unchanged — native_mandatory_read_chars, the review contract hash and the physical-operation binding never move with live exploration or owner clarification — and a historical wave without them is a typed source-unavailable result, never guessed values.

A fresh dispatch returns control at the dispatch barrier (ReviewRequest.drain_deadline, drain window 0): the wave is recorded open with typed pending_dispatch rows for the workers still running, whose custody stays with the recorded wave in the process (review_custody). A paid actor still physically in flight keeps the wave open as DEGRADED with review_late_result_pending even when the settled rows meet the arithmetic quorum, so a late blocking result cannot arrive after a false GREEN. When the last released slot settles, the settlement thread writes ONE system frame into the task's mailbox (owner_mailbox.write_task_message, provenance system) so the mind wakes exactly as for a child result; it never closes or aggregates the wave — collection is the sole wave writer (plan_review_collect). Task acceptance's twin announcements (its own quorum, completion) carry each reviewer's own verdict; the aggregate still comes only from collection.

Collection is the existing review_disposition mode addressed at the recorded wave (items may be []): it reconciles process-local custody with drain window 0 over the wave's own recorded inputs — never re-reading the evidence, never changing the wave's identity — closes or advances the wave, pays the cycle once from the settled rows, and never waits; waiting longer is only the identical envelope resubmitted. A reconciliation re-records the packet that was physically DISPATCHED (each slot's request_messages, session_task for a retrieving row), never a rebuilt one, so a directive that arrived after the dispatch reaches the reviewers only in the next physically sent packet. A NEW review envelope over an in-flight wave first collects what has settled at $0, then supersedes it (reconcile-before-supersede) instead of being refused; late settlers attach as historical supplements, accounted once without another send. A wave with custody pending is never compacted out of the hot index.

An answer that has not arrived is a gap, never a failure or a verdict. The STORED wave is the fail-closed floor: an unanswered slot stays an ok=False row no quorum counts, and the open DEGRADED + custody_pending pair holds every gate with no fifth aggregate for older builds to refuse. Renderers read plan_review_runtime.plan_wave_slot_census before wording a slot: awaiting is operation_state=pending_dispatch only; unresolved is in_flight/custody_lost, named as unresolved on the line and by raw state in Reviews, never called waiting (late_result_pending is true for both); uncollected has a settled supplement not collected yet; not_dispatched is a $0 refusal, not a failure. Verdict and failure wording and the task's degraded stamp belong to collected slots only: a finish over an only-awaited wave or panel keeps its disclosure and records awaiting (execution.plan_review, review axis; owner_hurry.plan_wave_only_awaited, _outcome_receipts.review_runs_only_awaited). owner_hurry.force_plan_decision makes ONE free collection of the current pending wave in every enforcement mode, hurry too, then projects the returned state (plan_review_collect.collect_before_gate). Advisory and hurry keep their local release; collection updates disclosed facts, never permission to proceed. Context health reads the canonical wave, never collecting: PLAN REVIEW WAVE OPEN names its fingerprint, pending-slot count and snapshot timestamp (not a physical start, not a liveness lease) and disappears once collection clears custody.

Closure follows the finding class (plan_spec.closure_after_disposition) at initial synthesis and at later dispositions: GREEN and note-only REVIEW_REQUIRED close immediately in either enforcement mode; outstanding need_evidence closes through a disposition-only plan_task call (no model call, no cost: accept = answered, reject, defer = deferred openly; an answer reaches the reviewers on the next paid cycle, a revised envelope supersedes the wave with its open requests); a below-quorum blocking finding stays open; REVISE_PLAN is never closed by disposition — a subsequent paid delta review may evaluate a changed spec or a justified rejection when another cycle is available. A closed note-only wave still accepts voluntary review_disposition annotations through the same writer, spec/verdict and paid-cycle count fixed and the immutable predecessor retained: optional reasoning history, not a new plan gate, never another panel.

Paid cycles are bounded by the shared OUROBOROS_REVIEW_MAX_CYCLES; cycles_paid and the typed cycles_exhausted state count PAID cycles only. A wave is paid iff at least one reviewer slot was physically dispatched, and the proof is a settled row: a slot released at the barrier is $0 until its row proves the send, so the cycle is paid at collection, never at the barrier, and a nothing-dispatched wave of typed $0 skip rows stays unpaid and never replaces a paid predecessor. A barrier-dispatched panel is nevertheless committed money: while its wave is custody-pending and the cap has no room for another panel, a revised envelope first collects EVERY other custody-pending wave at $0 (plan_review_collect.collect_before_supersede) and, if the cap still has no room, is HELD with the typed PLAN_REVIEW_IN_FLIGHT refusal (plan_review_collect.in_flight_hold) so the current wave stays collectible (the per-case cap arithmetic: plan_review_collect.py docstrings). An identical envelope replays the recorded wave free, with one exception: an open wave whose blocking findings all carry valid reject dispositions may dispatch exactly one subsequent paid delta panel when another cycle is available.

Before fan-out the engine captures one panel health snapshot (subagents.route_health, route-level evidence): a slot with positive structural evidence of a spent lane becomes a $0 typed skip row that stays in the denominator; unknown health dispatches (fail-open). The wave records the health epoch and reviewer-roster fingerprint, and a recorded DEGRADED wave replays free only under an identical envelope, matching epoch and unchanged roster — otherwise another paid panel needs remaining cycle capacity (replay/epoch casuistry: plan_review.py docstrings). When the wave's own typed rows prove the quorum structurally unreachable, it carries quorum_unreachable plus the earliest recorded reset, and under blocking enforcement the finalization gate RELEASES while the review stays open: the agent may finalize blocked_with_evidence (reason plan_review_quorum_unreachable), wait through a one-shot schedule_followup, or ask the owner — the host adds facts only, never an answer template.

Under blocking an open wave otherwise holds implementation and an exhausted cap escalates as typed review_cycles_exhausted (free disposition and exact pending custody keep their rules); under advisory the agent may proceed with the wave open, after typed owner-visible plan_review_advisory_open events (one per announced outcome: the dispatch snapshot, then each settled failure with its code and reported cause) and with the open wave recorded in the task's state and result; the model-facing guidance, the plan_task description and the cap head say that the agent's own answer is where the open wave is stated, never that a host narrator will state it for it. A forced rail hands its ONE model call the same limitations as typed key=value facts inside _prepare_forced_prompt, so they are priced by the existing wrap-up reservation probe and no no-call rail gains a call. A hold the agent was told about that the owner's hurry then released reaches it once before its last word; every other gate transition already arrives through the agent's own plan_task result, the settled-wave task message, or that forced prompt. Unavailability, invalid state, budget refusal and deadline rails stay typed non-authoritative attempts, never substitutes for GREEN, and a reviewer request the host cannot attach is never refused before dispatch: the only $0 exits are those attempts, not_dispatched slot rows and the pending_dispatch rows of a barrier-released wave. Every disclosure states the wave's CURRENT state, names the late-result fact while a slot can still settle, never calls an open wave ended, and renders a rail with no recorded reason as absence, not an internal token.

Deep self-review

The retained report (memory/deep_review.md) is historical evidence, not a current-tree verdict: it records its actual generated_at time in the provenance header, and source_revision=unknown is explicit because these deliveries capture no frozen reviewed Git revision — neither a live HEAD nor a file mtime can fill that gap. Context quotes a disclosed 8,000-character excerpt, labels unknown legacy provenance, and leaves the stored report intact.

A task with type=deep_self_review bypasses the ordinary tool loop and calls deep_self_review.run_deep_self_review on its configured deep_review row. An absent row is synthesized as API from OUROBOROS_MODEL_DEEP_SELF_REVIEW; the tool and agent read the effective row. Every row retrieves, with two deliveries under ReviewRequest(surface="deep_self_review") and the shared execution seam:

  • Native inspection: every api_chat row, bare or configured-subagent, runs NativeToolRoundReviewExecutor with the repository root and real runtime data root. The task carries the role/method, BIBLE.md and standing disclosures inline through the shared governance tiers, relevant rules within the inline share and reference-book navigation. Bounds are the route-sized working view, owner deadline and logical window, and paid ledger; exhaustion returns the retained draft with its actual incomplete end.
  • Delegated session: AgentSessionReviewExecutor receives the same task and report contract (markdown, not an output schema), and reads through its harness. Inline delivery is recorded separately from the weaker or unobserved tool-reading provenance.

Both deliveries inline the seven-file memory whitelist byte-exact: identity, scratchpad, registry, WORLD, full knowledge index, patterns and improvement backlog. Every entry carries inlined, missing, empty, oversized or read_error in the task, deep_review_memory usage and the report header. Memory is the subject of this review and is supplied directly, not left to discovery. Matching inline governance is delivered_inline; additional native coverage comes from executed receipts matched by opened root/path and source identity, with read, partial(fraction), missing and unobserved stated honestly. Reading gaps remain diagnostics and never invalidate a finished report.

Every delivered report has a sanitized host provenance comment and human line naming delivery, model, memory and omissions, coverage, incomplete status and attestation; native rows additionally carry rounds, calls, receipts, end reason, transcript and landing facts. incomplete reflects the delivery's actual interruption, independently of diagnostic reading gaps; session completeness stays unobserved. Each outcome (responded, empty or exception) records the row's last execution and memory fact, and persist_call retains prompt/response custody. Typed deep_self_review_unavailable or deep_self_review_error failures return execution_status=infra_failed, reach task/error records and leave the previous memory/deep_review.md intact; BudgetExceeded propagates to the agent's budget-pause rail.

Availability follows deep_review_route, never a window floor. An API row needs credentials for its actual routed model: a stored openai/<slug> can resolve to direct openai::<slug> when direct credentials exist and OPENAI_BASE_URL is unset, while a product -pro slug resolves to the direct default. Session rows require a healthy delegated route. These reviewers have no mutating tools and run no plan, acceptance or commit reviewers: the report is diagnostic memory, not implementation or publication authority. Retrieval enables a targeted whole-system survey across successive views without confusing an assembled packet with measured consideration of the whole system; exact sources remain accessible, and the record distinguishes supplied inline content, observed reads and unknown coverage.

Post-task reflection

Typed root post-task triggers decide whether a run warrants Experience Review. reflection.generate_reflection sends the Light route one open prompt with the EXACT initial text (never a prefix) and its host-recorded task_inputs.run_origin beside it (provenance, never by itself the accepted requirement) plus tool-use, error, review and child projections and the same frozen non-final cost snapshot the task summary uses; it runs outside the tool loop, records its own usage, and its failure never erases the delivered result or changes a review verdict. Its execution trace is the ALL-CALLS listing (build_trace_summary(all_calls=True)): every call in order with every argument, identical consecutive calls folded into one ×N row with their rounds, the first line of a failed or repeated call's result, and one header count of rounds whose every call was non-ok — no positional window and no literal cut, because the consolidation seam fits the call to the Light route whenever that route's window is known (an unknown window sends the prompt unchecked — the accepted residual of enlarging it); the STORED trace_summary (task card, parents, children) stays the bounded two-argument preview. The trace row carries the round_id of the model round that issued the call (absent when unknown), when the listing really cut an argument value or a failed/repeated call's answer, the redacted per-call record is retained through retain_memory_source and named in the prompt as OPTIONAL reading (never a required source), and its claim states the STORED bounds on both axes, not completeness: an argument already passed sanitize_tool_args_for_log (an oversized value carries a marker with its length and sha) and a result is the stored actor-visible cap — more than the listing, which shows only the first line of a failed or repeated answer — a partial one naming its own FULL_RESULT_SOURCE_JSON or FULL_RESULT_SOURCE_UNAVAILABLE, with a call's recorded manifest named only when it has one — claiming results "in full" or an unconditional manifest overstated a cognitive artifact. Cut detection reads the ONE shared marker list (artifacts.SANITIZER_OMISSION_MARKERS), because a width test over already-sanitized args measured the widest argument in the task as a small one and retained nothing at all, while a hand-rolled subset missed the _repr and _error shapes whose arguments survive only in the call blob. Unavailable source retention is disclosed, error details group by full redacted content before display clipping, and post-task synthesis — the reflection, its Pattern Register update and the episodic summary — thinks at the owner's Task / Chat effort (settings_scales.resolve_effort("task")), never a literal. Admission to the Pattern Register is typed, not a word scan: it opens on a call the loop recorded as errored (its stamped tool_result_code, or the recorded status for a legacy row), on a producer fact naming a failure the ok status cannot carry (a preserved commit whose post-commit tests failed publishes post_commit_tests), on typed codes already stored with an entry, or on a genuinely FAILED child — a cancelled, soft-landed best-effort or degraded child is not a failure. Those failed-child classes reach the root through the child evidence the synthesis walk already collects and make the run error-bearing for both the trigger and the prompt's error details: children do not reflect, so a short clean root that delegated the work is the only place its child's failure can be learned from at all. Deliberately not "a reason code exists", which would open the register on every terminal.

A reflection lands where it durably belongs: a non-project root appends the full entry to the canonical logs/task_reflections.jsonl; a project-scoped root appends the full entry to its project drive and the canonical log receives only a bounded pointer row — full project text never enters the canonical log, which feeds future global context. A project-bound task's context includes a bounded labeled tail of its own project's reflections; the headless mirror drive of a split root is never the reflection home, and the Pattern Register update stays canonical in both cases. Every entry carries task identity, evidence, lessons, backlog candidates, and validated memory actions. MEMORY_ACTIONS_JSON permits only scratchpad_append, knowledge_write, and identity_update_candidate, at bounded count and size. apply_memory_actions routes accepted actions through provenance-preserving memory and knowledge APIs. An identity_update_candidate is recorded in the scratchpad for review and is never auto-written to identity.md. For a project-scoped task, reflection applies knowledge actions only (the project store by default, explicit global allowed); its scratchpad and identity-candidate actions are skipped because this automatic Light pass lacks the conversation's full view — the conversation writes identity and scratchpad from any room through its own tools. Reflection may propose a future campaign or backlog item, but it cannot enqueue, review, commit, or enable one.

The Pattern Register writer (reflection._update_patterns) REPLACES the whole document, so every decision input it reads is complete — the full current register, the exact initial text (goal_exact, beside the bounded goal display) and its run_origin and the whole reflection text; a prefix of a decision input can never authorize the rewrite, because a clipped clause can record the inverse of what the reflection concluded. patterns is a reserved global-only topic alongside improvement-backlog and overview: whichever room writes it, the register has ONE home on the canonical drive, the only one its writer and its readers (context assembly, deep self-review, the headless copy) address. Re-read, exact compare, history append and atomic replacement are one critical section outside the Light call; a register that moved under a losing writer is preserved and that task's learning is DROPPED with a warning naming the task, never retried against a source it did not decide from. A row's count is bumped once per observed episode, so two roots of one owner request bump it twice — the number counts episodes, not distinct requests.

The advisory rows a reflection or summary reads are ATTRIBUTED. Advisory runs are scoped by repository and several tasks legitimately review one checkout, so collect_review_evidence renders as recent_advisory_runs only the rows this task owns plus legacy rows carrying no owner (which stay unknown, never re-attributed); another task's rows reach the prompt under their own heading naming the owning task ids, and every row carries its owning task, attempt, phase and complete snapshot hash. Repository readiness (current_repo, open obligations, commit-readiness debt, the exact-snapshot match) stays repository-scoped, because that is what "can this checkout be committed" means. The split is keyed on ROW IDENTITY, never on a scope key: an empty repository key widens the candidate list to every advisory run on the drive, and the installation's history is not one task's record.

Only roots synthesize; root_phase_checkpoint makes paid synthesis at-most-once across restart, while children contribute evidence. Durable-result persistence owes the answer as final:<tid>:<digest> in supervisor/terminal_delivery.py's bounded outbox (§5; normal/cancel/reap). send_message delivers immediately; the retained buffered copy shares its ID for durable dedupe. Replays use bounded backoff. Exhaustion/eviction preserves full text on disk, emits terminal_delivery_exhausted and a chat notice; external delivery remains at-least-once. Buffered task_done stays last to retain the slot/child drive during synthesis; a hung-synthesis reap need not lose the delivered answer. Project roots keep early answers in Project. Their canonical row and deferred Main mirror use terminal_projection.settle_terminal_projection via task-done/checkpoint/startup/maintenance; §3 "Main rows and host-stamped card rows" owns readiness, retirement and limits.

Synthesis receives a sealed final package from the durable result — the submitted final text, its artifact manifest and completion_observations. Full redacted action observations live in the canonical artifact store (task.budget_drive_root or drive_root), in the write-once source_handles/context_checkpoints store with verified task_source refs, before compact publication and outside deliverables and inferred readiness; their native reader get_task_result(include_completion_source=true) returns complete length/hash first, then explicit source_start_char/source_end_char ranges (artifacts.text_source_range_projection, the shared work-order range contract), with bytes, kind, path containment and SHA checked before any excerpt. Packet-only summary/reflection receive per-send-tool counts, each family's latest recorded return, and task-related skill readiness with coverage; full-source references are for later readers, not evidence the synthesizer has read. Positive observed facts correct error-trace impressions, while tool success does not prove owner receipt, empty material does not prove absence, and skill readiness does not attribute an owner's action to the task. Before context cleanup, agent_task_pipeline.emit_task_results also freezes review_evidence.task_inputs through post_task_synthesis.capture_task_inputs: run_origin, the existing task-local owner corpus, intact question/answer provenance and the canonical split-root verification-receipt union. Summary and reflection receive the same complete redacted content through reflection.task_inputs_prompt_section, separate from bounded trace/review excerpts. A zero return code is positive evidence; an unrelated later pass cannot resolve another check's failure. Peer proposals stay attributed, and unavailable input is not evidence that approval or verification never existed. Recovery uses these stored observations and inputs, not a later conversation. Summary trace pointers name the task and existing archive-aware reader (ouroboros tasks watch <task_id> --jsonl), not guessed flat log paths.

Pooled workers retain their slot until root post-task synthesis settles, for API-only and subscription tasks alike; early final-answer delivery keeps the response independent from that queue timing. Ordinary native post-work, including an inline Presence turn after its durable result is returned to the adapter, stays on its registered actor thread without a pooled worker slot: its TaskModelWait owner remains available through POST_TASK_SYNTHESIS_INFLIGHT after ordinary dialogue admission closes, detached server post-work binds its own live owner in the same registry, and the task mailbox stays available until the terminal post-task checkpoint. An open phase remains finalizing rather than appearing completed. Typed quota/auth waits resume only the unsettled call, stop or unknown outcomes degrade the phase without repeating finished stages, and restart recovery degrades an indeterminate running phase rather than replaying a possibly paid request.

Project registry and lease

A project is a focused working room, not an isolated sub-mind (§6 Durable memory and project focus). The projects registry (projects_registry.py, data/state/projects.json) owns immutable project identity, canonical chat id, optional working directory, lifecycle/tombstone state (active|deleting|tombstoned), routing generation and activity revision; admission persists the resolved project id in the task itself. An id minted from a DISPLAY name collapses dash runs and carries a short deterministic suffix when the name held characters the slug could not keep, so two different non-Latin names cannot share one project; the normalizer is unchanged, so existing ids are never re-slugged and stay reachable by their explicit id. project_lease.py serializes assignment of pooled roots by Project while allowing their own subagent trees; it is not a physical-folder lock and does not withhold tools from ordinary conversation. Binding/history files support routing and presentation, not the lease. Delete closes routing, cancels/quiesces the tree, and tombstones only after settlement, preserving everything for recovery.

Project binding by task and by origin

The durable binding is the SINGLE truth about a task's project: the in-task scope guard and project_facts.resolve_project_id read it FIRST, ahead of task["project_id"] and a worker's in-memory ctx.project_id, which are copies a mid-run conversion never reaches — a guard reading only the copy is how a task already bound to one project can mint a second, empty one. The same store also answers BY ORIGIN (projects_registry.project_id_for_origin, keyed on the ingress-captured (chat_id, client_message_id) stored in every binding's source_ref), because one owner message spawns several task ids — the turn that received it, the root it promoted, the timeout retry that replaced that root — and the project its WORK has belongs to all of them: a promote with no explicit target inherits the promoter's binding, then the origin's, before the in-memory scope copy, and its admission resolves the origin once more. A timeout retry is bound by the reaper inside its retry admission transaction (worker_promotion.bind_retry_to_origin_project, called by task_reaper._run_retry_admission_transaction): the predecessor's binding, then its origin, answer, and the retry reuses the stored origin by value, so retried work keeps its room instead of arriving in Main as a second convertible unit. That bind happens only after cancellation has lost the admission boundary, because a binding is immutable and a bound-but-never-admitted retry id would answer project_id_for_task forever; a retry suppressed by a cancelled or already-terminal root is never bound, and a refused bind or unreadable store leaves the retry unbound and discloses project_binding_failed rather than holding up the retry.

Every IMPLICIT claim — the UI conversion, that admission, the reaper's retry admission and the in-task ensure_project_scope — holds one process-local claim lock (projects_registry.origin_claim_lock) from the read of which Project the origin already names through to its own durable bind, while a naming model call stays OUTSIDE it, so two cards of one message clicked inside a naming window yield one Project and two bindings. Explicit project_name/project_id/route_to_project remain the model's choice (P13): a sibling that names a different room gets it, and the message's existing Project is never renamed from a task that does not belong to it. Legacy state where one origin names several active Projects resolves to the one whose task is still live, else the latest binding, disclosed as project_origin_ambiguous; continuations that carry no owner origin (schedule_followup, direct auto-resume, Presence promotion) stay per-task. An unreadable bindings store is disclosed once and read as "no binding" on every path: the refusal is reserved for the measured case, a readable binding to another project, and on BOTH conversion paths it precedes every side effect — no project row, lease mark, broadcast or announcement survives it: a UI conversion that named a DIFFERENT project id answers 409 naming the bound project by id and display name, the one-click conversion adopts that project instead, and a bind refused after the lease mark restores the lane with the same 409, so no lane keeps a project the binding does not name. One residual survives: a conflicting bind landing between the final re-read and the bind leaves behind the empty project row the request had already created, because the registry deliberately has no primitive that removes a row (delete tombstones the id permanently, which would cost the owner that id and display name forever); the row holds no task and no binding, and the owner deletes it like any other project. The lease mark is fill-only, except that a conversion which already holds the binding it is about to write moves the lane onto that binding.

In-task project scoping

ensure_project_scope (tools/control_delegation.py) can create or bind the current root to one project mid-execution: it persists the durable registry binding first, then marks the live queue/lease surface under the queue lock so the lease recognizes the running task as a lane occupant; it is idempotent for the same project, and a task BOUND elsewhere is renamed, not re-scoped — its scope call carries the requested display name to the project it already belongs to and creates nothing. The act rides the same receipt rail as the other routing verbs (§6 Task lifecycle) under its own synthetic agent-steer:<token> id: the tool scopes itself in memory at once (journal_write targets the project meanwhile) but its RESULT states only the durable outcome the supervisor recorded — delivered with the binding it re-reads, a typed rejected (bound elsewhere with the rename outcome, a refused bind, a registration failure), or unconfirmed when no receipt landed within the bounded wait; persisted on a routing receipt is true only when a row was written. A late bind after a lost acknowledgement is discoverable from the durable binding, which the next same-project call reads as already scoped. A planning obligation stays with the task on ensure. A project-SCOPED but unbound run keeps the older refusal, since there is no durable project to rename, and a child still cannot escape the inherited scope.

Durable memory and project focus

context.py assembles static governance, semi-stable memory, and dynamic task evidence without treating truncation as forgetting; the recent-activity sections are each task's OWN newest rows (progress 50 rendered; tools 20 selected, 10 rendered and 20 scanned for review markers; events 200 counted by type) through the bounded reader jsonl_tail.py (Memory.read_task_recent: a doubling live tail plus at most three newest archives), never a global tail filtered afterwards (issue #131), and their header's coverage line names the rows, the window and any unopened archives while read_file pages the rest; a subagent child gets the same three windows beside its ## Working sources block, its tools and events read from its own execution drive (its worker rows; host-side rows such as waits stay in the canonical log, as the header says) and progress from the canonical log; the Development context matrix and context_layout.py own which reference form is resident. When the rendered scratchpad exceeds SCRATCHPAD_SECTION_BUDGET_CHARS, context.py keeps the newest whole blocks that fit and drops the oldest behind an in-band gap marker naming memory/scratchpad.md as the live source; no block is retired by a context build, and scratchpad replacement keeps its explicit summary and source-journal provenance.

consolidator.py publishes a block and advances its generation-aware cursor only after complete room draft and correction; a missing generation appends [MEMORY GAP] instead of resetting. context_fit measures Light against fresh route/account capacity and calibrated density; llm_local owns the local output reserve, and absent evidence stays unknown. room_consolidation.py processes each room separately and assembles sections deterministically. Both knowledge stages receive the entire current note and source episode; the corrector's complete read, not the draft's, binds revised entries. Older episodes cannot negate newer facts; model judgment governs supported corrections. Source range reads and CAS preserve old/new history. Startup compaction is a no-op; earlier digests cannot be reversed. A failed correction withholds the chunk and cursor.

Oversized sources split without clipping, including within an entry. consolidation_retry records source hash and a smaller same-route bound, invalidated by source/route/capacity/reserve changes. Era compression regroups rooms deterministically; failed eras retain blocks and legacy provenance stays unknown. last_consolidation_error clears after a failure-free advance. pending_knowledge_nominations records each source entry BEFORE note publication; an unrelated successful batch cannot remove one. Legacy last_unpublished_nominations persists. Health shows three distinct abbreviated source+position IDs and omitted count; full proposals live in knowledge_history.jsonl. Unreadable meta preserves debts and warns in Health; Nano pressure records no-progress instead of aborting Main. Invalid legacy receipts warn separately. Old digests and debts need explicit resolution. Spend remains nullable; model-control errors follow propagate_model_error.

Consolidation labels every chronological source message through dialogue_provenance.RoomLabelResolver, using the actual chat_id, never lineage project_id; one read-only registry snapshot supplies the window. Main is named only for the actual Main id, a resolved project uses its current registry name and stable chat id, and missing, unknown or ambiguous rooms stay explicit. Ephemeral formatter offsets carry the original room/author/direction/transport header into split continuations without parsing message bodies or duplicating their bytes. Room draft and correction prompts require meaningful decisions, approvals, outcomes and unresolved commitments of that room, retaining source distinctions (who decided, what was authorized, what stays owed) and one first-person Ouroboros voice. Length adapts to content within the existing output-token ceiling; no per-room word quota, semantic gate or absent room is imposed. Labels establish provenance, not summary success. The mixed Main recent view opts into the same labels; focused Project rendering, membership and explicit chat_history retain their existing behavior and bytes.

knowledge.py owns note reads, revision-checked writes and indexing; tools/knowledge.py exposes them. knowledge_write(mode="edit", old_str=..., content=...) replaces ONE body occurrence under the current revision, preserving other bytes and frontmatter. Missing/ambiguous anchors, absent notes and stale revisions refuse before writing. Character/heading deltas reach results, bounded there but complete in history. Edit adds storage capability, not a semantic writer policy; overwrite/append stay unchanged. Blank revision creates only a missing note. Malformed legacy preambles stay readable with uncertain metadata; missing linked notes remain unwritten.

Ouroboros remains one identity across Main, project rooms, and Background Consciousness: unified dialogue memory remains available to the one agent, while an executing project task preferentially receives its own thread, journal, workpad, and project knowledge. project_facts.py routes project facts to projects/<id>/knowledge; subagents inherit the root's resolved project id and never derive a new one; identity and the scratchpad are one canonical pair written from every room, with no per-project copy. Project journal.jsonl records curated milestones and workpad.md retains active working context (tools/project_journal.py); focused context includes the workpad in full and recent journal rows with a visible pointer to older entries. On root completion, only high-signal blockers, questions, and interface contracts are mirrored once from the ephemeral task-tree ledger into the durable journal, and a finished root whose effective working tree is not the registered working_dir writes one typed "work lives at @ " journal row from facts the task record already holds. A project digest gives consciousness a concise completion signal without pretending to be the raw project memory.

Skills and extensions

Skill capability grows through independent gates, each with its own module: discovery and manifest parsing (skill_loader.py, contracts/skill_manifest.py), content-hash-bound review (skill_review.py / skill_review_runner.py), owner grants, dependency reconciliation (skill_dependencies.py, marketplace/isolated_deps.py), enablement, readiness (skill_readiness.py) and execution (tools/skill_exec.py). Discovery establishes identity, source, provenance (.self_authored.json), conflicts and hash; it confers no trust. Review status, grants, enablement and dependency health stay independent durable facts under data/state/skills/<name>/; the gate sequence, payload buckets and marketplace/hub layer are §13's. skill_readiness_for_execution() composes review, hash, enablement, grants, dependencies and peer conflicts into phase-specific next actions and reads the mode-aware gate_for projection (executable_review), never a raw verdict string. The owner may attest their own skill, or a hash-verified official hub payload, to skip the expensive LLM review (skill_owner_attestation.py); the deterministic preflight floor still gates it. The model-facing catalogue (list_skills, the per-turn Installed Skills section) projects the same live extension facts as /api/extensions; each enablement change best-effort appends a typed skill_enabled_changed row to logs/events.jsonl naming the actor (empty for writers not yet labelled), and an append failure is logged and never blocks the enablement change. Cyber review failures remain visible advice; visibility, successful execution and review PASS are still different facts.

A skill may declare a reviewed presence: behavior profile: instructions, knowledge topics, runtime defaults and portable capability requests that installation-local selections resolve to exact targets. Admission requires the bound behavior skill installed, enabled, freshly reviewed and complete for every required request, then compiles one immutable positive capability ceiling copied through task_contract, so mutable skill/Settings state cannot broaden a live turn. An optional owner-local workspace_root (chosen through configure_presence or the Skills settings form, validated by workspace_admission.validate_workspace_root, frozen into the task's workspace contract) changes the file/process target while retaining canonical shared memory; it neither derives Project scope nor forks the data drive, and later edits and promotion/follow-up tasks keep the admitted folder rather than rereading mutable settings. A folder that has become unusable refuses the promotion typed (workspace_unusable, the repair in detail naming the profile) and writes no chat row — the public conversation's synthetic chat id is never spoken to.

extension_loader.py and the isolated-dependency layer load only a ready, hash-matching extension — in-process via PluginAPIImpl, isolated-dep/native ones as child-process proxies — and a review PASS alone proves neither dependencies installed nor a widget loadable. Per-skill health at data/state/skills/<name>/health.json is process-qualified: the server's observation is authoritative, a worker's is a handoff-qualified view.

skill_lifecycle_queue.py (§13) exposes queued/running/succeeded/failed plus stale metadata — stale is recovery evidence, not a fake unlock of a still-running thread. Scheduled work is reconciled by resync_skill_schedules() and runs only after skill_readiness_for_execution(); schedule evaluation is a DST-aware system on the shared cron/timezone contract. Evolution remains hard-blocked in light runtime mode (§5). The existing payload binding (skill_payload_binding.resolve_skill_payload_base) fills an omitted skill name or bucket from a valid selected normal/repair TaskConstraint; explicit selectors still have to resolve to its same physical payload, and collision, native-mutation and child-profile restrictions remain independent. A selected-skill repair (skill_repair_admission.py, §13) is admitted against an immutable base_content_hash and verifies the observed expected_content_hash before each payload write, advancing it after opaque process work without claiming authorship; drift makes the repair STALE — no long shell lock, no rollback, other writers never blocked. These separations let skills expand capability without turning discovery, a UI toggle or old review state into execution authority.

Skill publication

The passive installed-skill projection never launches Betterleaks and never claims the current bytes are publication-ready (skill_publish_eligibility.py). Selecting Publish calls POST /api/skills/{skill}/publish-preflight (gateway/skill_publish.py, read-only), which resolves one current payload, captures and scans its bytes, recomputes review staleness and returns exactly one backend-authored state (ready, warnings, needs_attention, repairable, hard_block); the browser only renders those facts, and only hard_block prevents task creation. The authoritative flow is:

passive index (no scan) → selected preflight → explicit confirmation → ordinary managed task → immutable current capture with separate review provenance → payload scan → GitHub read-only planning → derived-output scans → first GitHub mutation → validated same-skill pull-request receipt → ordinary acceptance

Publication validation and immutable capture share skill_publish_eligibility.publication_author_acceptance: fresh publishable critic authority or qualified current Advisory author authority. Blocking keeps fresh critic requirements; owner attestation alone is not that independent-review/author chain. The host PR checklist states original critic and current author/published hashes, referenced review basis and author rationale, without claiming a second independent review. Recorded findings retain their actual severity, including critical and partial-quorum findings; catalog hashes and PR receipts never become review PASS. Passive and selected preflight use the same authority, with current captured bytes authoritative in selected preflight.

Every outbound byte derives from the capture (skill_publish_snapshot.py); the mutable live payload is neither reread nor rehashed to authorize the transaction (no time-of-check/time-of-use drift). Outside Cyber Pro, literal Betterleaks high findings (skill_publish_scanner.py) block the current outbound call; lower or unknown confidence is a redacted warning. Cyber keeps scanner findings/failures and review status as independent advice; an unavailable scan does not invalidate an observed publication. Packaged installs resolve the bundled betterleaks-standalone; source checkouts resolve the managed runtime installed explicitly with python -m ouroboros.betterleaks_runtime install — Publish never downloads it. A top-level skill_publish task (tools/skill_publish.py) is accepted only when pre-truncation metadata holds a validated pull-request receipt for the requested skill and configured Hub repository; the receipt proves the observed publication and never manufactures review PASS. A definite failure before any submission-branch request fails an explicit Publish objective even under degraded review; a last-confirmed stage never proves the next request unsent, so a previous unknown branch/PR attempt or valid same-target receipt is preserved.

A failed envelope (skill_publish_result.py) names its cause beside the stage: reason_code (the failed stage — fork_sync_failed and siblings), a repair_hint chosen from producer evidence, and the transport's sanitized error_detail with github_status/github_operation read only from gh's own error shapes, never inferred from prose. Fork synchronization is mandatory before any branch, commit or PR mutation; ambiguous PR settlement is a read-only exact lookup. Publication metadata is read from the producer JSON before the shared ToolResult host-note separator, so appended route or safety notes cannot hide diagnostics or a receipt.

Owner lifecycle actions share skill_lifecycle_actions.run_skill_action: grant/toggle run in the lifecycle lane, local delete uses skill_uninstall_state (uninstall tombstones and authorized local deletion keep separate retention), and attestation keeps its deterministic floor in skill_owner_attestation. The host supplies actor identity and checks exact resource/revision plus an existing member chat message, answered quiz or owner-mailbox record when owner intent is needed; the model interprets that source, and a generated edit-and-review request confers no grant, attestation, deletion or enable authority. An enable with no explicit chat/quiz/mailbox source resolves the original owner source itself (a client allow_enable flag authorizes nothing, Presence ceilings come first, a newer owner disable in enabled.json wins), and a Repair review never auto-enables — the model chooses the ordinary toggle_skill call. There is no permission ledger and no HTTP impersonation; the auto-grant policy stays separate from enablement, and an explicit owner disable survives review and free replay. skill_exec returns the revision captured before its script launch, and extension tool receipts carry the descriptor's content_hash from the same publication as extension_generation, only after physical dispatch; neither is review PASS or a semantic test verdict.

Marketplace update keeps the owner's selected version through retries; install.PayloadRollbackSnapshot captures payload/environment and the affected lifecycle-state quintet for update and adopt alike, and the single restore path verifies the required reload before claiming rolled_back, preserves independent enablement/history, and never deletes an already restored payload after a failed state write. Catalog updates disclose exact version strings and infer no ordering. The atomic publication-record owner (marketplace/provenance.py) also clears locally: it compares the displayed published object before setting that section to null, preserving unknown siblings, and changes no GitHub PR or installed bytes.

MCP and browser-facing external tools

mcp_client.py owns configured HTTP/SSE and local stdio MCP discovery and invocation (HTTP surface gateway/mcp.py; tools are named mcp_<server>__<tool>). HTTP/SSE entries validate URLs and auth headers, and URL userinfo is masked through the shared secret projection (secret_masking.py owns the exact MCP token placeholder shapes). Load-time placeholder repair runs before environment precedence, so a real environment credential is never mistaken for a wire mask, and only for recognized top-level Settings secrets — nested MCP values are never silently migrated. Stdio entries pass one executable command and an exact string args list to the MCP SDK without a shell; optional cwd, literal env and env_from_settings (environment name → existing setting key) resolve from the same saved configuration for discovery and calls. Unknown fields are retained with a visible not-applied warning while a valid server stays usable; invalid known fields or references produce MCP_CONFIG_ERROR; the response-only auth_configured flag never becomes configuration. Discovered tools join the selected initial capability envelope beside enabled, granted extension tools, behind the same network resource gate as on a managed task (schemas(), get_schema_by_name and execute agree on that), and a discovery failure is an explicit capability omission through list_available_tools, never a silent removal. The catalog lookup (MCPManager.resolve_tool_name) precedes the paid safety check: an unlisted name is UNKNOWN_TOOL, a disabled server or an unlisted catalog (health unknown) MCP_UNAVAILABLE, an allowlisted-out tool ACCESS_BLOCKED; none runs safety, transport or a refresh; hits recheck settings before Safety to honor revocation. Hints name only the addressed server's raw → callable pairs and an exact naming-rule identity (mcp_client.naming_rule_matches); nothing is ranked, aliased or dispatched, and raw names stay out of schema descriptions. Cyber metadata-target and configured allowed_tools filters do not veto an existing MCP call; descriptions and results remain untrusted data; every call still crosses registry, resource, safety, timeout and result-handling policy. Web-tool prohibition and network prohibition are distinct: web=false alone does not disable configured MCP or extensions, while an explicit network=false and tool disables are enforced at discovery and dispatch.

Browser tools are stateful and thread-sticky because Playwright sessions and greenlets have thread affinity; they cannot be scheduled as ordinary parallel stateless calls. A stateful-tool timeout therefore RETIRES the browser generation: the shared browser_state slot is replaced and the close runs on the owning thread once the hung call settles. Retired sessions are bounded in-process at _RETIRED_GENERATIONS_MAX per task, after which another session is a typed BROWSER_BACKLOG_RETIRED_SESSIONS refusal; generation isolation is best-effort under concurrent replacement (a fully closed class needs a process-isolated browser worker — disclosed future design). Model-driven in-page evaluation runs the supplied expression without substring guesses through _evaluate_bounded: Playwright's evaluate accepts no timeout, so the expression is raced against an in-page rejection, which bounds the ASYNC class honestly and no further — a synchronous event-loop block cannot be interrupted from inside the page, and the outer tool timeout remains the backstop. The caller's timeout also becomes the session default (page.set_default_timeout), floored on the action path so the five-second action default cannot strangle a capture. Chromium is the default; WebKit and device descriptors are targeted tools for a real Safari/iOS risk, not a universal acceptance matrix and not a claim that a narrow Chromium viewport is Safari-equivalent. First-party PR helpers are normal built-ins whose mutating operations stay subject to selected-root policy, runtime mode, delegated-child constraints, credentials and reviewed-publication authority.

Browser target policy is browser_policy.py: ordinary owner-control admission uses the actual HTTP request and three-valued service identity — never JavaScript substring guesses, no per-network-request LLM, no second policy engine. Cyber root/acting requests retain target and control reach; explicit read-only assignments keep their restricted contract. A live binding in state/server_port.bindings.json proves an endpoint ours; a service's missing snapshot can mean an older installation or failed publication, so its recorded facts still name what is EXPECTED — the integer state/server_port (the launcher's state/server_process.json can prove main), the Host Service configuration beside an expected main, and the local model's custody row with its live argv — and a matching expected endpoint whose process cannot be verified is refused as unknown, never treated as foreign (server_process.runtime_service_identity). Only a live snapshot of the same service supersedes its legacy expectation: main cannot vouch for Host Service. Unmeasurable POSIX identity never counts as proof. Every other port is an ordinary target, an unrelated application reusing an /api/owner/... pathname included — the owner-operation request shapes apply only at a proven or unknown Ouroboros endpoint — and a restricted target that DNS cannot classify is a typed BROWSER_POLICY_UNAVAILABLE refusal before navigation. Chromium and WebKit follow HTTP redirects natively and Playwright route callbacks see only the first URL of a chain, so tools/browser.py re-checks each document's redirected_from chain with the same target predicate before any page result: an allowed→blocked→allowed redirect withholds its content and refuses actions on that document, while the forbidden hop's request has already been dispatched — an owner-accepted, disclosed residual, not a pre-request or DNS-rebinding guarantee. Cookies, POST semantics, service workers and native redirects are untouched: no proxy, CDP fork, pre-probe or refetch.

Budget tracking

swarm_efficiency.fanout_count counts observed fan-out emissions and fanout_interval_sec_total sums the wall-clock gaps between them, intermediate parent work included; new events emit fanout_interval_sec, and readers tolerate the retired wave_count, inter_wave_latency_sec_total and inter_wave_latency_sec without rewriting stored bytes. These observations never infer semantic work waves or child-wait time.

The time/cost/intrinsic pacing checkpoints carry resource_facts: incremental per-tool call/error counts and producer-reported duration intervals (which may overlap; unmeasured calls stay unknown), plus own-task, tree and delegated-tree ledger buckets (own and delegated explicitly overlap the tree total) and the shared unreserved global remainder. No argv, stdout, sleep/poll classification or new stop rule selects behavior, and prompt and checkpoint show the same facts.

usage_accounting.py owns monetary policy over the physical-attempt ledger. Each send gets an ID and reserved → dispatched → settled | unresolved, or pre-dispatch reserved → released. Only typed proof of no sent bytes (connection/pool failure) permits dispatched → released; timeouts/unknown errors remain unresolved. Each retry is a new attempt; SDK/stream boundaries are labelled opaque. Main/direct/child/scout/review/safety/synthesis/reflection/consciousness work, retries and opaque SDK calls are covered; root scopes count tree and post-task/review work once. External scripts/extensions with model credentials stay unknown/unmetered at host-observed opaque boundaries absent authoritative settlement.

cost_projection.py owns producer amounts/openness: accounted_upper_bound_usd = settled+reserved+unresolved, null is unknown, finality is never invented, and COST_OPENNESS_FIELDS accompanies amounts. Retired cost_usd[_with_children] stays read-only. cost_presentation binds amount/facts to one bucket (§11.1): reconstruct_task_cost owns own scope; root terminal/heartbeat/checkpoint/synthesis paths use root-tree, never falling from null subtree to own zero. _usage_rows._summary's weighted priced_rows, tracked_nonfinal_rows, accounting_open_rows survive compaction. A price or retained bound evidences zero; empty/unpriced rows do not. Settled unpriced rows leave the known subtotal exact; estimates (zero included), retained bounds, open attempts and integrity gaps stay nonfinal. Terminal refresh/copyback, synthesis, heartbeat and history preserve scope; unknown never becomes sticky-final. Sums and cost_final are unchanged.

A proven abandoned send closes administratively as settled with settle_reason="abandoned", cost_usd=None and cost_final=false: its reservation bound stays accounted, not confirmed spend, and the price remains non-final. A real late receipt may correct the same attempt once (late_receipt); a positive never-started receipt may instead release it through the existing before_dispatch_failed: seam. Ordinary terminal settlements and releases remain immutable. Full and incremental ledger validation preserve the same correction eligibility; compaction retains unresolved and abandoned chains for that exact-id join, while existing historical baseline groups remain aggregates (the full storage contract is Usage-ledger compaction).

After cancellation custody, existing maintenance calls server_maintenance._reconcile_abandoned_usage for settled tasks with no live physical owner/open post-task synthesis; review attempts retain their review owner. llm_claudexor.recover_model_attempt reads retained terminal CAS or the exact operation via owned-only discovery, retains bytes before ACK and never starts/restarts/cancels bookkeeping work. Missing identity, live work or unreadable evidence defers. The money lock rechecks each row so real concurrent receipts win; no age cutoff, timer or ledger is added. Projection reconciliation includes compacted-baseline attribution, retaining retries after failed result writes/native late receipts. One bulk breakdown feeds task/root owners; foreign canonical budget roots stay separate. events_task_done._refresh_terminal_task_cost changes money fields only, without completion events or invented synthesis.

Summary/reflection/consolidation share a frozen ledger snapshot: settled subtree cost, reservations, unresolved bounds/counts, unknown exposure, integrity and non-final state. Final terminal checkpoint alone owns final cost; read failure is unavailable/null, never zero. No reconciliation LLM or parallel ledger exists. Cache-hit health uses only rounds carrying cached_tokens, and stays absent with insufficient samples: unmeasured is unknown, not zero, so providers lacking telemetry cannot dilute measured ratios. Unknown spend remains the existing unknown_unmetered field, not a third accounting status.

root_phase_checkpoint.accounting retains a cumulative ledger observation: root ID, accounted/reserved/unresolved amounts, non-final/unresolved counts, unknown exposure and integrity. Separate from own-cost and read-time TaskCostBreakdown, it survives public detail/recovery and canonical replica protection. The checkpoint owner refreshes it, including explicit null/unavailable on failure. These facts prove neither an invoice nor local paid-work closure; phase/execution ownership is independent.

Planning threshold (task_pacing.py). One unreserved CostCeiling (disabled, active, exhausted_soft_land, unknown) owns runtime in_task_cost_ceiling disclosure and loop checks. Roots use min(global-remainder percentage, root cap) minus margin; uncapped roots use starting wallet. Enabled descendants inherit that number, not later percentages. Disabled profiles disable only pacing; monetary admission stays independent and surfaces name the binding bound. Rooted tasks compare subtree spend/holds even without root caps; own-spend fallback is a disclosed lower bound.

loop_budget._finish_tool_round_budget and _finish_no_tool_round_budget share _check_budget_limits after unfinished spending rounds. No-tool tails, including cold Resume (resume_point.budget_tail), neither advance nanny baselines nor arm tool controls. Candidates survive; READY returns before this check without an extra paid final. Eligible actors pause, even before work on a soft threshold. Ineligible actors retain the priced terminal rail (budget_wrapup_unaffordable if it cannot fit); global rejection before work buys no call. Tool requests remain incomplete even beside replace.

Affordability probes reserve nothing; full-cap ledger admission decides. Unknown spend is not zero; unknown price fails open. _usage_cache_splits.py reuses settled provider/normalized-route/review-surface splits: cache loss reprices full write, expiry may under-reserve one write, competitors may consume room. Proxy stops/native images require exact transcript-copy pricing before service finalization. task_pacing.main_loop_wire_options owns wire options with no later payload additions. Identity drift refuses pre-send (forced_candidate_drift); terminal handling retries once unpredicated, still ledger-priced, so drift cannot lose the answer.

Exact budget pause and Resume

budget_pause.py pauses pooled/direct work under the SAME ID on global exhaustion, graceful ceiling, either last-fit stop, soft landing or refused dispatch, without a paid final. A missing task id/root/continuation owner keeps terminal behavior (exact_pause_unavailable). Increases never wake tasks.

Owner Ordered contract and reason
request_pause, usage_accounting.reserve_attempt Close DispatchFenced; review POST/replay and Light extraction also refuse (budget_pausing_no_send, budget_pausing_no_extraction). Persist pausing before waiting (direct RUNNING stub if absent) so crash custody cannot replay ambiguous work; the reaper derives the crash-retry fence from its one pause-row read (worker_health._complete_exact_budget_pause_after_death), failing closed on unreadable evidence.
local_producer_observation Hold the worker nonterminal until sent reviews settle through review_custody AND timed-out tool futures finish their late settlement callbacks (register_tool_future/hold_tool_settlement); done() alone is insufficient. Failed writes/producers publish budget_pause_hold on change and retry, never claim a pause or buy a final; a publication failing after quiescence retries the SAME prepared snapshot, discarded if a producer revives. Only Stop/Panic/deadline/cancel ends the hold as abandoned, via the loop's model-wait control rails (a no-call terminal, never a task exception); controls read the canonical budget root on split drives.
observe_task_runs Re-read custody (the loop side resolves its own custody root; the grant reads it directly); unreadable is custody_read=failed, not empty. Pre-terminal subscription coverage is unproven, so unsettled runs request cancel_and_verify, retaining typed outcomes. Unknown stop permits no second writer.
owner_wait.continuation_state Retains cognition, opaque acceptance/preparation, usage/rounds/clocks. resume_point keeps the budget tail and unanswered call IDs; Resume inserts execution-UNKNOWN host rows, never replays tools. Source/row precedes BudgetPauseRequested; no task_done/result/Main final.
events_budget.install_exact_budget_pause Validate event/source; park RUNNING or parkable_direct_task in PENDING as _budget_pause, persist snapshot, then mark paused. Root scope fences the tree. Park, projection and grant share the queue lock; pause-id/state CAS turns late publication into budget_pause_park_superseded. Direct records keep _is_direct_chat, omit inline image bytes and release the local fence once the actor unwinds (the durable row owns the hold). Resumed direct work uses a pooled worker; log_addressing.address_task_event keeps the lane. Projection carries ledger cost planes; a live-pause canonical row outranks the child replica and the queue mirror in task_status.effective_task_result.
worker_health._complete_exact_budget_pause_after_death, queue_snapshot._park_pausing_running_rows, workers.kill_workers Complete a saved pause after death/restart; shutdown leaves exact carriers/rows untouched (never pending_parent_interrupted); boot re-parks. Revoke an UNCONSUMED grant against its current identity and re-park; a CONSUMED grant takes terminal crash custody, never ordinary retry. Failed revocation keeps source/grant in a nonterminal hold (memory-only if unpersisted).
budget_pause_restore_refusal, events_budget.budget_hold_fact Pauses restore at any snapshot age/stamp validity without wake. Missing/unreadable/mismatched/terminal records or sources, root acceptance or malformed snapshot fences retain _budget_pause under _budget_pause_hold, never drop/cancel. Regrant revalidates.

Owner grant. POST /api/tasks/{id}/resume → queue_transitions.resume_budget_paused_task → budget_resume.grant_exact_budget_resume: require authoritative positive global/root headroom, readable source, clear cancel/restart/Panic/deadline/lifetime rails and an unpaused ancestor root; re-read failed custody. Root headroom is ONE fresh strict ledger read (usage_accounting.refresh_root_accounting(strict=True)): a cached snapshot after a failed read, an unreadable ledger or a degraded tree refuses typed; display readers keep a bounded-stale fallback. Pause ID/generation and resume_generation bind a single-use grant. revoke_exact_budget_resume re-parks after lost money/restart, keeps newer identities and retries unwritten revocation before regrant. model_wait.execution_elapsed_seconds subtracts quota-union and cumulative paused_duration_sec/budget_paused_sec from original started_at across continuations, owner waits included; None stays unlimited. Money/round/review wallets never reset.

Descendant selection. events_budget.hold_root_resume_descendants makes children eligible, not runnable: exact rows stay paused; zero-dispatch/fence-only siblings get root_fence_lifted_pending_selection, rebound to the current root grant on each Resume. resume_child_task (budget_resume_child, recorded selected_by) selects within lineage: root selects stored descendants, an intermediate parent only direct children; owner selection shares the path. With a live latch, budget_fence_selected binds one member to its fence generation, leaving siblings fenced; legacy root Resume selects only root; reserve_attempt admits the member selected against that exact fence. New root pauses invalidate pending child grants; assignment rechecks root grant and fence for exact grants and zero-dispatch selections (stale → unselected hold). Other roots/cancelled/completed/stopped members never revive.

Consumption. resume_paused_loop rejects spent/revoked/foreign grants and discloses drift/custody before effects; a consumption publication that fails HOLDS (resume_grant_consumption_unwritable; identities kept), a grant revoked underneath re-parks; never FAILED. Graceful Resume refreshes ledger-authorized headroom via the same strict tree read and admits ONE fitting reservation (_second_reservation_fits, budget_resume_last_fit_admitted), avoiding the early-margin re-pause. Hard exhaustion needs owner increase; full-cap admission binds. Browser/services/hidden state and crash recovery are not restored.

Monetary authority and projections

Live global limit. Agent readers share one resolver: saved settings first, re-parsed only on file change; failure (including refused benchmark pins) falls back to environment without caching failure. Workers refresh environment only at task start, so live file reads make every reservation/wallet follow current TOTAL_BUDGET; planning thresholds remain separate. Absent means shipped default; non-positive means unbounded and silences the loop global axis. Disclosed pre-existing gap: supervisor startup/reload still parse raw settings and treat absence as unlimited.

pricing.py uses bounded best-effort exact normalized-route lookup and provider catalog fields, never manual tables/prefix inheritance/fallback prices/admission allowlists. Unknown price admits while known spend fits: reserve None, settle from reported cost or later exact price, otherwise None/non-final. Only structural pre-generation evidence with zero usage permits confirmed-zero rejection/release, preventing phantom exhaustion during provider storms (_usage_response.py normalizes; adapters retain raw usage). review_wave_admission uses the same estimator and tighter global/root remainder as reserve_attempt before skill/plan/task/commit review; managed-update assisted apply reuses it before destructive merge. Unpayable slots bypass without model substitution. Unknown-priced slots disclose uncertainty without disabling priced siblings' admission.

Reservations retain applied global_limit_usd, explicit unbounded state and global_limit_source/global_limit_revision through transitions. Overrides inherit no foreign revision; legacy unknowns remain unknown. usage_compaction.py archives original rows beside the substrate, preserving sequence; commit requires identical NON-MONEY projections and decimal-identical money because six-decimal display rounding is not monetary equality. Policy abort emits usage_ledger_compaction_skipped with cause once per process/cause; name-tier refusal keeps usage_ledger_compaction_refused, making size warnings diagnosable.

usage_ledger.py owns one short cross-process lock for validation/reservation/transition/append/fsync; accounting imports it, never vice versa. Network I/O stays outside. Torn tails quarantine loudly; validated prefixes remain readable but degraded/non-final because paid work may be missing. Failed settlements retain dispatched/unresolved attempts; durable root-budget refusals clear on Resume only after proving no paid dispatch. _usage_rows_memo.py incrementally replays validated records, changing cost of reads, not meaning; append repairs a torn newline-less tail without losing earlier rows. _usage_rows.py owns pure arithmetic.

For a root task, GET /api/tasks/{id} derives cost_breakdown at read time from the same ledger (own, child, unattributed, disclosed delegated, subscription sessions, unknown/unmetered, finality, authority); it is never persisted, is not a third sum, and an unreadable or unattributable ledger omits the whole object rather than returning a confident zero. One physical reviewer send produces exactly one llm_usage row, emitted by the review substrate with that reviewer's wave and slot attribution (skill_review_usage.py is a read-only projection per (review_skill, review_wave_id), not a second ledger), and a delegated session row reports its own route provider and resolved model, never an inferred one. state.json, task results, llm_usage, /api/state and /api/cost-breakdown are compatibility projections only; startup's resumable importer (usage_legacy_import.py) records source hashes, imports only attributable usage, and represents ambiguous history explicitly without rewriting source logs or fabricating attempts.