diff --git a/docs/ARCHITECTURE.md b/docs/ARCHITECTURE.md index 4b39b7ff7..a2808013e 100644 --- a/docs/ARCHITECTURE.md +++ b/docs/ARCHITECTURE.md @@ -1117,6 +1117,10 @@ Provider death is the one forced rail that is NOT a best-effort completion: `_ha Host-enforced task acceptance is a root-owned completion coach, not the P3 commit gate. `off` disables it. Both `auto` and `required` review observable effects and typed deliverables/criteria. In `auto`, an explicit `task_acceptance_review` request also qualifies, including read-only research; queue membership alone does not. `required` retains its non-direct-root criterion. Ordinary conversation, exploration or cognitive-memory updates alone do not qualify in `auto`; no prose or tool-count classifier decides their meaning. Child reviews remain advisory evidence superseded by the root decision. +The root explicit call nominates the complete ready result. After the whole tool-result block, the host advances the same acceptance operation used by ordinary final delivery; early feedback does not seal the task. Main authors the effective criteria and can nominate material read observations through exact retained tool indices. Result bytes, those criteria and material effects form the paid subject; the full source packet and ingress generation remain separate forensic and ordering facts. Main acknowledges the source it actually received, so a status question can preserve a running review while a new criterion can request useful review of unchanged answer text. A released panel retains its exact request, resolved slot roster and existing physical-operation custody. Its settlement wakes the original Main through the mailbox; collection sends no new request, and waiting uses the existing continuation without fabricating a question or task id. Only actual final delivery seals ingress. Stop, missing custody and unfinished work retain their observed outcomes. + +A completed reviewer from an older plan wave is attached as a historical supplement through the existing locked task-result writer and exact producer CAS. It never rewrites the original actor verdict, aggregate, closure, author dispositions or current-wave pointer. The source ref, slot/operation binding and physical settlement facts remain available; a confirmed old dispatch can settle unknown historical cost once, without dispatching or charging another cycle. A terminal parent receives the supplement without starting another model turn. + Before an eligible panel is called, `supervisor/task_lifecycle.py` closes subtask admission under the queue lock and `task_status.find_child_tasks` proves the recursive subtree terminal and quiescent; revision reopens the fence, terminal or degraded completion seals it. Fence acknowledgement, subtree lookup, the timing telemetry stream, and the packet's mutation-attribution read use the canonical `budget_drive_root`; the one-shot `state/acceptance_fence_acks/` IPC sidecar is not a lifecycle authority. The reviewer packet preserves verbatim owner directives, the full task contract and criteria, canonical deliverable identity, terminal child state, verification receipts, artifact references, the host-attested lifecycle facts (review status, staleness, readiness, enablement) of every skill the task touched — visibility only, never a gate — and an explicit omissions manifest; a required component that cannot be assembled makes the affected actor `DEGRADED`, never a silently smaller prompt. The packet is SIZED against the review quorum's real windows — the same `reviewer_window`/`review_synthesis.quorum_input_token_limit` seam the triad and plan review use — resolved once per task and memoised on the acceptance context, so the packet bytes cannot drift between the binding build and the staleness rebuild. The same cached per-slot caps drive the pre-send fit check; dispatch never recalibrates them. Non-core sections then shed through a DISCLOSED ladder: the predecessor authority envelope first, then the trajectory tail and its results, artifact previews, agent-supplied evidence, and last a diff preview that keeps the durable `repo_diff_source_ref`. Each shed is a row in `omissions_manifest`. A slot whose own window cannot hold the rendered prompt is a typed `$0 not_dispatched` row while the rest of the panel reviews; a packet that still overflows after every shed stamps `__immutable_core_overflow__` naming the oversized sections and refuses the panel without spending anything. `__unresolved_partial_artifacts__` withholds the panel's packet rows only for a tool result whose exact source is genuinely `source_unavailable` (a retrieving row reads the exact source itself) — a budget shed with a durable, actor-resolvable source ref is an omission, never an unresolved partial. @@ -1543,7 +1547,7 @@ Read-only children can read/list the existing project-scoped knowledge store wit **Scheduling.** `schedule_subagent` requires `subagent_id`, a focused `objective`, and `expected_output`; the remaining public fields describe child-local context, constraints, memory, capability needs, write surface, narrower deadline, delegation budget, and acceptance claims. There is no model-visible lane/executor axis and no public `effort` override: the selected row is the complete execution choice, lineage/bounds/route/budget are host-derived, and omission never inherits the parent's acceptance claims. `subagent_runtime.select_subagent_snapshot` copies an immutable snapshot of the exact enabled row into the child task; an `api_model` row becomes an ordinary recursive API child on that exact model/effort, an `agent_session` row an ordinary recursive Ouroboros nanny on that exact external route. The cash side of a burst — each sibling launched before the first sibling's first response pays its own prefix write on cache-write-priced routes — is disclosed to the mind in the tool description as an affordance and deliberately not scheduled by the host. The old lane/executor resolver serves only old durable records; for those historical lane envelopes, `schedule_subagent` reports the requested lane only — effective facts remain on the dispatched child record — and a task carrying `configured_subagent` goes straight to `subagent_runtime`, so legacy policy cannot reinterpret an active selection. -**Waiting on children.** `wait_task`/`get_task_result` return the full single-child handoff (verification receipts red/masked-first, exact omitted count); `wait_tasks` stays batch-compact: `task_id, status, child_result_sha256, outcome_axes, result, terminal_host_notice when present, trace_summary, capability_delta when the child has something to disclose, duplicate_of`, plus the nullable cost-finality pair `accounted_upper_bound_usd`/`cost_final` and, when the child's envelope carries one, `execution_evidence` (§11.1). The retired `cost_usd` spelling is tolerated only when reading stored rows, never emitted. Both use `task_status.SETTLED_STATUSES`; a pending cancellation is the typed `cancel_state: "pending"` projection, never completion. A batch wait that expires with children still running discloses what it could not finish: the typed `wait_expired_with_live_children` block names the live child ids, the window that was requested and the clamp ceiling the schema already states. Facts only, with no advisory text in the payload and no host floor on the next window, because how long to wait is the mind's call (BIBLE P13); an id this tree never minted stays an `unknown_task_id` and is never counted as a live child, and the pinned per-child field list is unchanged. +**Waiting on children.** `wait_task`/`get_task_result` return the full single-child handoff (verification receipts red/masked-first, exact omitted count); `wait_tasks` stays batch-compact: `task_id, status, child_result_sha256, outcome_axes, result, terminal_host_notice when present, trace_summary, capability_delta when the child has something to disclose, duplicate_of`, plus the nullable cost-finality pair `accounted_upper_bound_usd`/`cost_final` and, when the child's envelope carries one, `execution_evidence` (§11.1). The retired `cost_usd` spelling is tolerated only when reading stored rows, never emitted. Both use `task_status.SETTLED_STATUSES`; a pending cancellation is the typed `cancel_state: "pending"` projection, never completion. A batch wait that expires with children still running discloses what it could not finish: the typed `wait_expired_with_live_children` block names the live child ids, the window that was requested and the clamp ceiling the schema already states. Facts only, with no advisory text in the payload and no host floor on the next window, because how long to wait is the mind's call (BIBLE P13); an id this tree never minted stays an `unknown_task_id` and is never counted as a live child, and the pinned per-child field list is unchanged. An optional `known_result_sha256` on the single-child reads, or `known_result_sha256_by_task` on batch wait, compares the existing join-ledger semantic result identity. An exact match omits repeated result/trace text and returns `result_unchanged` with an unconditional source call; current status, costs, outcome and custody facts remain. Missing or changed conditions return the normal full handoff. No persistent seen-state is inferred, and callers can always read the full result again after compaction. **What a delegated run costs.** Claudexor reports the amount in `summary.spendUsd` and its exactness in `summary.spendEstimated`; `delegate_custody.disclosed_spend` is the single reader of the pair, so the ledger row and the payload the nanny relays cannot tell different stories. Runs ask `authPreference: subscription` explicitly, because the engine default falls back to a paid key invisibly. Four cases: diff --git a/docs/CHECKLISTS.md b/docs/CHECKLISTS.md index 63ff163a8..04b23fe66 100644 --- a/docs/CHECKLISTS.md +++ b/docs/CHECKLISTS.md @@ -6,6 +6,15 @@ multi-model review prompt. When a new reviewable concern appears, add it here — not in prompts or docs. +**Application follows BIBLE P0/P3.** Review findings and failures are independent +facts in every mode. In Cyber Pro they inform Ouroboros and never prohibit an +action or require permission; the agent may configure its own subsequent work. +Configured enforcement remains recorded as selected, and an author decision does +not rewrite FAIL, pending, missing source or an unperformed effect as PASS. +Enforcement requirements below describe ordinary modes; Cyber applies the same +checks as advice. This product rule does not replace an external developer's +explicit work-order review obligations. + --- ## Advisory Pre-Review Workflow @@ -625,7 +634,7 @@ and do not return `PASS` for an item that also has a `FAIL` — the concrete `blockers` are executable by operator choice. This changes `executable_review` only; it does not rewrite the verdict, suppress findings, or change `skill_review_status` semantics. - - `pending` is never executable. A stale critic verdict does not authorize bytes; + - Outside Cyber Pro, `pending` is not executable. A stale critic verdict does not authorize bytes; under Advisory a separate current author acceptance may admit the payload after deterministic preflight. Blocking still requires fresh critic evidence. - Review state stores findings and computes the verdict at load time. Agents @@ -633,8 +642,8 @@ and do not return `PASS` for an item that also has a `FAIL` — the concrete not the raw status string, when deciding whether the skill is runnable. - A deterministic `skill_preflight` FAIL is a structural gate failure, not an LLM verdict: it persists and aggregates to `pending`, which is non-executable under - EVERY enforcement mode (advisory included) and in every readiness/execution - caller — the strongest fail-closed outcome, stronger than an overridable blocker. + ordinary enforcement mode (advisory included). Cyber keeps the failed check + and pending verdict visible while leaving the execution decision to Ouroboros. - Hard trust-boundary items are blocker findings on any FAIL regardless of reviewer-supplied severity: `manifest_schema`, `permissions_honesty`, `no_repo_mutation`, `path_confinement`, diff --git a/docs/DEVELOPMENT.md b/docs/DEVELOPMENT.md index d19e99756..d4a9aacf6 100644 --- a/docs/DEVELOPMENT.md +++ b/docs/DEVELOPMENT.md @@ -2609,9 +2609,11 @@ by "Provider Independence" above. Call-site imperatives: `ouroboros/loop_delivery.py`). A FORCED finalization resolves an armed control purely and without retry: valid keep/replace is honored, anything malformed preserves the retained candidate with a typed degraded reason, - and protocol JSON never reaches chat or the durable result. Owner - messages, tool effects, child results, and verification receipts advance - the evidence revision and require fresh delivery/acceptance binding; + and protocol JSON never reaches chat or the durable result. Main distinguishes + consumed owner source from changed requirements. Effective criteria and + material effects, including nominated read observations, define the reviewed + subject; ingress generations preserve unread-message ordering. Status text, + narration and a changed working view do not themselves buy another review; finalize task-scoped service outputs/errors before host acceptance. The control must not bypass verification, acceptance, safety, skill finalization, deadline, child handoff, the unconditional `FINAL ANSWER:` @@ -2635,11 +2637,14 @@ by "Provider Independence" above. Call-site imperatives: - Host task acceptance is root-only; eligibility uses structured facts (`outcomes.turn_has_reviewable_effects` plus a typed deliverable/criterion), never keywords (BIBLE P3/P5). The agent-callable - `task_acceptance_review` stores evidence but makes zero reviewer calls and - returns `deferred_to_host_acceptance`, `authoritative=false`. Before root - acceptance, atomically fence new descendants under the queue lock and - prove recursive subtree quiescence from the task-status SSOT; a revision - must explicitly reopen the fence, and terminal/degraded outcomes seal it. + `task_acceptance_review` records the full result nomination and returns + `deferred_to_host_acceptance`, `authoritative=false`. After the complete + tool-result block, the host advances the same operation as final delivery. + Freeze its request and resolved roster; use existing review custody and + mailbox continuation for pending work and free collection. The worker never + writes Main's live candidate or author decision. Early settlement does not + seal task ingress; actual final delivery does. Keep subtree/status facts + separate from reviewer findings and Cyber's authority under BIBLE P0. - Delivery-control JSON applies only to a final response with no tool calls. Retaining an answer leaves tools available for further work; changed evidence still requires the existing complete replacement. A requested file or diff diff --git a/ouroboros/tools/plan_review_runtime.py b/ouroboros/tools/plan_review_runtime.py index 0c76b49e8..94b3b6a79 100644 --- a/ouroboros/tools/plan_review_runtime.py +++ b/ouroboros/tools/plan_review_runtime.py @@ -803,6 +803,9 @@ def emit_plan_review_advisory_open( json.dumps(wave.get("health_epoch") or [], sort_keys=True, default=str)) if key in _ADVISORY_OPEN_SEEN: return + from ouroboros.config import get_review_enforcement + from ouroboros.tools.review_helpers import review_enforcement_blocks + row = { "type": "plan_review_advisory_open", "surface": "plan_review", @@ -813,7 +816,8 @@ def emit_plan_review_advisory_open( "paid": bool(wave.get("paid")), "cycles_paid": int(cycles_paid), "cap": cap, - "enforcement": "advisory", + "enforcement": get_review_enforcement(), + "decision_authority": "cyber_pro" if not review_enforcement_blocks("blocking") else "advisory", # Bounded per-slot typed facts: who failed, with what code, until when. "slots": [ {"slot_id": a.get("slot_id"), "ok": bool(a.get("ok")),