diff --git a/docs/CHECKLISTS.md b/docs/CHECKLISTS.md index 2420ca831..252adce99 100644 --- a/docs/CHECKLISTS.md +++ b/docs/CHECKLISTS.md @@ -886,8 +886,10 @@ authority, and prose outside the array is not parsed. paid cycle) or the slot that raised it no longer raising it in a later paid cycle. A REVIEW_REQUIRED wave whose open set empties is recorded GREEN. - **REVISE_PLAN** — blocking findings at quorum. A disposition can never close it: the agent - either changes the spec (a new fingerprint, the next paid cycle) or rejects a blocking finding - with a rationale that rides into that next cycle, where reviewers mark it resolved or still open. + either changes the spec (a new fingerprint, the next paid cycle) or answers a blocking finding + and sends the unchanged envelope with that answer, which asks again only the slot that raised + it (one paid cycle; every other slot keeps its recorded answer at $0) so it marks the finding + resolved or still open. - **DEGRADED** — no parseable quorum. Not a verdict, but the dispatched panel PAID its cycle: the wave records OPEN with each slot's typed failure state (code and reset time when known), the control line reports DEGRADED honestly, and the recorded result replays for free ONLY @@ -905,9 +907,12 @@ authority, and prose outside the array is not parsed. stays held), waiting via a one-shot `schedule_followup` and asking the owner stay open too. Paid cycles per task are bounded by the owner's `OUROBOROS_REVIEW_MAX_CYCLES` (default 2, -`unlimited` available). Replaying an identical envelope is free — identical including the -evidence the host attaches for reviewers' `need_evidence` requests, so a request received in the -last cycle makes the next envelope a new one. On cycle 2+ every reviewer sees +`unlimited` available). Replaying an identical envelope without answers is free — identical +including the evidence the host attaches for reviewers' `need_evidence` requests, so a request +received in the last cycle makes the next envelope a new one. Answers merge by `finding_id` +across calls; the identical envelope sent WITH `review_disposition` items is the addressed +re-ask (one paid cycle for the named slots only), and no host path buys a panel the agent did +not send. On cycle 2+ every reviewer sees all reviewers' findings from the previous cycle, the agent's dispositions and the spec delta, and its first duty is to adjudicate its OWN earlier findings (the rows whose finding_id starts with its panel seat): RESOLVED (the delta or the rationale answers it) and SUPERSEDED (the element it targeted diff --git a/docs/architecture/06-agent-core.md b/docs/architecture/06-agent-core.md index 0a8c6f91f..c87eb64c9 100644 --- a/docs/architecture/06-agent-core.md +++ b/docs/architecture/06-agent-core.md @@ -482,7 +482,7 @@ ONE structural fact tiers the governance pack: `constitutional` is true iff a de The ONE caller-facing strength axis is the envelope's optional `reviewer_effort`: the panel's effort for THIS plan, an ORDER that outranks each row's own pinned effort (`reviewer_slot_config.row_effort`: the order on every row but a compound Cursor/Agy route slug, whose encoded effort is the route's identity and stays; with no order, the row's explicit effort, then that compound effort, then the owner's `OUROBOROS_EFFORT_REVIEW`) and passed as an ARGUMENT of `plan_review_slots` only, so every other review surface keeps reading the untouched rows. Effort is roster identity: a different order on an OPEN review re-dispatches a paid panel within `OUROBOROS_REVIEW_MAX_CYCLES`, the same one replays free. Every wave records the order (`reviewer_effort`), the RESOLVED requested effort of each seat on its actor row (`effort`, `declared_effort` when the order set it; a compound seat that kept its own effort discloses `reviewer_effort_not_applied`), the per-seat owner baseline captured at dispatch and reused on collection (`owner_efforts`), and the typed `ordered_weaker` (seat, ordered effort, owner effort) that the verdict text, the Reviews card and the compact index all name when the panel was ordered weaker than the owner's setting. Against BIBLE P1 this is the review panel's strength for one order — the owner's setting is the default, Ouroboros may order stronger or weaker per envelope — never the core's own model, effort or horizon. Disclosed residual: a cheap panel's closed GREEN is earned authority for that envelope, and a later stronger order does not reopen it. -Declared evidence is resolved by `ouroboros/tools/plan_evidence.py` against exactly two roots — the active workspace and the system repository — with the shared sensitive-name policy on locator and target and the disclosed bounds `EVIDENCE_PER_ITEM_BYTES` / `EVIDENCE_TOTAL_BYTES`: every refused, missing, truncated, oversized, binary or URL locator becomes a typed omission row in the manifest (the host never fetches a URL); a locator may carry an exact range (`::lines=A-B`, `::bytes=A-B`, `::tail=N`, `::symbol=Name` for `.py` sources), and an oversized source is attached head-first with the cut named. A `need_evidence` locator a reviewer names is attached by the host on the next cycle through the same policy; what it cannot attach becomes a `[reviewer-requested]` omission row the panel is dispatched with, never a reason to run no reviewer. The manifest hash joins the spec hash and `constitutional` in the wave fingerprint, so changing what the reviewers can see changes the identity of the review — an attached locator makes the next envelope a new fingerprint, never an idempotent replay — while the root exploration log stays outside that identity. The evidence continuation uses a fresh full-packet dispatch only when no exact artifact reference exists, disclosed per slot as a `capability_delta`; an unreadable referenced artifact instead fails closed with `plan_review_exact_artifact_unavailable` and never mints replacement authority. +Declared evidence is resolved by `ouroboros/tools/plan_evidence.py` against exactly two roots — the active workspace and the system repository — with the shared sensitive-name policy on locator and target and the disclosed bounds `EVIDENCE_PER_ITEM_BYTES` / `EVIDENCE_TOTAL_BYTES`: every refused, missing, truncated, oversized, binary or URL locator becomes a typed omission row in the manifest (the host never fetches a URL); a locator may carry an exact range (`::lines=A-B`, `::bytes=A-B`, `::tail=N`, `::symbol=Name` for `.py` sources), and an oversized source is attached head-first with the cut named. A `need_evidence` locator a reviewer names is attached by the host on the next cycle through the same policy; what it cannot attach becomes a `[reviewer-requested]` omission row the panel is dispatched with, never a reason to run no reviewer. The manifest hash joins the spec hash and `constitutional` in the wave fingerprint, so changing what the reviewers can see changes the identity of the review — an attached locator makes the next envelope a new fingerprint, never an idempotent replay — while the root exploration log stays outside that identity. Every cycle after a paid one continues each reviewer: a packet slot sends its exact recorded transcript plus the new packet when both fit its window, otherwise a fresh packet disclosed on that slot as a `capability_delta`; a session slot resumes its sticky thread (`review_thread_continuity`), whose settled receipt is the disclosure. An unreadable referenced artifact fails closed with `plan_review_exact_artifact_unavailable` and never mints replacement authority. Own-room dialogue is automatic evidence: `ouroboros/dialogue_evidence.py` reads the complete retained room through `Memory.read_chat_generations` (rotated and consolidated history included) under the shared `project_dialogue.room_membership` predicate, plus the progress stream and addressed `owner_mailbox` entries — both speakers, options, quiz recommendations and answers, attachment names (the canonical manifest `label`), typed outcomes and sender/relay provenance; that complete capture is the SNAPSHOT, while the inline view a reviewer reads carries the conversation only. The winning quiz answer is written to canonical chat history as its structured block (after `record_answered`; an answered `quiz_closed` reply heals it), so eviction and mailbox cleanup never erase it and History shows it in the original card without a second bubble. Missing history, room binding and unstable capture remain explicit coverage gaps; acceptance freshness stays with the acceptance directives ledger, and planning consumes no separate directive section. @@ -492,7 +492,7 @@ Own-room dialogue is automatic evidence: `ouroboros/dialogue_evidence.py` reads Slots come from `reviewer_slot_config` over the `review_execution._review_route_executor` seam. Reviewers return ONLY a typed findings array — `blocking` with a `breaks` id, `note`, or `need_evidence` with a locator the host attaches (or a spec id in `breaks` as a question to the author, who answers, escalates or defers it openly in the disposition); the HOST validates membership, demotes invalid findings with disclosure, keeps failed slots in the quorum denominator, and computes the aggregate through `config.adaptive_quorum`: the honest `DEGRADED` below a parseable quorum, `REVISE_PLAN` when blocking findings reach it, `REVIEW_REQUIRED` while the wave's OPEN SET — every blocking finding not closed and every `need_evidence` without a valid disposition — is non-empty, and `GREEN` when a quorum parsed and the open set is empty; notes never change the verdict, so a note-only wave is GREEN. No reviewer emits GREEN as authority. A claim a reviewer cannot check as written is a question to the author or a note, never a blocker by construction. On cycle 2 and later the packet's convergence rule makes each reviewer adjudicate its OWN earlier findings first (resolved or superseded are not repeated; still open is re-emitted with the residual, and only while the goal is unchanged, which the host states as `Goal changed since cycle n` from the spec delta), the rubric carries a subtraction voice (what the spec could drop without losing the goal, as a note), and every slot's send ends with its own panel seat id after the cache-stable prefix (`plan_dialogue.dialogue_slot_inputs`). Premise criticism and simpler alternatives are optional brainstorming notes whose adoption is Ouroboros's; no mandatory competing plan, and repetition alone cannot promote advice to a blocker. -`plan_review_state` v2 inside the root task result is the bounded durable index: each wave records a hash-bound `task_source` reference to the full operative spec, the hashes, `constitutional`, validated findings, aggregate, dispositions and paid state (full waves, specs and manifests: `plan_review_artifacts.py`; review evidence stays redacted). State reads share the task-result writer's lock without rewriting JSON; source resolution starts after lock release. Current-authority readers restore the full spec before comparing plans or binding its `acceptance_claims` (`contracts/task_contract.effective_acceptance_claims`); compaction and child promotion retain both references, and an unavailable recorded source is a typed infrastructure failure, never empty claims or a new unpaid cycle. A v1 record is read-only (`legacy_v1_projection`; an OPEN v1 wave projects `legacy_open_requires_resubmission`, never auto-closed). A corrected full goal/plan/spec can be selected by an explicit `review_disposition.author_action` with author disposition and the referenced review fingerprint. `current_attempt.author_subject` holds its exact source and separate critic reference; `plan_review_artifacts.current_author_plan` restores it without a new wave, GREEN or paid cycle. Advisory finish can select its claims as `acceptance_claims_source=author_plan`, after ingress claims and through the existing receipt-support checks. Blocking can save and stop but cannot use the plan as approved; `closed_plan_review_wave` remains literal critic-closed authority. The gate and terminal disclosure follow the author's explicit `review_fingerprint` for historical critic outcome and pending custody as EVIDENCE INTO A GAP — a revised plan with no wave of its own stops reading `unavailable` — without granting that wave's closure to the revised plan: the enrichment fills only an ABSENT outcome and never moves the attempt's own lifecycle, because overwriting the status with `open` took the release back off an exhausted cap and off a degraded rail (an author permitted only to finish honestly was pushed back into a cycle it cannot buy), forcing `closed=False` contradicted `closed_plan_review_wave`, which still bound that wave, and replacing a current wave's own aggregate would let a historical verdict stand in for the one this plan earned. The projection therefore carries `historical_critic`, and both of `owner_hurry.py`'s renderers read it: the terminal disclosure labels that outcome as the earlier plan's instead of showing it as this plan's own, and the blocking reminder answers it with its own branch, because its outcome-keyed advice would otherwise send the agent to dispose a wave that cannot approve the revised bytes. A rail that forced finalization keeps its own sentence with the author clause appended, since returning the author decision alone hid `the cap is spent; the task ends blocked`; an author STOP carries `allow=True` under every enforcement, so it is excluded from the advisory-continuation sentence rather than described as work that proceeded — and because the projection's stop REPLACES the status, a rail that fired under a stop is not separately named, the stop being the fact that ended the task. The open-wave status reads `cycles_exhausted` from the attempt and the wave, because an author finish at a spent cap stamps only the attempt (`_apply_author_subject`, the one writer with no in-flight guard) and the gate otherwise held a task that may only finish honestly; that read is guarded by `custody_pending`, since a paid slot still working could yet close the wave and releasing then would trade an honest hold for a premature blocked terminal. Missing criticism stays unavailable, and author finish/stop stays separate. The author path publishes the critic wave's REAL `(aggregate, closed)` pair on its tool result — a revised plan gets no host-invented verdict — with `plan_review_historical_critic` set when the selected plan is not the reviewed one, and its text states that the earlier plan was that aggregate while this plan has no verdict of its own; `web/modules/review_presentation.js::planReviewGroupFromTaskDetail` labels such a group as the earlier plan's review and names the selected plan unreviewed. Disclosed gap: the same projector reads `rail_degraded` from the attempt but `cycles_exhausted` only from the wave, so an author finish at a spent cap on its own wave releases the gate and reads blocked in the terminal while the card still shows the wave's aggregate — a presentation gap needing its own change with browser verification, not a backend fact. Old-wave collection never replaces the selected author source. The first exact wave also stores the request policy and per-slot prepared inputs/fit sizes, which collection reuses unchanged — `native_mandatory_read_chars`, the review contract hash and the physical-operation binding never move with live exploration or owner clarification — and a historical wave without them is a typed source-unavailable result, never guessed values. +`plan_review_state` v2 inside the root task result is the bounded durable index: each wave records a hash-bound `task_source` reference to the full operative spec, the hashes, `constitutional`, validated findings, aggregate, dispositions and paid state (full waves, specs and manifests: `plan_review_artifacts.py`; review evidence stays redacted). State reads share the task-result writer's lock without rewriting JSON; source resolution starts after lock release. Current-authority readers restore the full spec before comparing plans or binding its `acceptance_claims` (`contracts/task_contract.effective_acceptance_claims`); compaction and child promotion retain both references, and an unavailable recorded source is a typed infrastructure failure, never empty claims or a new unpaid cycle. A v1 record is read-only (`legacy_v1_projection`; an OPEN v1 wave projects `legacy_open_requires_resubmission`, never auto-closed). A corrected full goal/plan/spec can be selected by an explicit `review_disposition.author_action` with author disposition and the referenced review fingerprint; the same call's `items` are recorded first. `current_attempt.author_subject` holds its exact source and separate critic reference; `plan_review_artifacts.current_author_plan` restores it without a new wave, GREEN or paid cycle. Advisory finish can select its claims as `acceptance_claims_source=author_plan`, after ingress claims and through the existing receipt-support checks. Blocking can save and stop but cannot use the plan as approved; `closed_plan_review_wave` remains literal critic-closed authority. The gate and terminal disclosure follow the author's explicit `review_fingerprint` for historical critic outcome and pending custody as EVIDENCE INTO A GAP — a revised plan with no wave of its own stops reading `unavailable` — without granting that wave's closure to the revised plan: the enrichment fills only an ABSENT outcome and never moves the attempt's own lifecycle, because overwriting the status with `open` took the release back off an exhausted cap and off a degraded rail (an author permitted only to finish honestly was pushed back into a cycle it cannot buy), forcing `closed=False` contradicted `closed_plan_review_wave`, which still bound that wave, and replacing a current wave's own aggregate would let a historical verdict stand in for the one this plan earned. The projection therefore carries `historical_critic`, and both of `owner_hurry.py`'s renderers read it: the terminal disclosure labels that outcome as the earlier plan's instead of showing it as this plan's own, and the blocking reminder answers it with its own branch, because its outcome-keyed advice would otherwise send the agent to dispose a wave that cannot approve the revised bytes. A rail that forced finalization keeps its own sentence with the author clause appended, since returning the author decision alone hid `the cap is spent; the task ends blocked`; an author STOP carries `allow=True` under every enforcement, so it is excluded from the advisory-continuation sentence rather than described as work that proceeded — and because the projection's stop REPLACES the status, a rail that fired under a stop is not separately named, the stop being the fact that ended the task. The open-wave status reads `cycles_exhausted` from the attempt and the wave, because an author finish at a spent cap stamps only the attempt (`_apply_author_subject`, the one writer with no in-flight guard) and the gate otherwise held a task that may only finish honestly; that read is guarded by `custody_pending`, since a paid slot still working could yet close the wave and releasing then would trade an honest hold for a premature blocked terminal. Missing criticism stays unavailable, and author finish/stop stays separate. The author path publishes the critic wave's REAL `(aggregate, closed)` pair on its tool result — a revised plan gets no host-invented verdict — with `plan_review_historical_critic` set when the selected plan is not the reviewed one, and its text states that the earlier plan was that aggregate while this plan has no verdict of its own; `web/modules/review_presentation.js::planReviewGroupFromTaskDetail` labels such a group as the earlier plan's review and names the selected plan unreviewed. Disclosed gap: the same projector reads `rail_degraded` from the attempt but `cycles_exhausted` only from the wave, so an author finish at a spent cap on its own wave releases the gate and reads blocked in the terminal while the card still shows the wave's aggregate — a presentation gap needing its own change with browser verification, not a backend fact. Old-wave collection never replaces the selected author source. The first exact wave also stores the request policy and per-slot prepared inputs/fit sizes, which collection reuses unchanged — `native_mandatory_read_chars`, the review contract hash and the physical-operation binding never move with live exploration or owner clarification — and a historical wave without them is a typed source-unavailable result, never guessed values. A fresh dispatch returns control at the dispatch barrier (`ReviewRequest.drain_deadline`, drain window 0): the wave is recorded open with typed `pending_dispatch` rows for the workers still running, whose custody stays with the recorded wave in the process (`review_custody`). A paid actor still physically in flight keeps the wave open as `DEGRADED` with `review_late_result_pending` even when the settled rows meet the arithmetic quorum, so a late blocking result cannot arrive after a false GREEN. When the last released slot settles, the settlement thread writes ONE system frame into the task's mailbox (`owner_mailbox.write_task_message`, provenance `system`) so the mind wakes exactly as for a child result; it never closes or aggregates the wave — collection is the sole wave writer (`plan_review_collect`). Task acceptance's twin announcements (its own quorum, completion) carry each reviewer's own verdict; the aggregate still comes only from collection. @@ -500,9 +500,9 @@ Collection is the existing `review_disposition` mode addressed at the recorded w An answer that has not arrived is a gap, never a failure or a verdict. The STORED wave is the fail-closed floor: an unanswered slot stays an `ok=False` row no quorum counts — on a same-spec cycle (equal `spec_hash`, whatever the roster, order or effort) it keeps the seat's still-open findings from the predecessor listed on the wave once its absence is terminal (a failed, unparseable or never-sent seat; a seat still awaiting carries nothing), stamped `carried_absent_answer` (`plan_spec.plan_standing_findings`, walked across pending predecessors by `plan_review_artifacts.standing_findings_lineage`), so silence never manufactures GREEN and casts no blocking-slot vote — and the open `DEGRADED` + `custody_pending` pair holds every gate with no fifth aggregate for older builds to refuse. Renderers read `plan_review_runtime.plan_wave_slot_census` before wording a slot: `awaiting` is `operation_state=pending_dispatch` only; `unresolved` is `in_flight`/`custody_lost`, named as unresolved on the line and by raw state in Reviews, never called waiting (`late_result_pending` is true for both); `uncollected` has a settled supplement not collected yet; `not_dispatched` is a $0 refusal, not a failure. Verdict and failure wording and the task's `degraded` stamp belong to collected slots only (the stamp keys on an open review, and a verdict-bearing wave is open only while its open set is non-empty): a finish over an only-awaited wave or panel keeps its disclosure and records `awaiting` (`execution.plan_review`, review axis; `owner_hurry.plan_wave_only_awaited`, `_outcome_receipts.review_runs_only_awaited`). `owner_hurry.force_plan_decision` makes ONE free collection of the current pending wave in every enforcement mode, hurry too, then projects the returned state (`plan_review_collect.collect_before_gate`). Advisory and hurry keep their local release; collection updates disclosed facts, never permission to proceed. Context health reads the canonical wave, never collecting: `PLAN REVIEW WAVE OPEN` names its fingerprint, pending-slot count and snapshot timestamp (not a physical start, not a liveness lease) and disappears once collection clears custody. -Closure follows the finding class over the wave's open set (`plan_spec.closure_after_disposition`, the ONE closure table, at initial synthesis and at later dispositions): GREEN — no blocking finding and no `need_evidence` without a disposition — closes immediately in either enforcement mode, so a note-only wave is GREEN; outstanding `need_evidence` closes through a disposition-only `plan_task` call (no model call, no cost: accept = answered, reject, defer = deferred openly; an answer reaches the reviewers on the next paid cycle, a revised envelope supersedes the wave with its open requests); a below-quorum blocking finding closes, per finding, under advisory enforcement by a reject with its rationale (accept or defer keep it open) and under blocking enforcement stays open until a changed spec is reviewed or the reviewer retires it in a later paid cycle (closure reads the CONFIGURED enforcement: an armed owner hurry projects advisory for the finalization gate only, never for closure); a REVIEW_REQUIRED wave whose open set empties is written GREEN on every write path with the closure note `closed_by_disposition`, while the four control validators keep accepting older closed REVIEW_REQUIRED rows; REVISE_PLAN is never closed by disposition — a subsequent paid delta review may evaluate a changed spec or a justified rejection when another cycle is available. A closed note-only wave still accepts voluntary `review_disposition` annotations through the same writer, spec/verdict and paid-cycle count fixed and the immutable predecessor retained: optional reasoning history, not a new plan gate, never another panel. +Closure follows the finding class over the wave's open set (`plan_spec.closure_after_disposition`, the ONE closure table, at initial synthesis and at later dispositions): GREEN — no blocking finding and no `need_evidence` without a disposition — closes immediately in either enforcement mode, so a note-only wave is GREEN; outstanding `need_evidence` closes through a disposition-only `plan_task` call (no model call, no cost: accept = answered, reject, defer = deferred openly; answers merge by `finding_id`, a later answer superseding only its own id; an envelope without items supersedes the wave, its answers kept); a below-quorum blocking finding closes, per finding, under advisory enforcement by a reject with its rationale (accept or defer keep it open) and under blocking enforcement stays open until a changed spec is reviewed or the reviewer retires it in a later paid cycle (closure reads the CONFIGURED enforcement: an armed owner hurry projects advisory for the finalization gate only, never for closure); a REVIEW_REQUIRED wave whose open set empties is written GREEN on every write path with the closure note `closed_by_disposition`, while the four control validators keep accepting older closed REVIEW_REQUIRED rows; REVISE_PLAN is never closed by disposition — a changed spec, or an addressed answer the raising slot re-judges, needs another paid cycle. A closed note-only wave still accepts voluntary `review_disposition` annotations through the same writer, spec/verdict and paid-cycle count fixed and the immutable predecessor retained: optional reasoning history, not a new plan gate, never another panel. -Paid cycles are bounded by the shared `OUROBOROS_REVIEW_MAX_CYCLES`; `cycles_paid` and the typed `cycles_exhausted` state count PAID cycles only. A wave is paid iff at least one reviewer slot was physically dispatched, and the proof is a settled row: a slot released at the barrier is $0 until its row proves the send, so the cycle is paid at collection, never at the barrier, and a nothing-dispatched wave of typed $0 skip rows stays unpaid and never replaces a paid predecessor. A barrier-dispatched panel is nevertheless committed money: while its wave is custody-pending and the cap has no room for another panel, a revised envelope first collects EVERY other custody-pending wave at $0 (`plan_review_collect.collect_before_supersede`) and, if the cap still has no room, is HELD with the typed `PLAN_REVIEW_IN_FLIGHT` refusal (`plan_review_collect.in_flight_hold`) so the current wave stays collectible (the per-case cap arithmetic: `plan_review_collect.py` docstrings). An identical envelope replays the recorded wave free, with one exception: an open wave whose blocking findings all carry valid reject dispositions may dispatch exactly one subsequent paid delta panel when another cycle is available. +Paid cycles are bounded by the shared `OUROBOROS_REVIEW_MAX_CYCLES`; `cycles_paid` and the typed `cycles_exhausted` state count PAID cycles only. A wave is paid iff at least one reviewer slot was physically dispatched, and the proof is a settled row: a slot released at the barrier is $0 until its row proves the send, so the cycle is paid at collection, never at the barrier, and a nothing-dispatched wave of typed $0 skip rows stays unpaid and never replaces a paid predecessor. A barrier-dispatched panel is nevertheless committed money: while its wave is custody-pending and the cap has no room for another panel, a revised envelope first collects EVERY other custody-pending wave at $0 (`plan_review_collect.collect_before_supersede`) and, if the cap still has no room, is HELD with the typed `PLAN_REVIEW_IN_FLIGHT` refusal (`plan_review_collect.in_flight_hold`) so the current wave stays collectible (the per-case cap arithmetic: `plan_review_collect.py` docstrings). An identical envelope replays the recorded wave free, and no host path buys a panel the mind did not send. The identical envelope carrying `review_disposition` items is an addressed answer (`plan_review_artifacts.addressed_slots`): the items are recorded first; when that wave was open before they landed, settled, quorum-bearing and from the current roster, only the slots the items name are dispatched (one paid cycle, same fingerprint, `addressed` lineage and `previous_wave_artifact`) while every other slot keeps its recorded row at $0 (`not_dispatched`, `replayed_from`) with its dispositions, and an addressed slot without a parseable answer keeps its findings. Any other case follows the ordinary rule with a typed `answers_not_addressed` note; a changed envelope with items is an ordinary wave. An escalated question stays open until the mind records its disposition (defer while the quiz is open, accept with the owner's decision); a quiz answer never closes a finding. Before fan-out the engine captures one panel health snapshot (`subagents.route_health`, route-level evidence): a slot with positive structural evidence of a spent lane becomes a $0 typed skip row that stays in the denominator; unknown health dispatches (fail-open). The wave records the health epoch and reviewer-roster fingerprint, and a recorded DEGRADED wave replays free only under an identical envelope, matching epoch and unchanged roster — otherwise another paid panel needs remaining cycle capacity (replay/epoch casuistry: `plan_review.py` docstrings). When the wave's own typed rows prove the quorum structurally unreachable, it carries `quorum_unreachable` plus the earliest recorded reset, and under blocking enforcement the finalization gate RELEASES while the review stays open: the agent may finalize `blocked_with_evidence` (reason `plan_review_quorum_unreachable`), wait through a one-shot `schedule_followup`, or ask the owner — the host adds facts only, never an answer template. diff --git a/docs/development/06-rules-by-change-class.md b/docs/development/06-rules-by-change-class.md index 469eda6d4..c1b900a5b 100644 --- a/docs/development/06-rules-by-change-class.md +++ b/docs/development/06-rules-by-change-class.md @@ -1092,6 +1092,14 @@ and what enforces each. author completion separately, never rewriting criticism or hiding independent failed effects, unaccepted review or unfinished stops (DEVELOPMENT §11; ARCHITECTURE §3). No task scope review or commit-gate reuse. +- Plan-review answers are durable statements merged by `finding_id`: a later answer + supersedes only its own id, and two entries for one id in ONE call stay contradictory and + open. Answering and re-asking are one optional call: an envelope with items records them + first, then reviews; the unchanged envelope asks only the named slots, and every other + slot's recorded row is replayed at $0 (`not_dispatched` + `replayed_from`, never a send), + keyed by fingerprint equality. An addressed slot without a parseable answer keeps its + findings; a quiz answer never closes a finding; no host path buys a panel the mind did + not send (`tests/test_plan_review_answer_channel.py`). - Keep reviewer DIALOGUE evidence: typed `disposition_kind`/`obligation_id` identifies obligations; disclose unknown re-raise ids as `new`. Reopen rows with arguments intact. A terminal critic vote cannot deny author reaction or choose its stop. Blocking may save corrections and stop; advancement needs fresh diff --git a/ouroboros/owner_hurry.py b/ouroboros/owner_hurry.py index 9514d0c6b..7cb184692 100644 --- a/ouroboros/owner_hurry.py +++ b/ouroboros/owner_hurry.py @@ -731,17 +731,19 @@ def plan_review_reminder(decision: Dict[str, Any]) -> str: ) if outcome == "REVIEW_REQUIRED": return ( - f"{tag} Blocking plan review remains REVIEW_REQUIRED. Re-call plan_task with a " - "complete review_disposition as the only field, naming the latest fingerprint, " - "then continue; do not rerun reviewers." + f"{tag} Blocking plan review remains REVIEW_REQUIRED. Open need_evidence requests close " + "with a $0 review_disposition naming the latest fingerprint; a blocking finding below quorum " + "stays open until its slot no longer raises it in a later paid cycle (the unchanged envelope " + "with your items asks only that slot) or a changed spec is reviewed without it. " + "Implementation stays held while the review is open." ) if outcome == "REVISE_PLAN": return ( - f"{tag} Blocking plan review requires a revised spec. Change the spec — it carries " - "affected_paths, the files the work will change ([] when none) — and call " - "plan_task again (or reject the blocking findings with a rationale via " - "review_disposition). Continue analysis and non-mutating preparation, but do not " - "begin the work before the review closes or a real task-wide rail fires." + f"{tag} Blocking plan review is REVISE_PLAN. A changed spec — with affected_paths, the files " + "the work will change ([] when none) — is a new envelope every slot reviews; the unchanged " + "envelope with review_disposition items asks only the slots those items name. Analysis and " + "non-mutating preparation remain open; the work starts after the review closes or a real " + "task-wide rail fires." ) return ( f"{tag} Call plan_task with a concrete goal, plan and spec, whose affected_paths lists " diff --git a/ouroboros/tools/plan_render.py b/ouroboros/tools/plan_render.py index bbaae5ccb..038b43980 100644 --- a/ouroboros/tools/plan_render.py +++ b/ouroboros/tools/plan_render.py @@ -164,6 +164,47 @@ def _quorum_unreachable_fact(wave: dict) -> str: ) +def _open_questions_line(wave: dict, enforcement: str) -> str: + """The open ``need_evidence`` ids, split into questions to the author (no locator) and + document requests (a locator), read from the ONE closure table; '' when none is open.""" + from ouroboros.tools import plan_spec + + findings = [f for f in wave.get("findings") or [] if isinstance(f, dict)] + open_ids = set(plan_spec.closure_after_disposition( + str(wave.get("aggregate") or ""), findings, wave.get("dispositions") or [], enforcement)["open_ids"]) + asks = [f for f in findings if f.get("class") == "need_evidence" + and str(f.get("finding_id") or f.get("id") or "") in open_ids] + questions = [f for f in asks if not str(f.get("locator") or "").strip()] + documents = [f for f in asks if str(f.get("locator") or "").strip()] + + def named(rows: list, field: str) -> str: + shown = [f"{f.get('finding_id') or f.get('id')} ({f.get(field) or '?'})" for f in rows[:8]] + return ", ".join(shown) + (f", +{len(rows) - 8} more" if len(rows) > 8 else "") + + return ((f"Open questions to you: {named(questions, 'breaks')}. " if questions else "") + + (f"Open document requests: {named(documents, 'locator')}. " if documents else "")) + + +def _answer_route(fp: str, *, at_cap: bool, cycles_paid: int, cap: Optional[int]) -> str: + """The route an answer takes, as facts: recorded at $0 and merged by id; beside the + unchanged envelope it asks again only the seats it names (one paid cycle, the rest kept at + $0); beside a changed one every seat reviews with the answers in view. Whether to buy that + cycle is the mind's; at the cap the answer stays recorded evidence.""" + recorded = (f"An answer is recorded at $0 by plan_task(review_disposition={{review_fingerprint: '{fp}', " + "items: [...]}) (accept = answered, the rationale is the answer; reject; defer = deferred " + "openly) and merged by finding_id") + if at_cap: + return recorded + ("; the cycle cap is reached, so no paid cycle remains to show it to a reviewer: " + "it stays recorded evidence. ") + return recorded + ( + ". The same call beside the unchanged goal/plan/spec asks only the slots its items name — one paid " + f"cycle {cycles_paid + 1}{'' if cap is None else f' of {cap}'}; every other slot keeps its recorded " + "answer at $0 — and a slot that no longer raises its finding retires it; beside a changed " + "goal/plan/spec every slot reviews the new envelope with the answers in view. Whether to buy that " + "cycle is yours. An envelope sent without items supersedes this wave (its recorded answers stay); " + "the identical one replays free. ") + + def _next_step(wave: dict, *, enforcement: str, cap: Optional[int], cycles_paid: int) -> str: aggregate = str(wave.get("aggregate") or "") fp = str(wave.get("request_fingerprint") or "") @@ -187,8 +228,8 @@ def _next_step(wave: dict, *, enforcement: str, cap: Optional[int], cycles_paid: author_note + "Cyber Pro: Ouroboros decides whether and how to continue. " "The recorded verdict, open findings and any unresolved physical reviewers remain " "independent facts; continuation does not close the wave or create a PASS. " - f"The existing $0 plan_task(review_disposition={{review_fingerprint: '{fp}', items: [...]}}) " - "can collect results or record a disposition without a new panel." + + ("" if bool(wave.get("closed")) else + _open_questions_line(wave, enforcement) + _answer_route(fp, at_cap=at_cap, cycles_paid=cycles_paid, cap=cap)) ) if bool(wave.get("closed")): if plan_review_notes_are_annotatable(wave): @@ -262,47 +303,31 @@ def _next_step(wave: dict, *, enforcement: str, cap: Optional[int], cycles_paid: text += _quorum_unreachable_fact(wave) elif aggregate == "REVIEW_REQUIRED": blocking = [f for f in wave.get("findings") or [] if f.get("class") == "blocking"] - text = author_note + ( - "Notes are optional. Open need_evidence requests (an evidence locator or " - "a question addressed to you by spec id) close with ONE $0 call: " - f"plan_task(review_disposition={{review_fingerprint: '{fp}', items: [...]}}) — accept = " - "answered (your rationale is the answer; " - + ("the cap leaves no further paid cycle to deliver it to reviewers), " if at_cap else - "it reaches reviewers on the next paid cycle), ") - + "reject, or defer = deferred openly; no reviewer call, no cycle. A revised envelope " - "supersedes this wave and its open requests can no longer be dispositioned. " - ) - if blocking: + text = (author_note + "Notes are optional and need no answer. " + _open_questions_line(wave, enforcement) + + _answer_route(fp, at_cap=at_cap, cycles_paid=cycles_paid, cap=cap)) + if blocking: # the closure table's per-finding rule, stated as a fact ids = ", ".join(str(f.get("finding_id") or f.get("id")) for f in blocking[:4]) - if enforcement == "advisory": # the closure table's per-finding rule, stated as a fact - text += ( - f"NOTE: {len(blocking)} BLOCKING finding(s) below quorum ({ids}): a reject with its " - "rationale closes each one; accept or defer keeps it open until a changed spec is reviewed. " - + ("" if not at_cap else "The cycle cap is reached; no further paid panel is available. ") - ) - else: - text += ( - f"NOTE: {len(blocking)} BLOCKING finding(s) below quorum ({ids}) stay OPEN whatever " - "you disposition: a changed spec, or a justified rejection judged in another paid " - "cycle, closes them. " - + ("" if not at_cap else "The cycle cap is reached; no further paid panel is available. ") - ) + text += ( + f"NOTE: {len(blocking)} BLOCKING finding(s) below quorum ({ids}): a reject with its " + "rationale closes each one; accept or defer keeps it open until a changed spec is reviewed. " + if enforcement == "advisory" else + f"NOTE: {len(blocking)} BLOCKING finding(s) below quorum ({ids}) stay open after any " + "disposition; they leave the review when the slot that raised them no longer raises it in a " + "later paid cycle, or when a changed spec is reviewed without them. " + ) + ("The cycle cap is reached; no further paid panel is available. " if at_cap else "") else: - text = author_note + ( - "Blocking findings: accept ⇒ change the spec and re-call plan_task (new fingerprint, " - f"{'the cap is reached — no further paid cycle' if at_cap else 'next paid cycle ' + str(cycles_paid + 1) + ('' if cap is None else f' of {cap}')}); " - "reject ⇒ record reject + rationale via review_disposition naming this fingerprint — " - + ("it is recorded as evidence; the cap leaves no further paid delta cycle to judge it. " - if at_cap else - "it rides into the next paid delta cycle where reviewers mark it resolved or still-open. ") - + "A disposition never closes REVISE_PLAN. " - ) + blocking = [f for f in wave.get("findings") or [] if f.get("class") == "blocking"] + ids = ", ".join(str(f.get("finding_id") or f.get("id")) for f in blocking[:6]) + text = (author_note + f"Blocking findings at quorum{f' ({ids})' if ids else ''}. A disposition records " + "an answer; a disposition never closes REVISE_PLAN. " + _open_questions_line(wave, enforcement) + + _answer_route(fp, at_cap=at_cap, cycles_paid=cycles_paid, cap=cap)) if enforcement == "blocking": text += ( _BLOCKING_HOLDS - + (" — the cycle cap is reached: exits are owner unstick (Swarm/hurry), a revised spec " - "once the owner raises OUROBOROS_REVIEW_MAX_CYCLES, or finalizing with " - "outcome_tier=blocked_with_evidence." if at_cap else ".") + + (" — the cycle cap is reached: no paid panel remains; finalization is RELEASED, and finalizing " + "now records outcome_tier=blocked_with_evidence with the review left OPEN and implementation " + "still held; a revised spec buys a paid cycle once the owner raises OUROBOROS_REVIEW_MAX_CYCLES " + "(owner authority, never reviewer approval)." if at_cap else ".") ) if wave.get("quorum_unreachable") and not bool(wave.get("closed")): # B2b facts, never imperatives: the honest exits that exist alongside diff --git a/ouroboros/tools/plan_review.py b/ouroboros/tools/plan_review.py index aa8236b41..0014aa028 100644 --- a/ouroboros/tools/plan_review.py +++ b/ouroboros/tools/plan_review.py @@ -230,15 +230,21 @@ _DISPOSITION_SCHEMA = { "type": "object", "additionalProperties": False, "description": ( - "Answer the findings of the wave (ordinary dispositions use only this field); " - "explicit author_action=finish|stop with author_disposition may also select a full current goal/plan/spec without a new reviewer. The wave is " - "named by review_fingerprint. While that wave is still open with reviewer slots in " - "flight, this call first COLLECTS what has settled at $0 without waiting (items may be " - "[]); to wait longer, re-submit the same envelope. Notes never hold the wave; need_evidence " - "closes at $0; under advisory a reject with rationale also closes a below-quorum blocking " - "finding; otherwise a blocking finding stays open. A subsequent paid delta review may consider a changed spec or " - "justified rejection when another paid cycle is available. Recording a disposition " - "consumes no cycle and never closes REVISE_PLAN." + "Answer the findings of the wave named by review_fingerprint. Alone, this field records the " + "answers at $0, merged by finding_id (a later answer to the same id supersedes the earlier one; " + "every other answer stays). Beside goal/plan/spec the answers are recorded first and the envelope " + "is then reviewed: the unchanged envelope asks again ONLY the slots whose findings the items name, " + "in one paid cycle, and every other slot keeps its recorded answer at $0; a changed envelope is " + "reviewed by every slot with the answers in view. With author_action=finish|stop the items are " + "recorded and the current goal/plan/spec may be selected without a new reviewer. While the wave " + "still has reviewer slots in flight this call first COLLECTS what has settled at $0 without " + "waiting (items may be []); to wait longer, re-submit the same envelope. A note or need_evidence " + "finding closes at $0 by its answer; a blocking finding stays open until the slot that raised it " + "no longer raises it or a changed spec is reviewed without it (under advisory a reject with its " + "rationale also closes a below-quorum blocking finding). Recording answers consumes no cycle and " + "never closes REVISE_PLAN. A question you escalate to the owner (escalate) stays open until you " + "record its answer here: defer (rationale names the quiz) while the quiz is open, accept with the " + "owner's decision in the rationale once decided. A quiz answer never closes a finding by itself." ), "properties": { "review_fingerprint": {"type": "string"}, @@ -247,12 +253,16 @@ _DISPOSITION_SCHEMA = { "description": "Default none means no author action, even if author_disposition is filled; collect or answer findings only. finish/stop explicitly select the current plan. stop permits unfinished finalization only; finish never overrides Blocking review."}, "items": { "type": "array", + "description": ("Answers by finding_id. Every item names its reviewer slot; beside the unchanged " + "envelope those slots are asked again."), "items": { "type": "object", "additionalProperties": False, "properties": { "finding_id": {"type": "string"}, "decision": {"type": "string", "enum": list(plan_spec.DISPOSITION_DECISIONS)}, - "rationale": {"type": "string"}, + "rationale": {"type": "string", "description": ( + "Your answer. For an escalated question: the quiz id while it is open (defer), " + "the owner's decision once answered (accept).")}, }, "required": ["finding_id", "decision", "rationale"], }, @@ -276,11 +286,14 @@ def get_tools(): "no open need_evidence (notes never change the verdict); need_evidence closes by " "review_disposition at no cost; under advisory enforcement a reject with its rationale " "also closes a blocking finding below quorum; REVISE_PLAN needs " - "a changed spec or justified rejection judged by a subsequent paid delta review " - "when another paid cycle is available. Cycles are bounded by the owner's Max review cycles; an unchanged " + "a changed spec, or your answer re-judged by the slot that raised the finding. " + "Cycles are bounded by the owner's Max review cycles; an unchanged " "envelope replays the recorded result for free (a locator a reviewer asked for " "with need_evidence is attached by the host next time and makes the envelope " - "new; on an OPEN review a different reviewer_effort re-dispatches the panel, a CLOSED review stands for its envelope). Under blocking enforcement an " + "new; on an OPEN review a different reviewer_effort re-dispatches the panel, a CLOSED review stands for its envelope); " + "the unchanged envelope together with review_disposition items records the answers and asks again " + "only the slots those items name (one paid cycle; the others keep their answers at $0), while a " + "changed envelope with items is reviewed by every slot. Under blocking enforcement an " "open review holds implementation. An explicit review_disposition.author_action=stop " "permits unfinished finalization only. Advisory author_action=finish may select a corrected " "goal+plan+spec in the same call without another panel, citing the earlier review_fingerprint " @@ -298,7 +311,7 @@ def get_tools(): "spec": _SPEC_SCHEMA, "reviewer_effort": {**_REVIEWER_EFFORT_SCHEMA, "enum": ["default", *_REVIEWER_EFFORT_SCHEMA["enum"]], "default": "default", - "description": "default uses the configured effort without an override; ignored when collecting or answering a recorded wave. " + _REVIEWER_EFFORT_SCHEMA["description"]}, + "description": "default uses the configured effort without an override; ignored when only collecting a recorded wave. " + _REVIEWER_EFFORT_SCHEMA["description"]}, "review_disposition": _DISPOSITION_SCHEMA, }, # Review, disposition-only, or explicit author selection with a current envelope. @@ -1037,11 +1050,12 @@ def _cycles_exhausted( ) if review_enforcement_blocks(enforcement): head += ( - "Blocking enforcement: the plan review stays OPEN, so implementation stays held — but " + "Blocking enforcement: the plan review stays OPEN, so implementation stays held — and " "finalization is RELEASED so the task can end honestly instead of waiting for a panel it " - "can no longer buy (owner decision D27). Your exits are an owner unstick (Swarm/hurry), a " - "revised spec once the owner raises OUROBOROS_REVIEW_MAX_CYCLES, or finalizing now with " - "outcome_tier=blocked_with_evidence. Do not start the work under an open blocking review." + "can no longer buy: finalizing now records outcome_tier=blocked_with_evidence with the " + "review left open. Recorded answers stay recorded evidence. Available authority beyond " + "yours: the owner may raise OUROBOROS_REVIEW_MAX_CYCLES (a revised spec then buys a paid " + "cycle) or perform an owner unstick — owner authority, never reviewer approval." ) elif not review_enforcement_blocks("blocking"): head += "Cyber Pro permits proceeding by Ouroboros's judgment; the open review and spent cycles remain recorded facts." diff --git a/ouroboros/tools/plan_review_runtime.py b/ouroboros/tools/plan_review_runtime.py index b1459aed1..f9e632a95 100644 --- a/ouroboros/tools/plan_review_runtime.py +++ b/ouroboros/tools/plan_review_runtime.py @@ -315,8 +315,9 @@ def plan_deadline_skip(ctx: ToolContext, *, emit: bool = False) -> str: f"review window of {int(scaled)}s (< {int(minimum)}s useful floor)." ) return ( - f"PLAN_TASK_SKIPPED_DEADLINE: {cause} Proceed with your own best plan " - "directly; do not re-call plan_task under this deadline." + f"PLAN_TASK_SKIPPED_DEADLINE: {cause} No reviewer was dispatched and no plan review is open " + f"for this envelope; remaining time {max(0, int(remaining))}s. The same call under this deadline " + "returns this rail again: the deadline is the owner's task bound, not a reviewer verdict." ) diff --git a/tests/test_plan_review_answer_texts.py b/tests/test_plan_review_answer_texts.py new file mode 100644 index 000000000..725d80695 --- /dev/null +++ b/tests/test_plan_review_answer_texts.py @@ -0,0 +1,155 @@ +"""The mind-facing texts of the answer channel state FACTS in every branch: the open +questions by id, the route an answer takes (recorded at $0, merged by id, the unchanged +envelope asks only the named slots), the escalated question that stays open until its +answer is recorded, and the deadline/cap producers that report state instead of advising +a plan of their own. Each pin fails with the old wording restored.""" +from __future__ import annotations + +import json + +from ouroboros.tools.plan_render import _next_step +from tests.test_plan_review_engine import ( # noqa: F401 + CLEAN, _call, _control, _finding, _state, harness, +) + +FP = "f" * 64 + + +def _wave(findings: list, aggregate: str = "REVIEW_REQUIRED", dispositions: list | None = None) -> dict: + return {"aggregate": aggregate, "closed": False, "request_fingerprint": FP, + "findings": findings, "dispositions": dispositions or []} + + +def _q(fid: str, breaks: str = "claim_1", locator: str = "") -> dict: + return {"finding_id": fid, "slot": fid.split(":")[0], "class": "need_evidence", "breaks": breaks, "locator": locator} + + +def test_next_step_names_open_questions_by_id(): + for aggregate in ("REVIEW_REQUIRED", "REVISE_PLAN"): + findings = [_q("s1:q1"), _q("s2:q2", "goal"), _q("s3:e1", "claim_1", "notes.md::tail=16"), + {"finding_id": "s1:f1", "slot": "s1", "class": "blocking", "breaks": "claim_1"}] + text = _next_step(_wave(findings, aggregate), enforcement="blocking", cap=5, cycles_paid=1) + assert "Open questions to you: s1:q1 (claim_1), s2:q2 (goal). " in text + assert "Open document requests: s3:e1 (notes.md::tail=16). " in text + # answered questions leave the line; with none open the line is absent (two-sided) + answered = [{"finding_id": fid, "decision": "accept", "rationale": "answered"} for fid in ("s1:q1", "s2:q2", "s3:e1")] + text = _next_step(_wave(findings, aggregate, answered), enforcement="blocking", cap=5, cycles_paid=1) + assert "Open questions to you" not in text and "Open document requests" not in text + many = [_q(f"s1:q{i}") for i in range(10)] + text = _next_step(_wave(many), enforcement="advisory", cap=5, cycles_paid=1) + assert "s1:q7 (claim_1), +2 more. " in text and "s1:q9" not in text + + +def test_next_step_states_the_answer_route_in_every_open_branch(monkeypatch): + from ouroboros.tools import plan_render + + blocking = [{"finding_id": "s1:f1", "slot": "s1", "class": "blocking", "breaks": "claim_1"}] + for aggregate in ("REVIEW_REQUIRED", "REVISE_PLAN"): + text = _next_step(_wave(blocking, aggregate), enforcement="blocking", cap=5, cycles_paid=1) + assert f"plan_task(review_disposition={{review_fingerprint: '{FP}', items: [...]}})" in text + assert "asks only the slots its items name — one paid cycle 2 of 5" in text + assert "every other slot keeps its recorded answer at $0" in text and "Whether to buy that cycle is yours" in text + assert "supersedes this wave (its recorded answers stay); the identical one replays free" in text + for stale in ("reaches reviewers on the next paid cycle", "rides into the next paid delta cycle", + "can no longer be dispositioned", "Swarm", "accept ⇒ change the spec"): + assert stale not in text, stale + below = _next_step(_wave(blocking), enforcement="blocking", cap=5, cycles_paid=1) + assert "stay open after any disposition; they leave the review when the slot that raised them no longer raises it" in below + advisory = _next_step(_wave(blocking), enforcement="advisory", cap=5, cycles_paid=1) + assert "a reject with its rationale closes each one; accept or defer keeps it open" in advisory + # Cyber Pro: the facts and the route ride the branch too, never for a closed wave. + monkeypatch.setattr(plan_render, "review_enforcement_blocks", lambda _mode: False) + cyber = _next_step(_wave([_q("s1:q1")]), enforcement="blocking", cap=5, cycles_paid=1) + assert cyber.startswith("Cyber Pro: Ouroboros decides") and "Open questions to you: s1:q1 (claim_1). " in cyber + assert "An answer is recorded at $0" in cyber + closed = _next_step({**_wave([]), "aggregate": "GREEN", "closed": True}, enforcement="blocking", cap=5, cycles_paid=1) + assert "An answer is recorded" not in closed + + +def test_cap_reached_texts_report_state_and_authority_without_an_escape_recipe(): + blocking = [{"finding_id": "s1:f1", "slot": "s1", "class": "blocking", "breaks": "claim_1"}] + text = _next_step(_wave(blocking, "REVISE_PLAN"), enforcement="blocking", cap=2, cycles_paid=2) + assert "no paid cycle remains to show it to a reviewer: it stays recorded evidence" in text + assert "finalization is RELEASED" in text and "implementation still held" in text + assert "owner authority, never reviewer approval" in text and "Swarm" not in text and "unstick" not in text + + +def test_disposition_schema_states_the_answer_route_and_escalation(): + from ouroboros.tools.plan_review import _DISPOSITION_SCHEMA, get_tools + + description = _DISPOSITION_SCHEMA["description"] + assert "merged by finding_id" in description + assert "the unchanged envelope asks again ONLY the slots whose findings the items name" in description + assert "a changed envelope is reviewed by every slot with the answers in view" in description + assert "A quiz answer never closes a finding by itself" in description + assert "defer (rationale names the quiz) while the quiz is open, accept with the owner's decision" in description + assert "subsequent paid delta review" not in description + items = _DISPOSITION_SCHEMA["properties"]["items"] + assert "those slots are asked again" in items["description"] + assert "the quiz id while it is open (defer)" in items["items"]["properties"]["rationale"]["description"] + tool = next(t for t in get_tools() if t.name == "plan_task") + desc = tool.schema["description"] + assert "your answer re-judged by the slot that raised the finding" in desc + assert "asks again only the slots those items name" in desc and "subsequent paid delta" not in desc + effort = tool.schema["parameters"]["properties"]["reviewer_effort"]["description"] + assert "ignored when only collecting a recorded wave" in effort and "answering" not in effort + + +def test_blocking_reminder_states_the_answer_route(): + from ouroboros.owner_hurry import plan_review_reminder + + held = plan_review_reminder({"outcome": "REVIEW_REQUIRED", "status": "open", "cycles_paid": 1}) + assert "REVIEW_REQUIRED" in held and "the unchanged envelope with your items asks only that slot" in held + assert "do not rerun reviewers" not in held and "Implementation stays held" in held + revise = plan_review_reminder({"outcome": "REVISE_PLAN", "status": "open", "cycles_paid": 1}) + assert "asks only the slots those items name" in revise and "a new envelope every slot reviews" in revise + assert "the work starts after the review closes or a real task-wide rail fires" in revise + + +def test_deadline_and_cap_producers_state_facts_not_advice(harness, monkeypatch): # noqa: F811 + """The deadline rail names remaining time and the absent review; the cap rail names the + open review, the released finalization and the owner's authority — neither advises the + mind to proceed with a plan of its own or to seek a reviewer-free escape.""" + from ouroboros.tools import plan_review as pr + from ouroboros.deadline_utils import utc_now + from ouroboros.tools.plan_review_runtime import plan_deadline_skip + + ctx = harness.make_ctx() + ctx.task_metadata["deadline_at"] = (utc_now()).isoformat() + monkeypatch.setattr(pr, "_plan_deadline_skip", plan_deadline_skip) + text = plan_deadline_skip(ctx) + assert text.startswith("PLAN_TASK_SKIPPED_DEADLINE:") and "remaining time 0s" in text + assert "no plan review is open" in text and "proceed with your own" not in text.lower() + monkeypatch.setenv("OUROBOROS_REVIEW_MAX_CYCLES", "1") + harness.install({"s1": json.dumps([_finding("f1", "blocking", breaks="claim_1")]), "s2": CLEAN, "s3": CLEAN}) + ctx2 = harness.make_ctx(task_id="task-cap") + _call(ctx2) + spent = _call(ctx2, plan="A revised outline.") + assert "PLAN_REVIEW_CYCLES_EXHAUSTED" in spent and "finalization is RELEASED" in spent + assert "outcome_tier=blocked_with_evidence" in spent and "owner authority, never reviewer approval" in spent + assert "Swarm" not in spent and "Do not start the work" not in spent + + +def test_quiz_answer_never_closes_a_finding(harness): # noqa: F811 + """A doc-only rule with no host shortcut to revert: an owner quiz answered through + owner_quiz leaves the plan-review state byte-identical and the question open; the mind's + later accept citing the decision is what closes it (two-sided).""" + from ouroboros import owner_quiz + from ouroboros.task_results import load_plan_review_state + from ouroboros.tools import plan_review as pr + + harness.install({"s1": json.dumps([_finding("q1", "need_evidence", breaks="claim_1", summary="which?")]), + "s2": CLEAN, "s3": CLEAN}) + ctx = harness.make_ctx() + assert _control(_call(ctx)) == {"outcome": "REVIEW_REQUIRED", "closed": False} + before = json.dumps(load_plan_review_state(harness.drive, ctx.task_id), sort_keys=True) + fp = _state(harness)["waves"][-1]["request_fingerprint"] + owner_quiz.record_asked(harness.drive, ctx.task_id, quiz_id="quiz-1", question="which?", + options=["the first", "the second"]) + answered = owner_quiz.record_answered(harness.drive, ctx.task_id, quiz_id="quiz-1", option_index=1, + request_id="req-1") + assert answered["ok"] + assert json.dumps(load_plan_review_state(harness.drive, ctx.task_id), sort_keys=True) == before + closed = pr._handle_plan_task(ctx, review_disposition={"review_fingerprint": fp, "items": [ + {"finding_id": "s1:q1", "decision": "accept", "rationale": "owner decided: the second (quiz-1)"}]}) + assert _control(closed) == {"outcome": "GREEN", "closed": True} diff --git a/tests/test_plan_review_epoch.py b/tests/test_plan_review_epoch.py index 2f2775dcb..0c3ce508e 100644 --- a/tests/test_plan_review_epoch.py +++ b/tests/test_plan_review_epoch.py @@ -568,27 +568,26 @@ def test_closed_and_pending_plan_states_do_not_advertise_new_review_at_cap(): def test_revise_plan_render_names_only_available_paid_cycles_and_exits(enforcement, cap): from ouroboros.tools.plan_render import _next_step - text = _next_step({"aggregate": "REVISE_PLAN", "closed": False}, - enforcement=enforcement, cap=cap, cycles_paid=2) - assert "A disposition never closes REVISE_PLAN" in text + text = _next_step({"aggregate": "REVISE_PLAN", "closed": False}, enforcement=enforcement, cap=cap, cycles_paid=2) + assert "never closes REVISE_PLAN" in text + assert "rides into the next paid delta cycle" not in text and "reaches reviewers on the next paid cycle" not in text if cap == 2: - assert "rides into the next paid delta cycle" not in text - assert "no further paid delta cycle" in text + assert "no paid cycle remains" in text and "asks only the slots" not in text if enforcement == "blocking": - assert "owner unstick (Swarm/hurry)" in text + assert "owner authority, never reviewer approval" in text and "Swarm" not in text assert "once the owner raises OUROBOROS_REVIEW_MAX_CYCLES" in text assert "outcome_tier=blocked_with_evidence" in text else: - assert "rides into the next paid delta cycle" in text + assert "asks only the slots its items name" in text and "one paid cycle 3" in text assert "OUROBOROS_REVIEW_MAX_CYCLES" not in text if enforcement == "advisory": assert "Advisory enforcement: you may proceed" in text assert "host discloses" not in text and "your own final answer" in text - evidence_text = _next_step({"aggregate": "REVIEW_REQUIRED", "closed": False}, - enforcement=enforcement, cap=cap, cycles_paid=2) - assert "ONE $0 call" in evidence_text and "no reviewer call, no cycle" in evidence_text - assert ("reaches reviewers on the next paid cycle" in evidence_text) is (cap != 2) - assert ("no further paid cycle" in evidence_text) is (cap == 2) + evidence_text = _next_step({"aggregate": "REVIEW_REQUIRED", "closed": False}, enforcement=enforcement, cap=cap, cycles_paid=2) + assert "An answer is recorded at $0" in evidence_text and "merged by finding_id" in evidence_text + assert "reaches reviewers on the next paid cycle" not in evidence_text + assert ("no paid cycle remains" in evidence_text) is (cap == 2) + assert ("asks only the slots its items name" in evidence_text) is (cap != 2) def test_loop_reminder_does_not_repromise_a_spent_panel(monkeypatch): diff --git a/tests/test_plan_review_w3.py b/tests/test_plan_review_w3.py index 4f67bb90f..dec7ee67a 100644 --- a/tests/test_plan_review_w3.py +++ b/tests/test_plan_review_w3.py @@ -591,10 +591,11 @@ def test_a_truncated_requested_document_dispatches_with_the_cut_named(harness): assert "cannot_verify" not in out -def test_a_compacted_paid_predecessor_degrades_to_a_fresh_dispatch(harness, monkeypatch): - """When the prior exact wave is gone (compacted out of the hot state), the evidence - continuation is a cache miss: the wave re-dispatches fresh and every slot row - discloses the typed `prior_exact_wave_missing` cause.""" +def test_a_predecessor_the_state_cannot_name_dispatches_fresh_without_a_delta(harness, monkeypatch): + """When no paid predecessor can be named at all, the cycle has nothing to continue: it + dispatches fresh like a first cycle and discloses nothing (a NAMED predecessor whose exact + record is gone is the per-slot `prior_exact_wave_ref_missing` cause, pinned in + ``tests/test_phase4_plan_review_continuity.py``).""" (harness.workspace / "notes.md").write_text("deck notes\n", encoding="utf-8") ask = json.dumps([_finding("f1", "need_evidence", breaks="goal", locator="notes.md", summary="I need the notes")]) @@ -606,15 +607,14 @@ def test_a_compacted_paid_predecessor_degrades_to_a_fresh_dispatch(harness, monk assert len(sub.calls) == 1 # dispatched, not refused wave = _state(harness)["waves"][-1] assert wave["paid"] is True - deltas = [d for row in wave["actors"] for d in row.get("capability_delta") or []] - assert deltas and all(d["kind"] == "capability_delta" for d in deltas) - assert {d["reason"] for d in deltas} == {"prior_exact_wave_missing"} + assert not [d for row in wave["actors"] for d in row.get("capability_delta") or []] assert "cannot_verify" not in out def test_first_cycle_and_no_request_delta_cycle_carry_no_continuation_delta(harness): - """The `reviewer_requested` guard is load-bearing: a cycle with no reviewer request - has no prior thread to continue, so it must not disclose a missing predecessor.""" + """A first cycle has nothing to continue and discloses nothing; a delta cycle with no + reviewer request CONTINUES every packet slot's recorded transcript, so it discloses + nothing either.""" sub = harness.install({"s1": CLEAN, "s2": CLEAN, "s3": CLEAN}) _call(harness.make_ctx()) wave = _state(harness)["waves"][-1] @@ -829,7 +829,7 @@ def test_reviewer_question_holds_the_wave_until_a_free_disposition_and_its_answe [finding] = wave["findings"] assert finding["class"] == "need_evidence" and finding["breaks"] == "claim_1" and finding["locator"] == "" assert _state(harness).get("need_evidence_seen", []) == [] # a question is not a locator the host attaches - assert "a question addressed to you by spec id" in first and "defer = deferred openly" in first + assert "Open questions to you: s1:q1 (claim_1)." in first and "defer = deferred openly" in first answered = pr._handle_plan_task(ctx, review_disposition={ "review_fingerprint": wave["request_fingerprint"], "items": [{"finding_id": "s1:q1", "decision": "accept", "rationale": "The board asked for five."}]}) diff --git a/tests/test_reference_book_budgets.py b/tests/test_reference_book_budgets.py index e6546a73e..74fa54403 100644 --- a/tests/test_reference_book_budgets.py +++ b/tests/test_reference_book_budgets.py @@ -240,7 +240,9 @@ CHAPTER_BYTE_BUDGETS: dict[str, int] = { # 321400 -> 321600 (measured 321394): the closure sentence names the CONFIGURED enforcement against the hurry-projected advisory. # 321600 -> 321800 (measured 321594): the unanswered-slot floor names the terminal-absence rule and the lineage walk. # 323400 -> 327400 (merge of the moved target into the plan-review branch, measured 327146): both sides' paragraphs land together. - "docs/architecture/06-agent-core.md": 327400, + # 327400 -> 328400 (measured 328111): the addressed answer replaces the automatic delta; answers merge by + # finding_id; continuation per slot; the escalated-question clause (each extended in place). + "docs/architecture/06-agent-core.md": 328400, # 36991 -> 37300: the facade paragraph names the three loop constants runtime_limits.py # gained (events batch bound, budget-projection retry interval); no older text to displace. # 37300 -> 38400 (PR #1207): the Z.ai (`zai::`) direct provider gets its own route @@ -373,7 +375,9 @@ CHAPTER_BYTE_BUDGETS: dict[str, int] = { # 98300 -> 98600: the plan-review bullet gains the numbered-conversation delivery and the # observed-source read fact (extended in place) and two test pointers. # 98600 -> 99300 (merge of the moved target into the plan-review branch, measured 99030): both sides' paragraphs land together. - "docs/development/06-rules-by-change-class.md": 99300, + # 99300 -> 99800 (measured 99712): one bullet — plan-review answers merge by finding_id, the addressed re-ask + # and its $0 replay rows, the quiz-answer rule; the base sat 270 bytes under. + "docs/development/06-rules-by-change-class.md": 99800, "docs/development/07-managed-update-rule.md": 4166, "docs/development/08-mutation-attribution-rule.md": 2899, "docs/development/09-process-custody-rule.md": 10028,