mirror of
https://github.com/razzant/ouroboros.git
synced 2026-10-03 04:07:04 +00:00
Delete the paid fence wait: a supervisor that did not answer is a gap, not a no
The acceptance gate used to buy a model round whenever the queue-owned admission fence did not come up: `[TASK ACCEPTANCE WAIT] … retry after the queue fence is available`, on every install that is not Cyber Pro. The model can do nothing about a supervisor that did not answer, so on the night of 2026-09-19 the paid loop only bought rounds. Owner decision: the branch is deleted. A begin that refused or never answered lets the panel run on the EXISTING rail `admission_fence_available=False` (Cyber Pro already does); the durable `supervisor_ack_unavailable` row from the transport package stays the only record. The final seal reads the queue's typed answer instead of a bool (`loop_delivery._seal_admission_before_delivery`; `_no_tool_final_answer` shrinks 294→287): `ok` seals; a generation mismatch — now also one seen locally while the transport was silent, `_end_task_acceptance_fence` records it whatever the queue answered — is a real owner follow-up and keeps today's revision path; a begin refused because the root is already `sealed` is the worker's own earlier seal whose ack was lost; `unknown` delivers with `admission_released=False`, never `owner_revision_required`, never a paid round, and on a blocking install whose reviewers approved this subject the decision becomes `accepted` / `admission_close_unconfirmed` (owner 2A). An advisory author finish likewise survives a silent end and is set aside only by a refusal. For the seal read the agent seam now refuses with the row's typed state (`sealed`, `active`, `released`) instead of the supervisor's prose; a malformed request keeps its error text. Rewritten pins: test_acceptance_fence_outcome (typed refusal reason), test_acceptance_optional_control ready/held_effect (it asserted the deleted wait text; now: a refused fence buys no round, the new subject is reviewed on the rail), the writer inventory (22→23). Docs: ARCH 06 "Task acceptance" REPLACED byte-negative (−16), DEV 06 pointer (+2), ARCH 10 invariant 25 now states the class rule and its consequence. Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
This commit is contained in:
parent
fec17479d6
commit
16f32b404a
11 changed files with 371 additions and 43 deletions
|
|
@ -30,17 +30,17 @@ Disclosed cancel-lifecycle residuals (deliberate): a cascade over a tree with no
|
|||
|
||||
Host-enforced task acceptance is a root-owned completion coach, not the P3 commit gate. `off` disables it; `auto` and `required` review observable effects, typed deliverables/criteria, and an explicit root `task_acceptance_review` nomination, read-only research included. Queue membership alone does not qualify; ordinary conversation, exploration or cognitive-memory updates alone do not qualify in `auto`, and no prose or tool-count classifier decides their meaning. Child reviews remain advisory evidence superseded by the root decision.
|
||||
|
||||
The explicit call nominates the complete ready result and returns `deferred_to_host_acceptance`, `authoritative=false`; after the whole tool-result block the host advances the same acceptance operation ordinary final delivery uses, and early feedback does not seal the task. Main authors the effective criteria, and the paid subject is result bytes, criteria and material effects, so a status question preserves a running review while a new criterion can buy review of unchanged text. A reviewer panel is advice for its author, never a signature on bytes the reviewers did not read: the wave wakes the original Main through the mailbox at its quorum and again when the last slot settles, each wake carrying each reviewer's own verdict (`acceptance_settlement.announce_acceptance_settlement`). The optional answer control is prepared before either ready-feedback shortcut, so a settled panel or queued wake skips parking without losing control provenance. Mailbox readiness is distinct from owner authority: transport and waits wake for all entries, while acceptance source capture, acknowledgement and final sealing use the drain's typed owner/control boundary. System, descendant and independent-task messages alone do not imply an owner revision; owner/principal messages, quiz answers and typed owner controls retain their handling. Only a still-pending panel may park the turn; `acceptance_settlement.awaited_panel_has_settled` lets the next round, including control repair, run once the panel is settled. Final delivery without re-nomination uses the same task-panel feedback whether it returned ready or pending; delivered feedback stays on its exact `review_runs` row, never in phantom pending state. For a pending panel Main waits (the default, and the only option under blocking enforcement; Cyber Pro keeps its own rule — Main's final response is its decision) or consciously finishes through the `pending_review` key of the delivery control; a panel that settled PASS on the earlier revision accepts the task on the reviewers' word (`previous_revision_accepted`; the owner row says the current version was not re-reviewed) only when nothing but the answer text changed — the same owner source, criteria and material evidence — while a changed subject and any other settled verdict hand delivery to the ordinary path with the collected verdicts in its dialogue history. Before a new panel is assembled and before any capacity refusal, every recorded still-pending panel of the same root is collected at $0 over its recorded request and roster (`review_dispatch.reconcile_pending_acceptance_runs`), so a subject re-authored mid-flight cannot discard verdicts the tree already bought; the collected verdicts enter the next panel's dialogue history outside the hashed material, so reading them mints no paid binding. Only actual final delivery seals ingress; stop, missing custody and unfinished work keep their observed outcomes.
|
||||
The explicit call nominates the complete ready result and returns `deferred_to_host_acceptance`, `authoritative=false`; after the whole tool-result block the host advances the same acceptance operation ordinary final delivery uses, and early feedback does not seal the task. Main authors the effective criteria, and the paid subject is result bytes, criteria and material effects, so a status question preserves a running review while a new criterion can buy review of unchanged text. A reviewer panel is advice for its author, never a signature on bytes the reviewers did not read: the wave wakes the original Main through the mailbox at its quorum and again when the last slot settles, each wake carrying each reviewer's own verdict (`acceptance_settlement.announce_acceptance_settlement`). The optional answer control is prepared before either ready-feedback shortcut, so a settled panel or queued wake skips parking without losing control provenance. Mailbox readiness is distinct from owner authority: transport and waits wake for all entries, while acceptance source capture, acknowledgement and final sealing use the drain's typed owner/control boundary. System, descendant and independent-task messages alone do not imply an owner revision; owner/principal messages, quiz answers and typed owner controls retain their handling. Only a still-pending panel may park the turn; `acceptance_settlement.awaited_panel_has_settled` lets the next round, including control repair, run once the panel is settled. Final delivery without re-nomination uses the same task-panel feedback whether it returned ready or pending; delivered feedback stays on its exact `review_runs` row, never in phantom pending state. For a pending panel Main waits (the default; the only option under blocking; Cyber Pro: Main's final response is its decision) or consciously finishes through the `pending_review` key of the delivery control; a panel that settled PASS on the earlier revision accepts the task on the reviewers' word (`previous_revision_accepted`, said on the owner row) only when nothing but the answer text changed — the same owner source, criteria and material evidence — while a changed subject and any other settled verdict hand delivery to the ordinary path with the collected verdicts in its dialogue history. Before a new panel or any capacity refusal, every recorded still-pending panel of the same root is collected at $0 over its recorded request and roster (`review_dispatch.reconcile_pending_acceptance_runs`), so a re-authored subject cannot discard verdicts the tree already bought; they enter the next panel's dialogue history. Stop, missing custody and unfinished work keep their observed outcomes.
|
||||
|
||||
A completed reviewer from an older plan wave is attached as a historical supplement through the locked task-result writer and exact producer CAS: it never rewrites the original verdict, aggregate, closure, author dispositions or current-wave pointer, settles its historical cost without another cycle, and reaches a terminal parent without a new model turn. Task acceptance has the same twin: a panel that settles after its task is terminal is collected at $0 over the recorded operation, republished on the task's review projection with a host-composed `late_settlement` note (the verdict, the revision it covered, settled after the terminal; honest about a reviewer whose physical outcome is still unknown) and announced once in the task's room as a System row stamped `card_row="reviews"` (`acceptance_settlement.attach_late_acceptance_settlement`, deduped by delivery id) — a timeline item of the card, read by the next turn from chat history; no model turn starts.
|
||||
|
||||
Before an eligible panel runs, `supervisor/queue_transitions.py` closes subtask admission under `_queue_lock` (a direct turn in-process, admission lock first, never the reverse; a pooled worker by event and a `state/acceptance_fence_acks/` ack, transport only) and `task_status.find_child_tasks` proves the subtree quiescent; early settlement releases that fence, only final delivery seals ingress, and a changed subject reopens review, keeping the earlier one. Reads use the canonical `budget_drive_root`. The reviewer packet carries verbatim owner directives, the full contract and criteria, canonical deliverable identity, terminal child state, verification receipts, artifact references, touched-skill lifecycle facts (visibility, never a gate) and an explicit omissions manifest; a required component that cannot be assembled makes the affected actor `DEGRADED`, never a silently smaller prompt. The packet is SIZED against the quorum's real windows (the `reviewer_window`/`review_synthesis.quorum_input_token_limit` seam the triad and plan review use), resolved once per task so the bytes cannot drift between the binding build and the staleness rebuild. Non-core sections shed through a DISCLOSED ladder — predecessor authority envelope, trajectory tail and its results, artifact previews, agent-supplied evidence, last a diff preview that keeps the durable `repo_diff_source_ref` — each shed a row in `omissions_manifest`. A slot whose own window cannot hold the rendered prompt is a typed `$0 not_dispatched` row while the rest of the panel reviews; a packet that still overflows after every shed stamps `__immutable_core_overflow__` naming the oversized sections and refuses the panel without spending. `__unresolved_partial_artifacts__` withholds packet rows only for a tool result whose exact source is genuinely `source_unavailable` (a retrieving row reads the source itself) — a budget shed with a durable, actor-resolvable source ref is an omission, never an unresolved partial.
|
||||
Before an eligible panel runs, `supervisor/queue_transitions.py` closes subtask admission under `_queue_lock` (transport: §10 invariant 25) and `task_status.find_child_tasks` proves the subtree quiescent; a fence that refused or never answered buys no model round — the panel runs on the disclosed rail `admission_fence_available=false`. Early settlement releases the fence; final delivery seals ingress by the queue's typed answer (`loop_delivery._seal_admission_before_delivery`: the worker's own earlier seal, a real owner follow-up, or no answer — delivered with `admission_released=false` and, where blocking reviewers approved, noted `admission_close_unconfirmed`); a changed subject reopens review, keeping the earlier one. Reads use the canonical `budget_drive_root`. The reviewer packet carries verbatim owner directives, the full contract and criteria, canonical deliverable identity, terminal child state, verification receipts, artifact references, touched-skill lifecycle facts (visibility, never a gate) and an explicit omissions manifest; a required component that cannot be assembled makes the affected actor `DEGRADED`, never a silently smaller prompt. The packet is SIZED against the quorum's real windows (the `reviewer_window`/`review_synthesis.quorum_input_token_limit` seam the triad and plan review use), resolved once per task so the bytes cannot drift between the binding build and the staleness rebuild. Non-core sections shed through a DISCLOSED ladder — predecessor authority envelope, trajectory tail and its results, artifact previews, agent-supplied evidence, last a diff preview that keeps the durable `repo_diff_source_ref` — each shed a row in `omissions_manifest`. A slot whose own window cannot hold the rendered prompt is a typed `$0 not_dispatched` row while the rest of the panel reviews; a packet that still overflows after every shed stamps `__immutable_core_overflow__` naming the oversized sections and refuses the panel without spending. `__unresolved_partial_artifacts__` withholds packet rows only for a tool result whose exact source is genuinely `source_unavailable` (a retrieving row reads the source itself) — a budget shed with a durable, actor-resolvable source ref is an omission, never an unresolved partial.
|
||||
|
||||
The packet is delivery-conditional; the FULL packet is not. Every triad row reaches the panel as configured (`reviewer_slot_config.triad_delivery_slots`): a packet (`api_chat`) row receives the assembled packet; a retrieving row receives the route-owned work order `loop_acceptance_review.acceptance_retrieving_work_order` writes onto `ReviewRequest.slot_session_tasks` — the same task-stable contract and output contract (`review_execution.review_output_contract` as `policy["output_contract"]`), absolute pointers to the task's ACTIVE workspace (`review_repo_dirs_for`'s subject root, never the governance repo, as `session_root`) and its result, artifacts and receipts, and the packet in the form its delivery can use. An agent-session row gets the FULL packet — its run is unobserved by the host, so the packet is its only attested view — with the disclosure that access outside the workspace is not guaranteed and a refused read is absence of evidence, not of the artifact; a native inspection row gets the packet WITHOUT its freely degradable tail (tool-trajectory rows and artifact previews, manifested as `retrieving_delivery` omissions) plus the real data root (`policy["native_data_root"]`), because its episode reads those sources itself under `host_file_read_attestation: host_observed`. Reviewer `evidence_refs` resolve against the FULL packet on every delivery, never the rendered projection, so a retrieving row citing a real receipt is clean exactly like a packet row. The wave budget gate prices API money only and DECIDES on one work-order send per paid row (a session row rides the owner's subscription and is not priced); that admission is the whole money rule — no rounds multiplier, no second read-only pricing pass — so a panel's cost is bounded at dispatch by the per-send wallet binding rather than predicted; only a floor that does not fit is refused `review_wave_budget_insufficient`, and a packet row's second physical send for format repair never buys a retrieving row a second episode or session. Child-task and `off`-mode acceptance stay advisory and run packet rows only.
|
||||
|
||||
Paid identity binds that semantic subject together with substantive nonempty obligation dispositions (`acceptance_paid_identity`); forensic source hashes and ingress counters alone do not buy a panel. A resubmit with the same paid identity reuses its recorded verdict for free: a clean replay can authorize acceptance, a non-clean replay keeps its verdict and the `identical_acceptance_refused` outcome — no repeated payment, and no cosmetic edit needed when real criteria or evidence change.
|
||||
|
||||
The configured slots are independent actors with adaptive quorum (`config.adaptive_quorum`: 2-of-N for N≥3, both for N=2, a single reviewer as loud `single_reviewer_no_diversity`; a fewer-responded shortfall stays a loud infra quorum failure); each receives one substantive interaction on its bound route — at most two physical sends for a packet row, one bounded episode for a native row, one delegated session for a session row — and a retrieving verdict is equally authoritative. Transport status, parse status, semantic verdict, criterion support, route, quorum contribution and binding hashes stay distinct, so an unavailable or malformed response cannot masquerade as a negative judgment; a panel that refuses before any transport projects `not_dispatched` on every row and on the panel, and a slot released at the dispatch barrier projects `awaiting` for transport and parse until it settles — distinct from `success`, `timeout` and `provider_transport_error`, never a failure or a verdict. `PASS`/`FAIL`/`DEGRADED` are reviewer verdicts; the host-owned completion decision is separately `accepted`, `revision_requested` or `finalized_unaccepted`, written only by `loop_acceptance._set_acceptance_decision`. A clean quorum supplies critic approval; an informed Advisory author finish supplies separate current-author authority. Neither changes the original verdict. Material FAIL and typed unavailable outcomes reach Main before it chooses how to respond, including after the last paid panel; a no-quorum outcome with real minority findings retains that partial feedback. Blocking without fresh approval remains unaccepted, and a terminal technical failure keeps `reason=review_degraded` when no permitted author completion follows.
|
||||
The configured slots are independent actors with adaptive quorum (`config.adaptive_quorum`: 2-of-N for N≥3, both for N=2, a single reviewer as loud `single_reviewer_no_diversity`; a fewer-responded shortfall stays a loud infra quorum failure); each receives one substantive interaction on its bound route — at most two physical sends for a packet row, one bounded episode for a native row, one delegated session for a session row — and a retrieving verdict is equally authoritative. Transport status, parse status, semantic verdict, criterion support, route, quorum contribution and binding hashes stay distinct, so an unavailable or malformed response cannot masquerade as a negative judgment; a panel that refuses before any transport projects `not_dispatched` on every row and on the panel, and a slot released at the dispatch barrier projects `awaiting` for transport and parse until it settles — distinct from `success`, `timeout` and `provider_transport_error`, never a failure or a verdict. `PASS`/`FAIL`/`DEGRADED` are reviewer verdicts; the host-owned completion decision is separately `accepted`, `revision_requested` or `finalized_unaccepted`, written only by `loop_acceptance._set_acceptance_decision`. A clean quorum supplies critic approval; an informed Advisory author finish supplies separate current-author authority. Neither changes the original verdict. Material FAIL and typed unavailable outcomes reach Main before it chooses how to respond, including after the last paid panel; a no-quorum outcome with real minority findings retains that partial feedback. A terminal technical failure keeps `reason=review_degraded` when no permitted author completion follows.
|
||||
|
||||
A clean criterion is evidence-resolved, not merely well argued: reviewer `evidence_refs` must be exact members of the packet's enumerable reference vocabulary, and a claim id resolves only through `acceptance_support_refs` linked to a passing host receipt for that claim. Agent-supplied, declared-intent, unattested and non-resolving sections never certify success; an OPEN plan wave binds nothing — its claims are disclosed as `acceptance_claims_source='none_open_plan_wave'` beside a non-binding `plan_claims_exhibit` inside `DECLARED_INTENT_SECTIONS`, so citing it never resolves and the task is distinguishable from one that never had claims. An unresolved reference keeps the actor's record for audit but removes its clean contribution (`criteria_refs_unresolved`). This total, fail-closed resolver is why the task cannot certify itself by echoing its expected outcome.
|
||||
|
||||
|
|
|
|||
|
|
@ -26,7 +26,7 @@ This chapter is the short list of properties the rest of the book must not contr
|
|||
22. **A host-owed round never parks.** A turn parks behind a review panel only when the panel is the sole thing it waits for; a turn in which the host has just spoken to the model never parks (`loop._finalize_loop_candidate`).
|
||||
23. **Every call has a bound, and a recorder speaks only for what it collected.** A reviewer's tool call runs under the loop's per-tool timeout narrowed by the inherited dispatch deadline; a call that outlives it is abandoned — its late value sources no receipt and no coverage. A deadline recorder reconciles the turn's own panel at $0 before it writes a terminal reason. Owners: `review_native_episode.py`, `loop_tool_execution.py`, `acceptance_settlement.py`.
|
||||
24. **The thread that answers workers runs only queue-bounded work.** Work that scales with history, the daemon or the network runs off-thread, reads its candidates before it reads liveness (one in-memory live source under `_queue_lock`), stops mutating when its loop generation ends and reaches the daemon attach-only once a stop is in flight. Named residuals on the loop thread: the 300-s zombie reconcile and the usage-ledger lock in the heartbeat handler. Owner: `ouroboros/server_maintenance.py`.
|
||||
25. **A missing supervisor answer is a gap, never a refusal.** A direct turn applies its acceptance fence in-process (admission lock, then `_queue_lock` — never the reverse); a pooled request is idempotent by token, acknowledged per request (`<token>.<req>.json`) and re-sent once; only `sealed` is a seal, and an absent row is never read as one. Owners: `ouroboros/agent.py`, `supervisor/queue_transitions.py`.
|
||||
25. **An answer that has not arrived is a gap — never a refusal, a failure, a verdict or an owner message.** A direct turn applies its acceptance fence in-process (admission lock, then `_queue_lock` — never the reverse); a pooled request is idempotent by token, acknowledged per request (`<token>.<req>.json`) and re-sent once; only `sealed` is a seal, an absent row is never read as one, and a fence that did not answer buys no model round: the panel runs as advice on `admission_fence_available=false`, delivery seals again, and a blocking install accepts a reviewer-approved answer with the typed note `admission_close_unconfirmed`. Owners: `ouroboros/agent.py`, `supervisor/queue_transitions.py`, `ouroboros/loop_delivery.py`.
|
||||
|
||||
### 10.1 Continuity data-flow map
|
||||
|
||||
|
|
|
|||
|
|
@ -931,7 +931,7 @@ and what enforces each.
|
|||
settles it.
|
||||
- Host acceptance: root-only, structured eligibility (`outcomes.turn_has_reviewable_effects`
|
||||
plus a typed deliverable/criterion), never keywords or authoritative agent nomination
|
||||
(BIBLE P3/P5; acceptance model, per-enforcement waiting, `previous_revision_accepted`,
|
||||
(BIBLE P3/P5; acceptance model, waiting, unanswered fence, `previous_revision_accepted`,
|
||||
`late_settlement`: ARCHITECTURE §6 "Task acceptance"). Freeze request/roster; existing
|
||||
review custody/mailbox handles pending/free collection. Before new-panel evidence or
|
||||
`review_cycles_exhausted`, reconcile every paid panel still running for that root: $0,
|
||||
|
|
|
|||
|
|
@ -283,8 +283,9 @@ class OuroborosAgent:
|
|||
ack = self._send_fence_event(event) or (request["action"] != "inspect" and self._send_fence_event(event)) or {}
|
||||
if not ack:
|
||||
raise TimeoutError(f"supervisor did not acknowledge acceptance fence {request['action']}")
|
||||
if not ack.get("ok", True) or str(ack.get("status") or "") not in accept:
|
||||
raise RuntimeError(str(ack.get("error") or f"acceptance fence {request['action']} failed"))
|
||||
status = str(ack.get("status") or "")
|
||||
if not ack.get("ok", True) or status not in accept: # the row's typed state is the reason; only a malformed request keeps its error text
|
||||
raise RuntimeError(status if status not in ("", "error") else str(ack.get("error") or f"acceptance fence {request['action']} failed"))
|
||||
return ack
|
||||
|
||||
def _send_fence_event(self, event: Dict[str, Any]) -> Dict[str, Any]:
|
||||
|
|
|
|||
|
|
@ -246,8 +246,8 @@ def _end_task_acceptance_fence(ctx: Any, *, outcome: str, admission_locked: bool
|
|||
if acquired:
|
||||
admission_lock.release()
|
||||
_drop_fence_binding(ctx) # also after a refusal or a gap: the next begin re-adopts or reopens
|
||||
ctx._task_acceptance_fence_generation_mismatch = generation_mismatch # local owner facts stand whatever the queue answered
|
||||
if result:
|
||||
ctx._task_acceptance_fence_generation_mismatch = generation_mismatch
|
||||
sealed = status == "sealed" or (not status and effective_outcome != "revision")
|
||||
ctx._task_acceptance_sealed_fence_token = token if sealed else None
|
||||
return result
|
||||
|
|
|
|||
|
|
@ -718,7 +718,8 @@ def _finish_advisory_author(ctx: _TaskAcceptanceContext) -> bool:
|
|||
capacity = project_task_acceptance_review_capacity(ctx.tools._ctx, task_id=ctx.task_id) if action == "stop" else {}
|
||||
terminal_reason = (REASON_REVIEW_CYCLES_EXHAUSTED if action == "stop" and capacity.get("reason") == REASON_REVIEW_CYCLES_EXHAUSTED
|
||||
else "author_stop" if action == "stop" else "author_finish")
|
||||
if not _loop()._end_task_acceptance_fence(ctx.tools._ctx, outcome="terminal"):
|
||||
ended = _loop()._end_task_acceptance_fence(ctx.tools._ctx, outcome="terminal")
|
||||
if ended.status == "refused": # a gap is not a refusal: the final seal asks again and discloses
|
||||
_loop()._supersede_task_acceptance_for_owner_followup(ctx.tools._ctx, ctx.llm_trace)
|
||||
return True
|
||||
ctx.tools._ctx._task_acceptance_reviewed = True
|
||||
|
|
@ -726,7 +727,8 @@ def _finish_advisory_author(ctx: _TaskAcceptanceContext) -> bool:
|
|||
_loop()._mark_root_acceptance_checkpoint(
|
||||
ctx.tools._ctx, ctx.llm_trace, status=author["reviewer_signal"].lower(), pass_index=ctx.passes_done,
|
||||
)
|
||||
ctx.llm_trace["review_decision"].update({"binding_hash": ctx.review_binding["binding_hash"], "author_finish": action == "finish"})
|
||||
ctx.llm_trace["review_decision"].update({"binding_hash": ctx.review_binding["binding_hash"], "author_finish": action == "finish",
|
||||
"admission_released": bool(ended)})
|
||||
_loop()._set_acceptance_decision(ctx.llm_trace, {
|
||||
"status": ACCEPTANCE_FINALIZED_UNACCEPTED, "reason": terminal_reason,
|
||||
"author_action": action, **({"review_capacity": capacity} if capacity else {}),
|
||||
|
|
@ -1349,19 +1351,9 @@ def _run_task_acceptance_review_once(
|
|||
set_decision=_loop()._set_acceptance_decision, emit_progress=emit_progress,
|
||||
):
|
||||
return False
|
||||
# A fence that answered no or not at all buys no model round: the panel runs on the
|
||||
# disclosed rail `admission_fence_available=False` and final delivery seals again.
|
||||
fence_ok, _fence_token = _loop()._begin_task_acceptance_fence(tools._ctx, task_id)
|
||||
if not fence_ok and review_enforcement_blocks("blocking"):
|
||||
llm_trace["review_decision"] = {
|
||||
"eligibility": "acceptance_fence_failed", "trigger": trigger,
|
||||
}
|
||||
_loop()._append_or_merge_user_message(
|
||||
messages,
|
||||
"[TASK ACCEPTANCE WAIT] The supervisor could not atomically close "
|
||||
"subtask admission. Do not finalize or spawn more work; retry after the "
|
||||
"queue fence is available.",
|
||||
)
|
||||
emit_progress("Task acceptance review waiting for the queue-owned admission fence.")
|
||||
return True
|
||||
quiescent, subtree_statuses = _loop()._task_acceptance_subtree_snapshot(
|
||||
tools._ctx, drive_root, task_id,
|
||||
)
|
||||
|
|
|
|||
|
|
@ -14,7 +14,7 @@ import queue
|
|||
from dataclasses import dataclass
|
||||
from typing import Any, Callable, Dict, List, Optional, Tuple
|
||||
from ouroboros.config import get_context_mode
|
||||
from ouroboros.outcomes import reviewable_effect_projection
|
||||
from ouroboros.outcomes import ACCEPTANCE_ACCEPTED, reviewable_effect_projection
|
||||
from ouroboros.task_finalization import set_terminal_host_notice
|
||||
from ouroboros.tools.registry import ToolRegistry
|
||||
from ouroboros.utils import sanitize_tool_result_for_log
|
||||
|
|
@ -1113,6 +1113,44 @@ def _plan_review_only_awaited(llm_trace: Dict[str, Any]) -> bool:
|
|||
return isinstance(plan_gate, dict) and plan_gate.get("review_only_awaited") is True
|
||||
|
||||
|
||||
def _seal_admission_before_delivery(tools: ToolRegistry, limit_ctx: Any, llm_trace: Dict[str, Any]) -> bool:
|
||||
"""Seal root admission once more right before delivery; False arms the owner-revision round.
|
||||
|
||||
The queue's typed answer decides. ``ok`` seals, and a generation mismatch (also one seen
|
||||
locally while the transport was silent) is a real owner follow-up. A begin ``refused``
|
||||
because the root is already ``sealed`` is the worker's own earlier seal whose ack was
|
||||
lost. Any other refusal keeps the revision path. ``unknown`` is a gap — never a refusal,
|
||||
a verdict or an owner message: the answer is delivered with ``admission_released=False``
|
||||
(the durable ``supervisor_ack_unavailable`` row is the record) and a blocking install
|
||||
whose reviewers approved this subject says so on the card (owner decision 2A).
|
||||
"""
|
||||
tool_ctx = tools._ctx
|
||||
opened, _token = _loop()._begin_task_acceptance_fence(tool_ctx, limit_ctx.task_id)
|
||||
answer = opened and _loop()._end_task_acceptance_fence(tool_ctx, outcome="terminal")
|
||||
own_seal = (opened.status, opened.reason) == ("refused", "sealed")
|
||||
if getattr(tool_ctx, "_task_acceptance_fence_generation_mismatch", False) or not (answer or own_seal or answer.status == "unknown"):
|
||||
_loop()._supersede_task_acceptance_for_owner_followup(tool_ctx, llm_trace)
|
||||
admission_lock = getattr(tool_ctx, "owner_message_admission_lock", None)
|
||||
admission_agent = getattr(tool_ctx, "owner_message_admission_agent", None)
|
||||
if admission_lock is not None and admission_agent is not None:
|
||||
with admission_lock:
|
||||
admission_agent._accepting_owner_messages = True
|
||||
_loop()._arm_delivery_control(tools, limit_ctx, llm_trace, control="owner_revision_required")
|
||||
return False
|
||||
if not answer and not own_seal:
|
||||
from ouroboros.review_projection import publish_acceptance_checkpoint
|
||||
from ouroboros.tools.review_helpers import review_enforcement_blocks
|
||||
|
||||
llm_trace.setdefault("review_decision", {})["admission_released"] = False
|
||||
decision = llm_trace.get("acceptance_decision") if isinstance(llm_trace.get("acceptance_decision"), dict) else {}
|
||||
if (decision.get("status") == ACCEPTANCE_ACCEPTED and review_enforcement_blocks()
|
||||
and decision.get("reason") in ("clean_pass", "clean_pass_obligations_closed")):
|
||||
_loop()._set_acceptance_decision(llm_trace, {**decision, "status": ACCEPTANCE_ACCEPTED, "reason": "admission_close_unconfirmed",
|
||||
"rationale": "Quorum PASS accepted the deliverable; the supervisor did not confirm that task admission was closed."})
|
||||
publish_acceptance_checkpoint(tool_ctx, llm_trace)
|
||||
return True
|
||||
|
||||
|
||||
def _no_tool_final_answer(
|
||||
content: Any,
|
||||
limit_ctx: _RoundLimitContext,
|
||||
|
|
@ -1386,16 +1424,9 @@ def _no_tool_final_answer(
|
|||
_loop()._publish_delivery_candidate(tools, candidate, llm_trace)
|
||||
content = candidate.full_text
|
||||
if (getattr(tools._ctx, "_task_acceptance_reviewed", False)
|
||||
and not getattr(tools._ctx, "_task_acceptance_sealed_fence_token", None)):
|
||||
opened, _token = _loop()._begin_task_acceptance_fence(tools._ctx, limit_ctx.task_id)
|
||||
sealed = opened and _loop()._end_task_acceptance_fence(tools._ctx, outcome="terminal")
|
||||
if not sealed or getattr(tools._ctx, "_task_acceptance_fence_generation_mismatch", False):
|
||||
_loop()._supersede_task_acceptance_for_owner_followup(tools._ctx, llm_trace)
|
||||
if admission_lock is not None and admission_agent is not None:
|
||||
with admission_lock:
|
||||
admission_agent._accepting_owner_messages = True
|
||||
_loop()._arm_delivery_control(tools, limit_ctx, llm_trace, control="owner_revision_required")
|
||||
return None
|
||||
and not getattr(tools._ctx, "_task_acceptance_sealed_fence_token", None)
|
||||
and not _seal_admission_before_delivery(tools, limit_ctx, llm_trace)):
|
||||
return None
|
||||
if isinstance(getattr(tools._ctx, "_presence_completion", None), dict):
|
||||
# Only this successful common exit accepts the requested outcome. Holds,
|
||||
# owner controls and budget exits must not inherit an earlier silent/send.
|
||||
|
|
|
|||
|
|
@ -118,10 +118,9 @@ def test_refused_begin_is_typed_refused_with_its_reason(monkeypatch, tmp_path, s
|
|||
finally:
|
||||
supervisor.stop()
|
||||
assert not outcome and token is None
|
||||
assert outcome.status == "refused" and "already sealed" in outcome.reason
|
||||
assert (outcome.status, outcome.reason) == ("refused", "sealed") # the row's typed state, not prose
|
||||
rows = _unavailable_rows(tmp_path)
|
||||
assert [(row["op"], row["outcome"]) for row in rows] == [("begin", "refused")]
|
||||
assert "already sealed" in rows[0]["reason"]
|
||||
assert [(row["op"], row["outcome"], row["reason"]) for row in rows] == [("begin", "refused", "sealed")]
|
||||
|
||||
|
||||
def test_begin_with_stale_token_rebinds_through_fresh_begin(tmp_path):
|
||||
|
|
|
|||
|
|
@ -258,6 +258,8 @@ def test_return_order_preserves_feedback_identity_and_new_subjects(full_loop, mo
|
|||
revised = ANSWER + " Budget: $12."
|
||||
_order_acceptance_feedback(f, monkeypatch, ANSWER, order)
|
||||
if next_action == "held_effect" and order == "ready":
|
||||
# The new subject's own admission fence is REFUSED (no token). That buys no model
|
||||
# round any more: its panel runs on the disclosed rail and the final seal asks again.
|
||||
begins = []
|
||||
def begin(**_kwargs):
|
||||
begins.append(True)
|
||||
|
|
@ -282,21 +284,19 @@ def test_return_order_preserves_feedback_identity_and_new_subjects(full_loop, mo
|
|||
assert any(source == {"task_id": f.run_args["task_id"], "run_index": index,
|
||||
"binding_hash": offered["binding_hash"]}
|
||||
for message in messages for source in message.get("review_feedback", []))
|
||||
if next_action == "effect":
|
||||
if next_action == "effect" or (next_action == "held_effect" and order == "ready"):
|
||||
return {"content": "", "tool_calls": [call("write_file", {
|
||||
"root": "task_drive", "path": "new-effect.txt", "content": "Additional evidence.",
|
||||
}, "effect")]}, 0.0
|
||||
if next_action == "nominate":
|
||||
return {"content": "", "tool_calls": [call("task_acceptance_review", {"claim": revised}, "second")]}, 0.0
|
||||
if f.model_step == 3 and next_action == "held_effect":
|
||||
if order == "ready":
|
||||
assert "supervisor could not atomically close" in str(messages)
|
||||
else:
|
||||
assert f.waits # The real pending-panel wait held the rewritten answer.
|
||||
if f.model_step == 3 and next_action == "held_effect" and order == "pending":
|
||||
assert f.waits # The real pending-panel wait held the rewritten answer.
|
||||
return {"content": "", "tool_calls": [call("write_file", {
|
||||
"root": "task_drive", "path": "new-effect.txt", "content": "Additional evidence.",
|
||||
}, "held-effect")]}, 0.0
|
||||
if (f.model_step == 2 or (f.model_step == 3 and next_action == "effect")
|
||||
or (f.model_step == 3 and next_action == "held_effect" and order == "ready")
|
||||
or (f.model_step == 4 and next_action == "held_effect")):
|
||||
subject = {"owner_source_sha256": f.ctx._acceptance_observation["owner_source_sha256"]}
|
||||
if next_action == "criterion":
|
||||
|
|
@ -319,6 +319,13 @@ def test_return_order_preserves_feedback_identity_and_new_subjects(full_loop, mo
|
|||
if next_action in {"effect", "held_effect"}:
|
||||
effect = f.ctx.drive_root / "task_drives" / f.run_args["task_id"] / "new-effect.txt"
|
||||
assert effect.read_text(encoding="utf-8") == "Additional evidence."
|
||||
if next_action == "held_effect" and order == "ready":
|
||||
# The refused fence bought no round: the new subject was reviewed at once on the
|
||||
# disclosed rail (the fourth round is the ordinary post-PASS control round, as for
|
||||
# `effect`), and the answered re-seal closed admission as usual.
|
||||
assert f.model_step == 4 and "TASK ACCEPTANCE WAIT" not in str(f.model_inputs)
|
||||
assert len(begins) == 3 # the nomination's, the refused one, the answered re-seal
|
||||
assert trace["acceptance_decision"]["reason"] == "clean_pass"
|
||||
if next_action == "rewrite":
|
||||
decision = trace["acceptance_decision"]
|
||||
assert decision["reason"] == "previous_revision_accepted"
|
||||
|
|
|
|||
|
|
@ -10,14 +10,26 @@ fence wait is gone; the final seal reads the queue's typed answer) lives below i
|
|||
|
||||
from __future__ import annotations
|
||||
|
||||
import copy
|
||||
import json
|
||||
import pathlib
|
||||
import queue as stdqueue
|
||||
import threading
|
||||
from types import SimpleNamespace
|
||||
|
||||
import pytest
|
||||
|
||||
from ouroboros import loop
|
||||
from ouroboros.loop_acceptance import ACCEPTANCE_DECISION_REASONS
|
||||
from ouroboros.outcomes import (
|
||||
ACCEPTANCE_ACCEPTED, ACCEPTANCE_FINALIZED_UNACCEPTED, OBJECTIVE_PASS, OUTCOME_TIER_BLOCKED,
|
||||
OUTCOME_TIER_SOLVED, _objective_axis,
|
||||
)
|
||||
from ouroboros.project_dialogue import TASK_CAUSE_PHRASES, _completion_verdict, outcome_phase
|
||||
from tests.test_acceptance_async_loop import ANSWER, _terminal_record, call, full_loop as _full_loop, keep # noqa: F401
|
||||
from tests.test_acceptance_fence_transport import _pooled_agent
|
||||
|
||||
full_loop = _full_loop # noqa: F811 - shared real-loop fixture
|
||||
|
||||
NOTE = "admission_close_unconfirmed"
|
||||
SENTENCE = "Reviewers approved this answer; the supervisor did not confirm that task admission was closed."
|
||||
|
|
@ -51,3 +63,288 @@ def test_an_accepted_answer_with_the_note_is_done_and_never_blocked():
|
|||
assert blocked["outcome_tier"] == OUTCOME_TIER_BLOCKED and blocked["reason"] == refusal
|
||||
# A clean accepted decision keeps rendering nothing: the note is the only accepted cell that speaks here.
|
||||
assert _completion_verdict(_record(ACCEPTANCE_ACCEPTED, "clean_pass"), {}) == ""
|
||||
|
||||
|
||||
# --- the behaviour half: no paid wait, the final seal reads the queue's typed answer -------
|
||||
|
||||
|
||||
def _unavailable_rows(ctx):
|
||||
from ouroboros.task_pacing import acceptance_timing_events_path
|
||||
|
||||
path = acceptance_timing_events_path(ctx)
|
||||
if not path.is_file():
|
||||
return []
|
||||
rows = [json.loads(line) for line in path.read_text(encoding="utf-8").splitlines() if line.strip()]
|
||||
return [row for row in rows if row.get("type") == "supervisor_ack_unavailable"]
|
||||
|
||||
|
||||
def _gap(**_kwargs):
|
||||
raise TimeoutError("supervisor did not acknowledge acceptance fence")
|
||||
|
||||
|
||||
def _silent_supervisor(f, *, begin=_gap, end=_gap):
|
||||
"""The pooled seam raised its gap: no answer arrived within the wait."""
|
||||
f.ctx.begin_acceptance_fence, f.ctx.end_acceptance_fence, f.ctx.inspect_acceptance_fence = begin, end, _gap
|
||||
|
||||
|
||||
def _nominate_then_keep(f, monkeypatch, *, nominate=True):
|
||||
"""Main nominates (or simply writes) its complete answer, gets the verdict, then keeps it."""
|
||||
if nominate: # the explicit nomination settles synchronously; a plain answer parks on its panel
|
||||
f.release.set()
|
||||
f.ctx.owner_wait_callback = None
|
||||
|
||||
def main(_llm, messages, *_a, **_kw):
|
||||
f.model_inputs.append(copy.deepcopy(messages))
|
||||
f.model_step += 1
|
||||
if f.model_step == 1:
|
||||
if not nominate:
|
||||
return {"content": ANSWER}, 0.0
|
||||
return {"content": "", "tool_calls": [call("task_acceptance_review", {"claim": ANSWER}, "review")]}, 0.0
|
||||
assert f.model_step < 5, ("the fence bought model rounds", f.progress)
|
||||
return keep(f), 0.0
|
||||
|
||||
monkeypatch.setattr(loop, "call_llm_with_retry", main)
|
||||
|
||||
|
||||
def test_blocking_never_acked_begin_and_seal_accepts_with_the_note(full_loop, monkeypatch):
|
||||
"""Owner 2A: the reviewers approved, the supervisor never confirmed the admission close —
|
||||
not at the panel's begin, not at the final seal. The answer is ACCEPTED with the note,
|
||||
the rail is disclosed, no round was bought and the gate never became an owner follow-up."""
|
||||
f = full_loop
|
||||
_silent_supervisor(f)
|
||||
_nominate_then_keep(f, monkeypatch)
|
||||
result, _usage, trace = f.run()
|
||||
assert result == ANSWER and f.model_step == 2 and len(f.review_sends) == 1
|
||||
assert (trace["acceptance_decision"]["status"], trace["acceptance_decision"]["reason"]) == ("accepted", NOTE)
|
||||
assert trace["review_decision"]["admission_fence_available"] is False
|
||||
assert trace["review_decision"]["admission_released"] is False
|
||||
assert "TASK ACCEPTANCE WAIT" not in str(f.model_inputs)
|
||||
assert trace["review_decision"]["eligibility"] != "pending_owner_followup"
|
||||
rows = _unavailable_rows(f.ctx)
|
||||
assert rows and {(row["op"], row["outcome"]) for row in rows} == {("begin", "unknown")}
|
||||
record = _terminal_record(trace)
|
||||
assert outcome_phase(record, {}) == "done" and _completion_verdict(record, {}) == SENTENCE
|
||||
assert record["outcome_axes"]["objective"]["outcome_tier"] == OUTCOME_TIER_SOLVED
|
||||
|
||||
|
||||
def test_blocking_never_acked_begin_but_a_confirmed_final_seal_is_an_ordinary_clean_pass(full_loop, monkeypatch):
|
||||
"""The natural second chance: the panel ran on the disclosed rail, the final seal's own
|
||||
begin+end were answered. No note — the admission close WAS confirmed."""
|
||||
f = full_loop
|
||||
ends: list = []
|
||||
|
||||
def begin(**_kwargs):
|
||||
if not getattr(f.ctx, "_task_acceptance_reviewed", False):
|
||||
raise TimeoutError("no ack while the panel was owed")
|
||||
return {"token": "final-fence", "owner_message_generation": 0}
|
||||
|
||||
def end(**kwargs):
|
||||
ends.append(dict(kwargs))
|
||||
return {"ok": True, "status": "sealed"}
|
||||
|
||||
_silent_supervisor(f, begin=begin, end=end)
|
||||
_nominate_then_keep(f, monkeypatch)
|
||||
result, _usage, trace = f.run()
|
||||
assert result == ANSWER and f.model_step == 2
|
||||
assert trace["acceptance_decision"]["reason"] == "clean_pass"
|
||||
assert ends == [{"token": "final-fence", "outcome": "terminal", "expected_generation": 0}]
|
||||
assert f.ctx._task_acceptance_sealed_fence_token == "final-fence"
|
||||
assert _completion_verdict(_terminal_record(trace), {}) == ""
|
||||
|
||||
|
||||
def test_blocking_fail_verdict_is_never_laundered_by_the_note(full_loop, monkeypatch):
|
||||
"""The live one-cycle shape with a silent supervisor: FAIL on the first draft, the rewrite
|
||||
is refused by the cap. Today's outcome stands — `review_cycles_exhausted`, Failed, BLOCKED —
|
||||
the answer still goes out (no bought round) and the note is nowhere near it."""
|
||||
f = full_loop
|
||||
f.reviewer_verdict = "FAIL"
|
||||
monkeypatch.setenv("OUROBOROS_REVIEW_MAX_CYCLES", "1")
|
||||
_silent_supervisor(f)
|
||||
reauthored = ANSWER + " Budget: $12."
|
||||
|
||||
def main(_llm, messages, *_a, **_kw):
|
||||
f.model_inputs.append(copy.deepcopy(messages))
|
||||
f.model_step += 1
|
||||
if f.model_step == 1:
|
||||
return {"content": "", "tool_calls": [call("task_acceptance_review", {"claim": ANSWER}, "first-review")]}, 0.0
|
||||
if f.model_step == 2:
|
||||
assert f.entered.wait(5) and not f.release.is_set()
|
||||
observation = f.ctx._acceptance_observation
|
||||
return {"content": json.dumps({"delivery_control": "replace", "full_answer": reauthored,
|
||||
"acceptance_subject": {"owner_source_sha256": observation["owner_source_sha256"]}})}, 0.0
|
||||
assert f.model_step < 5, ("the fence bought model rounds", f.progress)
|
||||
return keep(f), 0.0
|
||||
|
||||
monkeypatch.setattr(loop, "call_llm_with_retry", main)
|
||||
result, _usage, trace = f.run()
|
||||
assert result == reauthored and len(f.review_sends) == 1
|
||||
assert trace["acceptance_decision"]["reason"] == "review_cycles_exhausted"
|
||||
assert trace["review_decision"]["admission_released"] is False
|
||||
record = _terminal_record(trace)
|
||||
assert outcome_phase(record, {}) == "error"
|
||||
assert record["outcome_axes"]["objective"]["outcome_tier"] == OUTCOME_TIER_BLOCKED
|
||||
assert SENTENCE not in _completion_verdict(record, {})
|
||||
|
||||
|
||||
def test_advisory_never_acked_delivers_with_the_loud_row_and_no_note(full_loop, monkeypatch):
|
||||
f = full_loop
|
||||
monkeypatch.setenv("OUROBOROS_REVIEW_ENFORCEMENT", "advisory")
|
||||
_silent_supervisor(f)
|
||||
_nominate_then_keep(f, monkeypatch)
|
||||
result, _usage, trace = f.run()
|
||||
assert result == ANSWER and f.model_step == 2
|
||||
assert trace["acceptance_decision"]["reason"] == "clean_pass" # advice on a soft install: no card note
|
||||
assert trace["review_decision"]["admission_fence_available"] is False
|
||||
assert trace["review_decision"]["admission_released"] is False
|
||||
assert _unavailable_rows(f.ctx), "the loud durable row is the soft install's record"
|
||||
|
||||
|
||||
def test_a_healthy_supervisor_seals_as_before_without_a_row_or_a_note(full_loop, monkeypatch):
|
||||
f = full_loop
|
||||
begins: list = []
|
||||
ends: list = []
|
||||
_silent_supervisor(
|
||||
f,
|
||||
begin=lambda **_kw: begins.append(1) or {"token": f"fence-{len(begins)}", "owner_message_generation": 0},
|
||||
end=lambda **kw: ends.append(kw["outcome"]) or {"ok": True, "status": "sealed" if kw["outcome"] == "terminal" else "released"},
|
||||
)
|
||||
_nominate_then_keep(f, monkeypatch)
|
||||
result, _usage, trace = f.run()
|
||||
assert result == ANSWER and f.model_step == 2
|
||||
assert trace["acceptance_decision"]["reason"] == "clean_pass"
|
||||
assert ends == ["revision", "terminal"] and f.ctx._task_acceptance_sealed_fence_token == "fence-2"
|
||||
assert trace["review_decision"]["admission_fence_available"] is True
|
||||
assert _unavailable_rows(f.ctx) == []
|
||||
|
||||
|
||||
def test_a_seal_applied_after_the_wait_expired_is_read_as_the_workers_own_seal(full_loop, monkeypatch):
|
||||
"""Main writes its answer, the panel passes and its end(terminal) loses the ack, but the
|
||||
supervisor applied it late: the final seal's own begin is refused because the root is
|
||||
already `sealed`. That is our seal, not a failure and not a gap — an ordinary clean pass,
|
||||
no note, no owner follow-up."""
|
||||
f = full_loop
|
||||
|
||||
def begin(**_kwargs):
|
||||
if getattr(f.ctx, "_task_acceptance_reviewed", False):
|
||||
raise RuntimeError("sealed") # the queue's typed refusal: already sealed for this root
|
||||
return {"token": "panel-fence", "owner_message_generation": 0}
|
||||
|
||||
def end(**_kwargs):
|
||||
raise TimeoutError("ack lost; the supervisor sealed the row after the wait")
|
||||
|
||||
_silent_supervisor(f, begin=begin, end=end)
|
||||
_nominate_then_keep(f, monkeypatch, nominate=False)
|
||||
result, _usage, trace = f.run()
|
||||
assert result == ANSWER and f.model_step == 2
|
||||
assert trace["acceptance_decision"]["reason"] == "clean_pass"
|
||||
assert trace["review_decision"].get("admission_released") is not False
|
||||
# (a silent `inspect` during the panel leaves its own gap row; the seal story is end -> begin)
|
||||
assert [(row["op"], row["outcome"], row["reason"]) for row in _unavailable_rows(f.ctx) if row["op"] != "inspect"] == [
|
||||
("end", "unknown", "TimeoutError"), ("begin", "refused", "sealed")]
|
||||
|
||||
|
||||
def _seal_context(tmp_path, owner_changed=False):
|
||||
ctx = SimpleNamespace(
|
||||
task_id="root-1", drive_root=tmp_path, task_metadata={"root_task_id": "root-1", "budget_drive_root": str(tmp_path)},
|
||||
_task_acceptance_fence_token=None, _task_acceptance_sealed_fence_token=None, _task_acceptance_fence_generation=None,
|
||||
_task_acceptance_queue_descendants=[], _task_acceptance_reviewed=True,
|
||||
_task_acceptance_fence_generation_mismatch=owner_changed, _owner_directives=[],
|
||||
)
|
||||
return SimpleNamespace(_ctx=ctx), SimpleNamespace(task_id="root-1")
|
||||
|
||||
|
||||
@pytest.mark.parametrize("enforcement,decision,answer,owner_changed,delivered,reason", [
|
||||
("blocking", "clean_pass", "unknown", False, True, NOTE),
|
||||
("blocking", "clean_pass_obligations_closed", "unknown", False, True, NOTE),
|
||||
("blocking", "clean_pass", "sealed", False, True, "clean_pass"), # the worker's own earlier seal
|
||||
("blocking", "clean_pass", "active", False, False, "owner_followup"), # another holder: today's revision path
|
||||
("blocking", "clean_pass", "unknown", True, False, "owner_followup"), # a locally seen owner change is real
|
||||
("advisory", "clean_pass", "unknown", False, True, "clean_pass"), # soft install: the row, not the note
|
||||
("blocking", "identical_acceptance_refused", "unknown", False, True, "identical_acceptance_refused"),
|
||||
])
|
||||
def test_the_final_seal_reads_the_queues_typed_answer(tmp_path, monkeypatch, enforcement, decision, answer, owner_changed, delivered, reason):
|
||||
from ouroboros.loop_delivery import _seal_admission_before_delivery
|
||||
|
||||
monkeypatch.setenv("OUROBOROS_REVIEW_ENFORCEMENT", enforcement)
|
||||
tools, limit_ctx = _seal_context(tmp_path, owner_changed)
|
||||
tools._ctx.begin_acceptance_fence = _gap if answer == "unknown" else (lambda **_kw: (_ for _ in ()).throw(RuntimeError(answer)))
|
||||
status = ACCEPTANCE_ACCEPTED if decision.startswith("clean_pass") else ACCEPTANCE_FINALIZED_UNACCEPTED
|
||||
llm_trace = {"review_decision": {"eligibility": "eligible"}, "review_runs": [],
|
||||
"acceptance_decision": {"status": status, "reason": decision, "source": "task_acceptance_review"}}
|
||||
assert _seal_admission_before_delivery(tools, limit_ctx, llm_trace) is delivered
|
||||
assert llm_trace["acceptance_decision"]["reason"] == reason
|
||||
assert llm_trace["acceptance_decision"]["status"] == (status if delivered else "revision_requested")
|
||||
if delivered and answer == "unknown":
|
||||
assert llm_trace["review_decision"]["admission_released"] is False
|
||||
|
||||
|
||||
def test_a_locally_seen_owner_change_survives_a_silent_end(tmp_path):
|
||||
"""`refused(generation_mismatch)` is a real owner follow-up even when the transport went
|
||||
silent: the local comparison decides the flag, not the missing ack. The quiet direction:
|
||||
an unchanged owner generation with the same silent end leaves the flag down."""
|
||||
from ouroboros.loop import _end_task_acceptance_fence
|
||||
|
||||
def ctx(generation):
|
||||
return SimpleNamespace(
|
||||
task_metadata={"root_task_id": "root-1"}, task_id="root-1", drive_root=tmp_path,
|
||||
_task_acceptance_fence_token="token-1", _task_acceptance_sealed_fence_token=None,
|
||||
_task_acceptance_fence_generation=0, _task_acceptance_queue_descendants=[],
|
||||
_task_acceptance_owner_generation=0, owner_message_admission_lock=threading.RLock(),
|
||||
owner_message_admission_agent=SimpleNamespace(_owner_message_generation=generation),
|
||||
end_acceptance_fence=_gap,
|
||||
)
|
||||
|
||||
changed, same = ctx(1), ctx(0)
|
||||
assert _end_task_acceptance_fence(changed, outcome="terminal").status == "unknown"
|
||||
assert changed._task_acceptance_fence_generation_mismatch is True
|
||||
assert _end_task_acceptance_fence(same, outcome="terminal").status == "unknown"
|
||||
assert same._task_acceptance_fence_generation_mismatch is False
|
||||
|
||||
|
||||
@pytest.mark.parametrize("status,finished", [("unknown", True), ("refused", False)])
|
||||
def test_an_advisory_author_finish_survives_a_silent_supervisor_but_not_a_refusal(tmp_path, monkeypatch, status, finished):
|
||||
from ouroboros import loop as loop_mod
|
||||
from ouroboros.loop_acceptance import FenceOutcome
|
||||
from ouroboros.loop_acceptance_review import _finish_advisory_author
|
||||
from ouroboros.loop_delivery import delivery_evidence_fingerprint
|
||||
from tests.test_acceptance_delivery import _acceptance_ctx
|
||||
|
||||
ctx = _acceptance_ctx(tmp_path)
|
||||
ctx.tools._ctx._owner_directives = [{"source": "initial_user", "content": "goal"}]
|
||||
binding = ctx.review_binding["binding_hash"]
|
||||
ctx.llm_trace["review_runs"] = [{"authority": "host_root", "feedback_delivered": True,
|
||||
"binding_hash": binding, "aggregate_signal": "FAIL"}]
|
||||
ctx.llm_trace["review_decision"] = {}
|
||||
ctx.llm_trace["acceptance_decision"] = {
|
||||
"agent_disposition": "accepted", "agent_rationale": "done",
|
||||
"agent_finish_intent": {"review_binding_hash": binding, "tool_count": 0, "owner_directives": 1,
|
||||
"evidence_fingerprint": delivery_evidence_fingerprint(ctx.tools._ctx, ctx.llm_trace)},
|
||||
}
|
||||
monkeypatch.setattr(loop_mod, "get_review_enforcement", lambda: "advisory")
|
||||
answer = FenceOutcome(status, "end", "TimeoutError" if status == "unknown" else "sealed", 0.8)
|
||||
monkeypatch.setattr(loop_mod, "_end_task_acceptance_fence", lambda *_a, **_k: answer)
|
||||
assert _finish_advisory_author(ctx) is True
|
||||
decision = ctx.llm_trace["acceptance_decision"]
|
||||
if finished:
|
||||
assert decision["reason"] == "author_finish" and ctx.tools._ctx._task_acceptance_reviewed is True
|
||||
assert ctx.llm_trace["review_decision"]["admission_released"] is False
|
||||
else:
|
||||
assert decision["reason"] == "owner_followup"
|
||||
|
||||
|
||||
@pytest.mark.parametrize("op,ack,reason", [
|
||||
("begin", {"ok": False, "status": "sealed", "error": "acceptance fence already sealed for root root-1"}, "sealed"),
|
||||
("begin", {"ok": False, "status": "error", "error": "invalid acceptance fence event"}, "invalid acceptance fence event"),
|
||||
("inspect", {"ok": True, "status": "released", "token": "t", "row_absent": True}, "released"),
|
||||
])
|
||||
def test_the_seam_refuses_with_the_rows_typed_status(tmp_path, op, ack, reason):
|
||||
"""A refusal carries the row's typed state (`sealed`, `active`, `released`) as its reason;
|
||||
only a malformed request keeps the error text. The worker reads the state, never the prose."""
|
||||
agent = _pooled_agent(tmp_path, stdqueue.Queue())
|
||||
agent.fence_transition = lambda **_kwargs: ack
|
||||
with pytest.raises(RuntimeError) as refused:
|
||||
if op == "begin":
|
||||
agent._begin_acceptance_fence(root_task_id="root-1", task_id="root-1")
|
||||
else:
|
||||
agent._inspect_acceptance_fence(token="t")
|
||||
assert str(refused.value) == reason
|
||||
|
|
|
|||
|
|
@ -163,7 +163,8 @@ def test_every_host_acceptance_writer_emits_a_canonical_status_and_typed_reason(
|
|||
]
|
||||
# Include the separate infrastructure-outcome handback; it requests an
|
||||
# author response without manufacturing a critic capsule or reviewer PASS.
|
||||
assert len(starts) == 22, f"writer inventory changed: {len(starts)} call sites"
|
||||
# ... and the final seal's `admission_close_unconfirmed` note (owner 2A) in the delivery leaf.
|
||||
assert len(starts) == 23, f"writer inventory changed: {len(starts)} call sites"
|
||||
allowed_status = {
|
||||
"ACCEPTANCE_ACCEPTED", "ACCEPTANCE_REVISION_REQUESTED",
|
||||
"ACCEPTANCE_FINALIZED_UNACCEPTED",
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue