diff --git a/docs/ARCHITECTURE.md b/docs/ARCHITECTURE.md index 4b39b7ff7..a2808013e 100644 --- a/docs/ARCHITECTURE.md +++ b/docs/ARCHITECTURE.md @@ -1117,6 +1117,10 @@ Provider death is the one forced rail that is NOT a best-effort completion: `_ha Host-enforced task acceptance is a root-owned completion coach, not the P3 commit gate. `off` disables it. Both `auto` and `required` review observable effects and typed deliverables/criteria. In `auto`, an explicit `task_acceptance_review` request also qualifies, including read-only research; queue membership alone does not. `required` retains its non-direct-root criterion. Ordinary conversation, exploration or cognitive-memory updates alone do not qualify in `auto`; no prose or tool-count classifier decides their meaning. Child reviews remain advisory evidence superseded by the root decision. +The root explicit call nominates the complete ready result. After the whole tool-result block, the host advances the same acceptance operation used by ordinary final delivery; early feedback does not seal the task. Main authors the effective criteria and can nominate material read observations through exact retained tool indices. Result bytes, those criteria and material effects form the paid subject; the full source packet and ingress generation remain separate forensic and ordering facts. Main acknowledges the source it actually received, so a status question can preserve a running review while a new criterion can request useful review of unchanged answer text. A released panel retains its exact request, resolved slot roster and existing physical-operation custody. Its settlement wakes the original Main through the mailbox; collection sends no new request, and waiting uses the existing continuation without fabricating a question or task id. Only actual final delivery seals ingress. Stop, missing custody and unfinished work retain their observed outcomes. + +A completed reviewer from an older plan wave is attached as a historical supplement through the existing locked task-result writer and exact producer CAS. It never rewrites the original actor verdict, aggregate, closure, author dispositions or current-wave pointer. The source ref, slot/operation binding and physical settlement facts remain available; a confirmed old dispatch can settle unknown historical cost once, without dispatching or charging another cycle. A terminal parent receives the supplement without starting another model turn. + Before an eligible panel is called, `supervisor/task_lifecycle.py` closes subtask admission under the queue lock and `task_status.find_child_tasks` proves the recursive subtree terminal and quiescent; revision reopens the fence, terminal or degraded completion seals it. Fence acknowledgement, subtree lookup, the timing telemetry stream, and the packet's mutation-attribution read use the canonical `budget_drive_root`; the one-shot `state/acceptance_fence_acks/` IPC sidecar is not a lifecycle authority. The reviewer packet preserves verbatim owner directives, the full task contract and criteria, canonical deliverable identity, terminal child state, verification receipts, artifact references, the host-attested lifecycle facts (review status, staleness, readiness, enablement) of every skill the task touched — visibility only, never a gate — and an explicit omissions manifest; a required component that cannot be assembled makes the affected actor `DEGRADED`, never a silently smaller prompt. The packet is SIZED against the review quorum's real windows — the same `reviewer_window`/`review_synthesis.quorum_input_token_limit` seam the triad and plan review use — resolved once per task and memoised on the acceptance context, so the packet bytes cannot drift between the binding build and the staleness rebuild. The same cached per-slot caps drive the pre-send fit check; dispatch never recalibrates them. Non-core sections then shed through a DISCLOSED ladder: the predecessor authority envelope first, then the trajectory tail and its results, artifact previews, agent-supplied evidence, and last a diff preview that keeps the durable `repo_diff_source_ref`. Each shed is a row in `omissions_manifest`. A slot whose own window cannot hold the rendered prompt is a typed `$0 not_dispatched` row while the rest of the panel reviews; a packet that still overflows after every shed stamps `__immutable_core_overflow__` naming the oversized sections and refuses the panel without spending anything. `__unresolved_partial_artifacts__` withholds the panel's packet rows only for a tool result whose exact source is genuinely `source_unavailable` (a retrieving row reads the exact source itself) — a budget shed with a durable, actor-resolvable source ref is an omission, never an unresolved partial. @@ -1543,7 +1547,7 @@ Read-only children can read/list the existing project-scoped knowledge store wit **Scheduling.** `schedule_subagent` requires `subagent_id`, a focused `objective`, and `expected_output`; the remaining public fields describe child-local context, constraints, memory, capability needs, write surface, narrower deadline, delegation budget, and acceptance claims. There is no model-visible lane/executor axis and no public `effort` override: the selected row is the complete execution choice, lineage/bounds/route/budget are host-derived, and omission never inherits the parent's acceptance claims. `subagent_runtime.select_subagent_snapshot` copies an immutable snapshot of the exact enabled row into the child task; an `api_model` row becomes an ordinary recursive API child on that exact model/effort, an `agent_session` row an ordinary recursive Ouroboros nanny on that exact external route. The cash side of a burst — each sibling launched before the first sibling's first response pays its own prefix write on cache-write-priced routes — is disclosed to the mind in the tool description as an affordance and deliberately not scheduled by the host. The old lane/executor resolver serves only old durable records; for those historical lane envelopes, `schedule_subagent` reports the requested lane only — effective facts remain on the dispatched child record — and a task carrying `configured_subagent` goes straight to `subagent_runtime`, so legacy policy cannot reinterpret an active selection. -**Waiting on children.** `wait_task`/`get_task_result` return the full single-child handoff (verification receipts red/masked-first, exact omitted count); `wait_tasks` stays batch-compact: `task_id, status, child_result_sha256, outcome_axes, result, terminal_host_notice when present, trace_summary, capability_delta when the child has something to disclose, duplicate_of`, plus the nullable cost-finality pair `accounted_upper_bound_usd`/`cost_final` and, when the child's envelope carries one, `execution_evidence` (§11.1). The retired `cost_usd` spelling is tolerated only when reading stored rows, never emitted. Both use `task_status.SETTLED_STATUSES`; a pending cancellation is the typed `cancel_state: "pending"` projection, never completion. A batch wait that expires with children still running discloses what it could not finish: the typed `wait_expired_with_live_children` block names the live child ids, the window that was requested and the clamp ceiling the schema already states. Facts only, with no advisory text in the payload and no host floor on the next window, because how long to wait is the mind's call (BIBLE P13); an id this tree never minted stays an `unknown_task_id` and is never counted as a live child, and the pinned per-child field list is unchanged. +**Waiting on children.** `wait_task`/`get_task_result` return the full single-child handoff (verification receipts red/masked-first, exact omitted count); `wait_tasks` stays batch-compact: `task_id, status, child_result_sha256, outcome_axes, result, terminal_host_notice when present, trace_summary, capability_delta when the child has something to disclose, duplicate_of`, plus the nullable cost-finality pair `accounted_upper_bound_usd`/`cost_final` and, when the child's envelope carries one, `execution_evidence` (§11.1). The retired `cost_usd` spelling is tolerated only when reading stored rows, never emitted. Both use `task_status.SETTLED_STATUSES`; a pending cancellation is the typed `cancel_state: "pending"` projection, never completion. A batch wait that expires with children still running discloses what it could not finish: the typed `wait_expired_with_live_children` block names the live child ids, the window that was requested and the clamp ceiling the schema already states. Facts only, with no advisory text in the payload and no host floor on the next window, because how long to wait is the mind's call (BIBLE P13); an id this tree never minted stays an `unknown_task_id` and is never counted as a live child, and the pinned per-child field list is unchanged. An optional `known_result_sha256` on the single-child reads, or `known_result_sha256_by_task` on batch wait, compares the existing join-ledger semantic result identity. An exact match omits repeated result/trace text and returns `result_unchanged` with an unconditional source call; current status, costs, outcome and custody facts remain. Missing or changed conditions return the normal full handoff. No persistent seen-state is inferred, and callers can always read the full result again after compaction. **What a delegated run costs.** Claudexor reports the amount in `summary.spendUsd` and its exactness in `summary.spendEstimated`; `delegate_custody.disclosed_spend` is the single reader of the pair, so the ledger row and the payload the nanny relays cannot tell different stories. Runs ask `authPreference: subscription` explicitly, because the engine default falls back to a paid key invisibly. Four cases: diff --git a/docs/CHECKLISTS.md b/docs/CHECKLISTS.md index 63ff163a8..04b23fe66 100644 --- a/docs/CHECKLISTS.md +++ b/docs/CHECKLISTS.md @@ -6,6 +6,15 @@ multi-model review prompt. When a new reviewable concern appears, add it here — not in prompts or docs. +**Application follows BIBLE P0/P3.** Review findings and failures are independent +facts in every mode. In Cyber Pro they inform Ouroboros and never prohibit an +action or require permission; the agent may configure its own subsequent work. +Configured enforcement remains recorded as selected, and an author decision does +not rewrite FAIL, pending, missing source or an unperformed effect as PASS. +Enforcement requirements below describe ordinary modes; Cyber applies the same +checks as advice. This product rule does not replace an external developer's +explicit work-order review obligations. + --- ## Advisory Pre-Review Workflow @@ -625,7 +634,7 @@ and do not return `PASS` for an item that also has a `FAIL` — the concrete `blockers` are executable by operator choice. This changes `executable_review` only; it does not rewrite the verdict, suppress findings, or change `skill_review_status` semantics. - - `pending` is never executable. A stale critic verdict does not authorize bytes; + - Outside Cyber Pro, `pending` is not executable. A stale critic verdict does not authorize bytes; under Advisory a separate current author acceptance may admit the payload after deterministic preflight. Blocking still requires fresh critic evidence. - Review state stores findings and computes the verdict at load time. Agents @@ -633,8 +642,8 @@ and do not return `PASS` for an item that also has a `FAIL` — the concrete not the raw status string, when deciding whether the skill is runnable. - A deterministic `skill_preflight` FAIL is a structural gate failure, not an LLM verdict: it persists and aggregates to `pending`, which is non-executable under - EVERY enforcement mode (advisory included) and in every readiness/execution - caller — the strongest fail-closed outcome, stronger than an overridable blocker. + ordinary enforcement mode (advisory included). Cyber keeps the failed check + and pending verdict visible while leaving the execution decision to Ouroboros. - Hard trust-boundary items are blocker findings on any FAIL regardless of reviewer-supplied severity: `manifest_schema`, `permissions_honesty`, `no_repo_mutation`, `path_confinement`, diff --git a/docs/DEVELOPMENT.md b/docs/DEVELOPMENT.md index d19e99756..d4a9aacf6 100644 --- a/docs/DEVELOPMENT.md +++ b/docs/DEVELOPMENT.md @@ -2609,9 +2609,11 @@ by "Provider Independence" above. Call-site imperatives: `ouroboros/loop_delivery.py`). A FORCED finalization resolves an armed control purely and without retry: valid keep/replace is honored, anything malformed preserves the retained candidate with a typed degraded reason, - and protocol JSON never reaches chat or the durable result. Owner - messages, tool effects, child results, and verification receipts advance - the evidence revision and require fresh delivery/acceptance binding; + and protocol JSON never reaches chat or the durable result. Main distinguishes + consumed owner source from changed requirements. Effective criteria and + material effects, including nominated read observations, define the reviewed + subject; ingress generations preserve unread-message ordering. Status text, + narration and a changed working view do not themselves buy another review; finalize task-scoped service outputs/errors before host acceptance. The control must not bypass verification, acceptance, safety, skill finalization, deadline, child handoff, the unconditional `FINAL ANSWER:` @@ -2635,11 +2637,14 @@ by "Provider Independence" above. Call-site imperatives: - Host task acceptance is root-only; eligibility uses structured facts (`outcomes.turn_has_reviewable_effects` plus a typed deliverable/criterion), never keywords (BIBLE P3/P5). The agent-callable - `task_acceptance_review` stores evidence but makes zero reviewer calls and - returns `deferred_to_host_acceptance`, `authoritative=false`. Before root - acceptance, atomically fence new descendants under the queue lock and - prove recursive subtree quiescence from the task-status SSOT; a revision - must explicitly reopen the fence, and terminal/degraded outcomes seal it. + `task_acceptance_review` records the full result nomination and returns + `deferred_to_host_acceptance`, `authoritative=false`. After the complete + tool-result block, the host advances the same operation as final delivery. + Freeze its request and resolved roster; use existing review custody and + mailbox continuation for pending work and free collection. The worker never + writes Main's live candidate or author decision. Early settlement does not + seal task ingress; actual final delivery does. Keep subtree/status facts + separate from reviewer findings and Cyber's authority under BIBLE P0. - Delivery-control JSON applies only to a final response with no tool calls. Retaining an answer leaves tools available for further work; changed evidence still requires the existing complete replacement. A requested file or diff diff --git a/ouroboros/config.py b/ouroboros/config.py index 68e979cea..a88028090 100644 --- a/ouroboros/config.py +++ b/ouroboros/config.py @@ -20,7 +20,7 @@ import time from typing import Any, Optional, Sequence # noqa: F401 from ouroboros.context_mode_compat import ( - normalize_and_persist_context_mode_compat, normalize_context_mode, owner_declared_low, + VALID_CONTEXT_MODES, normalize_and_persist_context_mode_compat, normalize_context_mode, owner_declared_low, ) from ouroboros.platform_layer import pid_lock_acquire as _compat_pid_lock_acquire, pid_lock_release as _compat_pid_lock_release from ouroboros.provider_models import compute_direct_review_models_fallback, fallback_candidate_targets, local_only_review_route_env, migrate_model_value, resolve_model_target, review_model_uses_local as review_model_uses_local # noqa: F401 @@ -441,7 +441,7 @@ def get_safety_mode() -> str: def get_context_mode() -> str: - """The EFFECTIVE working-context mode (low | max) used by context sizing: owner selection or + """The EFFECTIVE working-context mode (nano | low | max) used by context sizing: owner selection or an explicitly forwarded benchmark/operator value. The P3 scope gate reads get_owner_context_mode instead so a bare env Low cannot author owner intent. No boot-pin: hot-applies on the next task; the key is dropped from the agent-reachable /api/settings POST (P1).""" @@ -454,8 +454,9 @@ def get_owner_context_mode() -> str: auto-Low is retired, but a bare forwarded env ``low`` still lacks owner provenance and keeps P3 at Max; only explicit ``low`` + tombstone ``false`` means owner Low. Raw persisted legacy ambiguity is normalized before env projection, so this matters only for env-only runs.""" - if get_context_mode() != "low": - return "max" + mode = get_context_mode() + if mode != "low": + return mode return "low" if owner_declared_low(runtime_setting("OUROBOROS_CONTEXT_MODE_AUTO_LOW", "")) else "max" @@ -483,9 +484,9 @@ def _guard_context_mode_lowering(settings: dict, *, allow_context_lowering: bool return previous_mode = normalize_context_mode(_settings_file_value("OUROBOROS_CONTEXT_MODE", "max")) next_mode = normalize_context_mode(settings.get("OUROBOROS_CONTEXT_MODE", previous_mode)) - if previous_mode == "max" and next_mode == "low" and not allow_context_lowering: + if VALID_CONTEXT_MODES.index(next_mode) < VALID_CONTEXT_MODES.index(previous_mode) and not allow_context_lowering: raise PermissionError( - "OUROBOROS_CONTEXT_MODE lowering refused: 'max' -> 'low'. " + f"OUROBOROS_CONTEXT_MODE lowering refused: {previous_mode!r} -> {next_mode!r}. " "Context mode is owner-controlled — use the dedicated owner endpoint/UI/CLI." ) if allow_context_lowering or "OUROBOROS_CONTEXT_MODE_AUTO_LOW" not in settings: diff --git a/ouroboros/context_mode_compat.py b/ouroboros/context_mode_compat.py index 6bc519886..95dde57a9 100644 --- a/ouroboros/context_mode_compat.py +++ b/ouroboros/context_mode_compat.py @@ -10,7 +10,7 @@ from ouroboros.utils import write_text_atomic log = logging.getLogger(__name__) -VALID_CONTEXT_MODES = ("low", "max") +VALID_CONTEXT_MODES = ("nano", "low", "max") _MIGRATION_WARNED_PATHS: set[str] = set() @@ -40,7 +40,7 @@ def normalize_context_mode_compat( raw_mode = normalize_context_mode(normalized.get(mode_key)) marker_is_false = owner_declared_low(normalized.get(marker_key)) ambiguous_low = raw_mode == "low" and not marker_is_false - normalized[mode_key] = "low" if raw_mode == "low" and marker_is_false else "max" + normalized[mode_key] = raw_mode if raw_mode == "nano" or (raw_mode == "low" and marker_is_false) else "max" normalized[marker_key] = "false" warning_key = str((settings_path or Path("")).resolve(strict=False)) if ambiguous_low and warn_ambiguous and warning_key not in _MIGRATION_WARNED_PATHS: diff --git a/ouroboros/gateway/settings.py b/ouroboros/gateway/settings.py index e46948886..57bb6bff0 100644 --- a/ouroboros/gateway/settings.py +++ b/ouroboros/gateway/settings.py @@ -737,7 +737,7 @@ def _review_capability_notices(settings: Dict[str, Any]) -> list: @owner_write_guard async def api_owner_context_mode(request: Request) -> JSONResponse: - """Persist the owner-selected context mode (low/max). + """Persist the owner-selected working context mode. Ordinary modes use this owner path; Cyber can also author a generic save. The choice applies to subsequent tasks, so no restart is required. @@ -757,15 +757,15 @@ def _api_owner_context_mode_sync(request: Request, body: Any) -> JSONResponse: from ouroboros.context_mode_compat import VALID_CONTEXT_MODES if raw_mode not in set(VALID_CONTEXT_MODES): - return unsaved_error("'mode' must be one of: low, max", 400) + return unsaved_error("'mode' must be one of: " + ", ".join(VALID_CONTEXT_MODES), 400) next_mode = _config.normalize_context_mode(raw_mode) digest = settings_document_digest() previous_mode = _config.get_owner_context_mode() cyber = runtime_mode_at_least(_config.get_runtime_mode(), "cyber_pro") - if not cyber and previous_mode == "max" and next_mode == "low" and _has_running_agent_tasks(): + if not cyber and VALID_CONTEXT_MODES.index(next_mode) < VALID_CONTEXT_MODES.index(previous_mode) and _has_running_agent_tasks(): return unsaved_error( "Context mode can only be lowered while Ouroboros is idle. " - "Wait until no queued or running work remains, then switch Low/Max.", + "Wait until no queued or running work remains, then choose the working context.", 409, ) @@ -783,10 +783,10 @@ def _api_owner_context_mode_sync(request: Request, body: Any) -> JSONResponse: # digest alone cannot attest idleness. Cyber can choose the next mode # during work; existing task snapshots retain their original settings. previous_mode = _config.get_owner_context_mode() - if not cyber and previous_mode == "max" and next_mode == "low" and _has_running_agent_tasks(): + if not cyber and VALID_CONTEXT_MODES.index(next_mode) < VALID_CONTEXT_MODES.index(previous_mode) and _has_running_agent_tasks(): return unsaved_error( "Context mode can only be lowered while Ouroboros is idle. " - "Wait until no queued or running work remains, then switch Low/Max.", + "Wait until no queued or running work remains, then choose the working context.", 409, ) # This endpoint IS the author of both keys, so they persist even at the shipped default. diff --git a/ouroboros/loop_forced_finalization.py b/ouroboros/loop_forced_finalization.py index 2ef62a3aa..5a2834345 100644 --- a/ouroboros/loop_forced_finalization.py +++ b/ouroboros/loop_forced_finalization.py @@ -501,9 +501,22 @@ _FORCED_BEST_EFFORT_TAIL = ( def _prepare_forced_prompt( ctx: _RoundLimitContext, prompt: str, llm_trace: Dict[str, Any], ) -> str: + _loop()._drain_forced_owner_directives(ctx, llm_trace) _loop()._finalize_forced_services(ctx, llm_trace) tools_ctx = getattr(getattr(ctx, "tools", None), "_ctx", None) - return prompt + _loop()._forced_delegation_note(tools_ctx, llm_trace) + return prompt + _loop()._forced_delegation_note(tools_ctx, llm_trace) + _forced_subject_prompt(ctx, llm_trace) + + +def _forced_subject_prompt(ctx: _RoundLimitContext, llm_trace: Dict[str, Any]) -> str: + """Capture the source before pricing/sending, never after a reply arrives.""" + from ouroboros.loop_acceptance import capture_acceptance_observation, acceptance_observation_prompt + + tools_ctx = getattr(getattr(ctx, "tools", None), "_ctx", None) + if tools_ctx is None: + return "" + observed = capture_acceptance_observation(tools_ctx, llm_trace, ctx.incoming_messages) + rendered = acceptance_observation_prompt(tools_ctx, observed) + return "\n\n" + rendered if rendered else "" def _finalize_forced_services( @@ -851,6 +864,7 @@ def _forced_fallback_result( def _resolve_forced_delivery_control( tools_ctx: Any, extracted: str, + *, ctx: Optional[_RoundLimitContext] = None, llm_trace: Optional[Dict[str, Any]] = None, ) -> Tuple[str, str, bool, bool]: """Resolve forced control; returns text, degradation, retained, replaced.""" if tools_ctx is None or not extracted: @@ -867,6 +881,29 @@ def _resolve_forced_delivery_control( ) if consumed: tools_ctx._delivery_control_required = False + from ouroboros.loop_delivery import _parse_delivery_control_body, apply_delivery_subject_decision + + parsed, duplicate, embedded = _parse_delivery_control_body(extracted) + if (not duplicate and not embedded and isinstance(parsed, dict) + and "acceptance_subject" in parsed and not degraded): + applied, error = ( + apply_delivery_subject_decision(ctx.tools, ctx, llm_trace, parsed["acceptance_subject"]) + if ctx is not None and llm_trace is not None else + (False, "forced subject has no current source observation context") + ) + if llm_trace is not None: + llm_trace["forced_acceptance_subject"] = {"applied": applied, "reason": error} + if not applied: + degraded = True + if ctx is not None and llm_trace is not None and isinstance(candidate, _loop().DeliveryCandidate): + candidate.acceptance_binding = _loop()._forced_unaccepted_binding( + ctx.tools, candidate, REASON_DELIVERY_CONTROL_DEGRADED, + ) + tools_ctx._task_acceptance_reviewed = False + _loop()._set_acceptance_decision(llm_trace, { + "status": "finalized_unaccepted", "reason": REASON_DELIVERY_CONTROL_DEGRADED, + "source": "forced_acceptance_subject", "rationale": error, + }) return ( resolved, REASON_DELIVERY_CONTROL_DEGRADED if degraded else "", @@ -929,7 +966,7 @@ def _forced_final_answer( reason_code, source="provider_outcome_unknown_no_resend", ) - if attempt == 1: + if single_semantic_turn or attempt == 1: return _loop()._forced_fallback_result( ctx, llm_trace, @@ -941,7 +978,8 @@ def _forced_final_answer( _loop()._finalize_forced_services(ctx, llm_trace) _loop()._append_or_merge_user_message( ctx.messages, - "[FORCED_OWNER_REFRESH] Answer all current directives; ignore the stale draft.", + "[FORCED_OWNER_REFRESH] Answer all current directives; ignore the stale draft." + + _forced_subject_prompt(ctx, llm_trace), ) # Control resolution runs BEFORE the incomplete branch: a retained candidate @@ -949,7 +987,7 @@ def _forced_final_answer( # and a stale-evidence retention keeps its own reason (#447/issue-449). incomplete = bool(extracted) and forced_response_is_incomplete(response_meta) extracted, control_degraded, retained, replaced = _resolve_forced_delivery_control( - tools_ctx, extracted, + tools_ctx, extracted, ctx=ctx, llm_trace=llm_trace, ) current = _loop()._current_delivery_candidate(ctx, llm_trace) if retained and current is None: diff --git a/ouroboros/review_state.py b/ouroboros/review_state.py index f88679e7d..c392b8f92 100644 --- a/ouroboros/review_state.py +++ b/ouroboros/review_state.py @@ -439,13 +439,18 @@ def advisory_commit_ready( ) -> bool: """SSOT for every ``repo_commit_ready`` projection (H5, capinv-447). - Mirrors the real advisory gate: fresh/bypassed/skipped coverage, or a + Mirrors the real advisory gate: Cyber retains action authority; otherwise + fresh/bypassed/skipped coverage, or a typed technical failure permitted under owner-selected advisory enforcement. ``matching_run`` is supplied only after the caller matches current repo/hash; permission never changes its failure status or makes it fresh. Obligations and debt block only under blocking enforcement. Triad, scope, custody and every other commit requirement remain independent. """ + from ouroboros.tools.review_helpers import review_enforcement_blocks + + if not review_enforcement_blocks("blocking"): + return True if not effectively_fresh: from ouroboros.config import get_review_enforcement from ouroboros.tools.commit_gate import review_failure_is_technical diff --git a/ouroboros/review_state_custody.py b/ouroboros/review_state_custody.py index 583f5f77d..6563f456f 100644 --- a/ouroboros/review_state_custody.py +++ b/ouroboros/review_state_custody.py @@ -93,7 +93,11 @@ def checkpoint_pending_review_invocation( The existing advisory-review state remains the only ledger. This narrow locked patch is deliberately stricter than a whole-attempt merge: triad and scope start concurrently, so each may update only its exact reserved row - and may never overwrite the other surface's token. + and may never overwrite the other surface's token. A commit or review-only + action may already have finished while its reserved reviewer is starting. + Its logical status does not revoke that physical operation's custody; the + paid attempt, retry key and exact in-flight slot below are the authority. + This checkpoint records no verdict and never changes the commit status. """ expected = { "review_retry_key": str(review_retry_key or ""), @@ -116,7 +120,6 @@ def checkpoint_pending_review_invocation( ) if ( current is None - or current.status != "reviewing" or not current.paid or current.review_retry_key != expected["review_retry_key"] ): diff --git a/ouroboros/skill_loader.py b/ouroboros/skill_loader.py index 0e674ebf3..f01b50999 100644 --- a/ouroboros/skill_loader.py +++ b/ouroboros/skill_loader.py @@ -319,11 +319,13 @@ def _iter_payload_files( """Return files hashed for review freshness. The hash covers every regular runtime-reachable file under ``skill_dir`` - except metadata/cache/sensitive paths, lifecycle control files + except metadata/cache paths, lifecycle control files (``HASH_EXEMPT_CONTROL_FILENAMES``), and symlink escapes. Manifest entry points are re-added only when confined, keeping executable and reviewed surfaces aligned. ``include_control_files=True`` reproduces the legacy pre-v6.31 hash (control files included) for one-shot state migration. + Sensitive-looking filenames refuse ordinary loading; Cyber includes them + in the same byte hash and review pack rather than silently omitting them. """ out: List[pathlib.Path] = [] resolved_root = skill_dir.resolve() @@ -354,6 +356,10 @@ def _iter_payload_files( _SENSITIVE_EXTENSIONS, _SENSITIVE_NAMES, ) + from ouroboros.config import get_runtime_mode + from ouroboros.runtime_mode_policy import runtime_mode_at_least + + cyber = runtime_mode_at_least(get_runtime_mode(), "cyber_pro") def _is_sensitive(path: pathlib.Path) -> bool: lowered = path.name.lower() @@ -385,7 +391,7 @@ def _iter_payload_files( and resolved_root.parent.name == "native" ): continue - if _is_sensitive(path): + if _is_sensitive(path) and not cyber: # Fail closed: a reviewed skill could still read a skipped # credential-shaped file at runtime. raise SkillPayloadUnreadable( diff --git a/ouroboros/skill_review_packs.py b/ouroboros/skill_review_packs.py index 7c8ec57f2..6c644010f 100644 --- a/ouroboros/skill_review_packs.py +++ b/ouroboros/skill_review_packs.py @@ -83,9 +83,9 @@ def _read_skill_file( path: pathlib.Path, *, relpath: str = "" ) -> tuple[Optional[str], bytes, Optional[Dict[str, Any]]]: """Read one skill file: ``(text, sha256_digest, descriptor)`` — exactly one set. - Loadable executables (CONTENT magic bytes, never filename) hard-block review; - WebAssembly (``WASM_MAGIC``, even when its bytes decode as UTF-8) and other - non-UTF-8 files yield a typed descriptor instead of raw bytes.""" + Native executable magic blocks ordinary review; Cyber retains these bytes + through a descriptor. Binary formats never become text merely because they + decode as UTF-8. Unreadable source remains an actual read failure.""" try: data = path.read_bytes() except OSError as exc: @@ -98,11 +98,18 @@ def _read_skill_file( text = None kind = executable_magic_kind(data, is_utf8_text=text is not None) if kind: - raise _SkillBinaryPayload(rel, len(data), kind) + from ouroboros.config import get_runtime_mode + from ouroboros.runtime_mode_policy import runtime_mode_at_least + + if not runtime_mode_at_least(get_runtime_mode(), "cyber_pro"): + raise _SkillBinaryPayload(rel, len(data), kind) digest = hashlib.sha256(data).digest() - if text is not None and not data.startswith(WASM_MAGIC): + if text is not None and not kind and not data.startswith(WASM_MAGIC): return text, digest, None - return None, digest, binary_file_descriptor(rel, data, filename=path.name) + descriptor = binary_file_descriptor(rel, data, filename=path.name) + if kind: + descriptor["format_from_magic"] = kind + return None, digest, descriptor def _build_skill_file_packs( diff --git a/ouroboros/skill_review_status.py b/ouroboros/skill_review_status.py index a5f447127..ff8088c70 100644 --- a/ouroboros/skill_review_status.py +++ b/ouroboros/skill_review_status.py @@ -200,14 +200,14 @@ def skill_review_gate( ) -> Dict[str, Any]: """Structured, agent-facing explanation of whether a review is executable. - A current author acceptance admits changed bytes only in Advisory; stale - still describes the original reviewer evidence, never the author hash. + Advisory author acceptance binds changed bytes; Cyber retains judgment even + without a prior verdict. Stale describes reviewer evidence, not author intent. Author fields are optional and appear only with a valid author disposition; callers without one retain the frozen gate key set. Deterministic hard-gate failures (e.g. skill_preflight) are persisted as STATUS_PENDING by `_run_deterministic_preflight`, so they are non-executable - here under every enforcement mode without needing per-caller findings — only + here outside Cyber without needing per-caller findings — only LLM blocker verdicts are overridable by advisory enforcement. ``findings`` is optional: a caller that has the persisted findings gets a @@ -224,8 +224,7 @@ def skill_review_gate( from ouroboros.review_records import validate_author_disposition author = validate_author_disposition(author_disposition) - author_current = bool(current_hash and author and author["subject_hash"] == current_hash - and author["enforcement"] == "advisory") + author_current = bool(current_hash and author and author["subject_hash"] == current_hash) raw_status = normalize_skill_review_status(status) if enforcement is None: try: @@ -234,11 +233,19 @@ def skill_review_gate( except Exception: enforcement = "blocking" enforcement = str(enforcement or "blocking").lower() - if raw_status == STATUS_PENDING: + from ouroboros.tools.review_helpers import review_enforcement_blocks + + cyber = not review_enforcement_blocks("blocking") + if cyber: + executable = True + reason = "cyber_authority" + summary = ("Cyber Pro permits acting on this payload by Ouroboros's judgment; " + "review status, staleness and failures remain independent evidence, not a PASS.") + elif raw_status == STATUS_PENDING: executable = False reason = "review_pending" summary = "Review is pending or did not produce an executable verdict." - elif author_current and enforcement == "advisory": + elif author_current and author["enforcement"] == "advisory" and enforcement == "advisory": executable = True reason = "author_accepted_advisory" summary = "The author accepted the current payload under Advisory; the original reviewer verdict and hash are unchanged." @@ -270,7 +277,7 @@ def skill_review_gate( return { "status": raw_status or STATUS_PENDING, "stale": bool(stale), - **({"author_accepted": reason == "author_accepted_advisory", + **({"author_accepted": reason == "author_accepted_advisory" or (cyber and author_current), "author_disposition": author} if author else {}), "executable_review": bool(executable), "blocking_reason": reason, diff --git a/ouroboros/task_results.py b/ouroboros/task_results.py index fa46fdcd7..4dbaad8bc 100644 --- a/ouroboros/task_results.py +++ b/ouroboros/task_results.py @@ -1134,19 +1134,15 @@ def plan_review_gate_projection( *, hard_rail: str = "", ) -> Dict[str, Any]: - """Project one plan-review finalization decision from existing authority. + """Project finalization permission without changing the durable review facts. - ``plan_review_state`` is the durable SSOT; the ``current_attempt`` pointer keeps a - newer fingerprint from falling back to an older closed wave. Statuses: ``closed`` - (allow) · ``rail_degraded`` (a task-wide rail released the hold — allow) · - ``advisory_open`` (advisory enforcement proceeds under loud disclosure) · - ``cycles_exhausted`` (the shared cap is spent on an OPEN wave: finalization is - released so the task can terminalize honestly as blocked — owner D27 — while - the wave itself stays open) · ``open`` / ``unavailable`` / ``pending`` / - ``legacy_open_requires_resubmission`` (blocking hold; EXCEPT an ``open`` wave - whose ``quorum_unreachable`` typed fact holds — B2b — which releases - finalization the same honest-blocked way while staying open) · ``absent``. Accepts a v2 - state, a loaded v1 wrapper, or a raw v1 record (read-only projection).""" + The current-attempt pointer prevents an older closed wave authorizing new + work. In ordinary Blocking, open/unavailable/pending/legacy-open reviews hold + finalization; spent cycles (D27), unreachable quorum (B2b) or a hard rail + release it for an honest blocked outcome. Advisory releases an open review. + Cyber retains judgment even with missing evidence: allow never implies closed + or PASS. Accepts v2 state, a v1 wrapper or raw v1 as a read-only projection. + """ policy = "blocking" if str(enforcement or "").lower() == "blocking" else "advisory" control: Dict[str, Any] = {} attempted = False @@ -1202,8 +1198,13 @@ def plan_review_gate_projection( status = str(control.get("status") or "unavailable") closed = bool(control.get("closed")) + from ouroboros.tools.review_helpers import review_enforcement_blocks + + cyber = not review_enforcement_blocks("blocking") if status == "closed" and closed: gate_status, allow = "closed", True + elif cyber: + gate_status, allow = "advisory_open", True elif hard_rail or status == "rail_degraded": gate_status, allow = "rail_degraded", True elif policy == "advisory" and status in { @@ -1225,6 +1226,7 @@ def plan_review_gate_projection( gate_status, allow = status, False return { "enforcement": policy, + **({"decision_authority": "cyber_pro", "review_status": status} if cyber else {}), "status": gate_status, "allow": allow, "attempted": attempted, diff --git a/ouroboros/tools/claude_advisory_review.py b/ouroboros/tools/claude_advisory_review.py index 1d19a04d3..29e76825a 100644 --- a/ouroboros/tools/claude_advisory_review.py +++ b/ouroboros/tools/claude_advisory_review.py @@ -769,6 +769,16 @@ def _next_step_guidance(latest: Optional["AdvisoryRunRecord"], state: "AdvisoryR one unbindable case stays as before: an uncomputable current hash cannot establish a mismatch either way. """ + from ouroboros.tools.review_helpers import review_enforcement_blocks + + if not review_enforcement_blocks("blocking"): + return ( + f"Cyber Pro: preflight status={getattr(latest, 'status', 'missing')}; " + f"stale={bool(stale_from_edit or not effective_is_fresh)}. " + "Ouroboros decides whether to continue or request more feedback. " + "Original findings, missing evidence and pending operations remain recorded; this is not a PASS." + ) + def _debt_hint() -> str: parts = [] if open_obs: diff --git a/ouroboros/tools/commit_gate.py b/ouroboros/tools/commit_gate.py index fde7edd2b..dbbe9f4d3 100644 --- a/ouroboros/tools/commit_gate.py +++ b/ouroboros/tools/commit_gate.py @@ -18,6 +18,7 @@ from ouroboros.review_state import ( infer_review_phase, ) from ouroboros.tools.registry import ToolContext +from ouroboros.tools.review_helpers import review_enforcement_blocks from ouroboros.utils import ( truncate_review_artifact as _truncate_review_reason, ) @@ -624,11 +625,12 @@ def _record_commit_attempt( if author_disposition is not None: subject = pre_review_fingerprint or str(getattr(existing, "pre_review_fingerprint", "") or "") author_record = validate_author_disposition(author_disposition, subject_hash=subject) or {} - if (not subject or not getattr(existing, "paid", False) + cyber = not review_enforcement_blocks("blocking") + if (not subject or (not cyber and not getattr(existing, "paid", False)) or subject != getattr(existing, "pre_review_fingerprint", "") or (post_review_fingerprint and post_review_fingerprint != subject) - or get_review_enforcement() != "advisory" - or author_record.get("enforcement") != "advisory"): + or review_enforcement_blocks(get_review_enforcement()) + or (not cyber and author_record.get("enforcement") != "advisory")): author_record = {} attempt = CommitAttemptRecord( ts=_utc_now(), @@ -708,6 +710,8 @@ def _record_commit_attempt( getattr(existing, "review_owner_pid", 0) or 0 ), ) + if status != "reviewing" and "late_result_pending" not in legacy_kwargs and not review_enforcement_blocks("blocking"): + attempt.late_result_pending = bool(getattr(existing, "late_result_pending", False)) or _attempt_has_active_review_custody(attempt) stamp_paid_review_owner(attempt, paid=bool(paid)) state.record_attempt(attempt, semantic_redirects=_obligation_redirects) @@ -809,6 +813,7 @@ def _check_overlapping_review_attempt(ctx: ToolContext) -> Optional[str]: expiration_window = _REVIEW_ATTEMPT_TTL_SEC + _REVIEW_ATTEMPT_GRACE_SEC ctx._review_resume_pending = False ctx._pending_review_attempt = None + ctx._review_cyber_pending = "" def _mutate(state): state.expire_stale_attempts(now_ts=_utc_now()) @@ -821,6 +826,9 @@ def _check_overlapping_review_attempt(ctx: ToolContext) -> Optional[str]: active_attempts = update_state(pathlib.Path(ctx.drive_root), _mutate) except Exception as e: log.warning("Failed to check overlapping review attempts: %s", e) + if not review_enforcement_blocks("blocking"): + ctx._review_cyber_pending = f"Review custody is unreadable: {e}. No new reviewer will be dispatched." + return None return ( "⚠️ REVIEW_STATE_UNAVAILABLE: active paid-review custody could not " "be verified, so no reviewer dispatch was started. Retry after the " @@ -828,6 +836,13 @@ def _check_overlapping_review_attempt(ctx: ToolContext) -> Optional[str]: ) if not active_attempts: return None + if not review_enforcement_blocks("blocking"): + ctx._review_cyber_pending = ( + "Existing review custody remains active: " + + ", ".join(f"{item.tool_name}#{item.attempt}" for item in active_attempts) + + ". No new reviewer will be dispatched; the original attempts remain collectible." + ) + return None task_id = str(getattr(ctx, "task_id", "") or "") tool_name = _current_review_tool_name(ctx) @@ -930,6 +945,20 @@ def _check_advisory_freshness(ctx: ToolContext, commit_message: str, _record_advisory_override(ctx, warning) ctx._review_advisory = list(getattr(ctx, "_review_advisory", []) or []) + [warning, *matching_run.items] + if not review_enforcement_blocks("blocking"): + from ouroboros.tools.review import _record_advisory_override + + if not fresh or open_obs or open_debts: + warning = ("Cyber Pro: preflight status=" + str(getattr(matching_run, "status", "missing")) + + ("; current" if fresh else "; stale or unavailable") + + ". Ouroboros may continue; this does not create review evidence.\n" + + str(getattr(matching_run, "raw_result", "") or "") + + "\n" + "\n".join([*_render_obligations(), *_render_debts()])) + ctx._last_review_block_reason = "advisory_cyber_authority" + _record_advisory_override(ctx, warning) + ctx._review_advisory = list(getattr(ctx, "_review_advisory", []) or []) + [warning] + return None + if (fresh or technical_failure) and not open_obs and not open_debts: return None diff --git a/ouroboros/tools/git.py b/ouroboros/tools/git.py index 1c9699cf6..b7be7f9be 100644 --- a/ouroboros/tools/git.py +++ b/ouroboros/tools/git.py @@ -16,6 +16,7 @@ import time from typing import Any, Dict, List, Optional, Tuple from ouroboros.config import get_runtime_mode # noqa: F401 +from ouroboros.tools.review_helpers import review_enforcement_blocks from ouroboros.runtime_mode_policy import ( core_patch_notice, # noqa: F401 format_protected_paths, # noqa: F401 @@ -117,6 +118,8 @@ def _free_cycle_gate( disclosure and WITHOUT buying another review.""" from ouroboros.config import get_review_enforcement + if getattr(ctx, "_review_cyber_pending", "") and not review_enforcement_blocks("blocking"): + return {"advisory_replay": ctx._review_cyber_pending, "replay_reason": "review_pending"} fp = pre_fingerprint.get("fingerprint", "") rebuttal_sha = compute_rebuttal_sha256(review_rebuttal) contract_fp = commit_review_contract_fingerprint() @@ -169,7 +172,7 @@ def _free_cycle_gate( cycles_paid=int(ceiling["cycles_paid"]), cap=int(ceiling["cap"]), enforcement=enforcement, root_task_id=root_task_id, fingerprint=str(fp), ) - if enforcement != "blocking": + if not review_enforcement_blocks(enforcement): # ADVISORY: neither state hard-blocks a commit — disclose loudly (typed # event + result message) and reuse the recorded outcome for free. # The identical-replay half of this branch is structurally near-dead @@ -573,6 +576,12 @@ def _advisory_and_tests_gate( ctx, runner=lambda c, **kw: _run_review_preflight_tests(c, **kw)) if test_err: msg = _tests_preflight_block_message(_managed_needs_proof, test_err) + if not review_enforcement_blocks("blocking"): + from ouroboros.tools.review import _handle_review_block_or_warning + + ctx._last_review_block_reason = "tests_preflight_blocked" + _handle_review_block_or_warning(ctx, True, msg, "") + return None try: run_cmd(["git", "reset", "HEAD"], cwd=ctx.repo_dir) except Exception: diff --git a/ouroboros/tools/git_review_cycle.py b/ouroboros/tools/git_review_cycle.py index 5dee0caaf..ccd7cee14 100644 --- a/ouroboros/tools/git_review_cycle.py +++ b/ouroboros/tools/git_review_cycle.py @@ -1,14 +1,6 @@ -"""Staging, advisory/triad/scope review and reviewed-material binding for the -commit gate, split out of ``ouroboros/tools/git.py`` (v7 module-size -discipline). Every span is extracted VERBATIM from the parent's tip bytes by -scripts/v7next_transplant.py; the parent re-exports every moved name. -Parent-scope helpers the monolith read as module globals — including the -post-cutoff paid-cycle gate family — are read through the call-time handle -``_git()`` — never a from-import — so the facade binding stays the one tests -monkeypatch. ``_sanitize_git_error`` is the one f-string-read exception (the -byte gate cannot rewrite f-string internals): it binds the plumbing owner at -import time. -""" +"""Commit staging, review, continuation and binding. ``tools.git`` re-exports +these functions; ``_git()`` preserves its patchable facade bindings, while +neutral plumbing imports bind their own owner.""" from __future__ import annotations @@ -23,22 +15,15 @@ import time from typing import Any, Dict, List, Optional from ouroboros.tools.registry import ToolContext -from ouroboros.tools.git_plumbing import _sanitize_git_error -from ouroboros.tools.git_plumbing import _publish_git_error, _publish_review_blocked +from ouroboros.tools.git_plumbing import _sanitize_git_error, _publish_git_error, _publish_review_blocked +from ouroboros.tools.review_helpers import review_enforcement_blocks -# The parent's logger name is pinned so moved log records keep their %(name)s -# in server.log/stdout — the same logger object the parent binds. +# Keep the public facade's logger name in server/stdout records. log = logging.getLogger("ouroboros.tools.git") def _git(): - """The parent module, read at call time. - - The parent owns the rebindable module state and the members tests - monkeypatch there; reading them through the module at each call keeps - one binding, where a from-import would freeze the value this leaf saw - at import time (the owner-approved D18/D33 mechanical exception). - """ + """Read the public facade at call time so monkeypatches keep one binding.""" from ouroboros.tools import git return git @@ -276,8 +261,11 @@ def _finalize_pending_review( *, pre_fingerprint: Dict[str, Any], post_fingerprint: Dict[str, Any], -) -> str: - """Persist the non-terminal wave and leave its exact retry fail-closed.""" +) -> Optional[str]: + """Retain non-terminal custody; Cyber may continue without closing the wave.""" + if getattr(ctx, "_review_cyber_pending", "") and not review_enforcement_blocks("blocking"): + # The pending row belongs to an earlier invocation, not this free continuation. + return None custody_lost = bool(getattr(ctx, "_review_custody_lost", False)) message = ( "⚠️ REVIEW_CUSTODY_LOST: the paid review wave is still unresolved, but " @@ -312,6 +300,13 @@ def _finalize_pending_review( degraded_reasons=list(getattr(ctx, "_review_degraded_reasons", []) or []), review_retry_key=str(getattr(ctx, "_current_review_retry_key", "") or ""), ) + if not review_enforcement_blocks("blocking"): + from ouroboros.tools.review import _handle_review_block_or_warning + + ctx._last_review_block_reason = "review_late_result_pending" + _handle_review_block_or_warning(ctx, True, + "Physical review remains pending; its source and invocation are retained for collection.", "") + return None # The index is part of this live wave's identity. Retain it for exact # reconciliation; rebuilding it from the worktree could review new bytes. return message @@ -571,6 +566,9 @@ def _reconcile_advisory_before_preparation(ctx, commit_message, *, goal, scope, from ouroboros.tools.preflight_review_run import pending_advisory_execution ctx._advisory_reconciled = False + if not review_enforcement_blocks("blocking"): + # Commit continuation leaves each old critic's source/custody with its invocation. + return "" try: execution, _ = pending_advisory_execution( ctx, commit_message, goal=goal, scope=scope, paths=paths, review_rebuttal=review_rebuttal, @@ -774,7 +772,9 @@ def _run_reviewed_stage_cycle( # ships without a fresh review and the disclosure must say so. review_err, scope_result, triad_block_reason, triad_advisory = None, None, "", [] replay_reason = str(advisory_replay.get("replay_reason") or "") - if replay_reason == _git().IDENTICAL_DIFF_BLOCK_REASON: + if not review_enforcement_blocks("blocking"): + progress_note = "Cyber Pro: continuing without a new reviewer dispatch; original review facts are retained." + elif replay_reason == _git().IDENTICAL_DIFF_BLOCK_REASON: progress_note = ( "Max Review Cycles: identical staged diff — reusing the recorded " "review verdict, no paid triad+scope dispatch." @@ -790,6 +790,8 @@ def _run_reviewed_stage_cycle( f"this commit ({replay_reason}); no fresh automatic preflight was bought. " + str(advisory_replay.get("advisory_replay") or "") ) + if not review_enforcement_blocks("blocking"): + disclosure = "Cyber Pro: proceeding without a new review. " + str(advisory_replay.get("advisory_replay") or "") advisory_list = getattr(ctx, "_review_advisory", None) if isinstance(advisory_list, list): advisory_list.append(disclosure) @@ -826,16 +828,12 @@ def _run_reviewed_stage_cycle( if isinstance(advisory_list, list): advisory_list.extend(scope_advisory) post_fingerprint = _git()._fingerprint_staged_diff(pathlib.Path(ctx.repo_dir)) - if _git()._review_custody_pending(ctx): + if _git()._review_custody_pending(ctx) and (pending_message := _git()._finalize_pending_review( + ctx, commit_message, commit_start, + pre_fingerprint=pre_fingerprint, post_fingerprint=post_fingerprint)): return { "status": "blocked", - "message": _git()._finalize_pending_review( - ctx, - commit_message, - commit_start, - pre_fingerprint=pre_fingerprint, - post_fingerprint=post_fingerprint, - ), + "message": pending_message, "block_reason": ( "review_custody_lost" if bool(getattr(ctx, "_review_custody_lost", False)) @@ -980,10 +978,13 @@ def _run_non_committing_review_cycle( ) ctx._scope_review_history = {} outcome["message"] = ( + "Cyber Pro: review-only operation completed; independent failures and pending work remain recorded. " + if not review_enforcement_blocks("blocking") else "Review-only cycle completed under advisory enforcement; failed or missing review remains recorded. " if "review_technical_failure_advisory" in (getattr(ctx, "_review_degraded_reasons", []) or []) else "Review-only cycle passed. " - ) + "Commit was not created and the index was unstaged." + ) + ("Commit was not created; the index is retained while review custody is pending." + if _git()._review_custody_pending(ctx) else "Commit was not created and the index was unstaged.") return outcome finally: try: diff --git a/ouroboros/tools/parallel_review.py b/ouroboros/tools/parallel_review.py index 6f1794287..b176250a1 100644 --- a/ouroboros/tools/parallel_review.py +++ b/ouroboros/tools/parallel_review.py @@ -11,7 +11,7 @@ import time from ouroboros.utils import run_cmd from ouroboros.review_substrate import scope_reviewer_slots -from ouroboros.tools.review_helpers import build_scope_actor_record, format_review_history_entry +from ouroboros.tools.review_helpers import build_scope_actor_record, format_review_history_entry, review_enforcement_blocks from ouroboros.tools.scope_review import ( run_scope_review, ScopeReviewResult, @@ -160,7 +160,7 @@ def _format_scope_advisory_msg(scope_result) -> str: """Format advisory scope findings as a readable message (advisory enforcement path).""" parts = [] if scope_result.critical_findings: - parts.append("Scope advisory findings (enforcement=advisory):\n" + + parts.append("Scope review findings:\n" + "\n".join(f" • {f['item']}: {f.get('reason', '')}" for f in scope_result.critical_findings)) if scope_result.advisory_findings: @@ -493,7 +493,7 @@ def _run_scope(ctx, commit_message, scope_rows, dispatch, *, goal, scope, ) if partial_quorum_shortfall: from ouroboros.config import get_review_enforcement - if get_review_enforcement() == "blocking": + if review_enforcement_blocks(get_review_enforcement()): blocked = True block_messages.append(_qmsg) # Surface any non-blocking shortfall LOUDLY (advisory, never a silent @@ -742,7 +742,7 @@ def run_parallel_review( from ouroboros.config import get_review_enforcement blocking_review = bool((triad_prepared or {}).get( - "blocking_review", get_review_enforcement() == "blocking")) + "blocking_review", review_enforcement_blocks(get_review_enforcement()))) and review_enforcement_blocks("blocking") if not hasattr(ctx, "_review_degraded_reasons"): ctx._review_degraded_reasons = [] ctx._review_degraded_reasons.append( @@ -969,20 +969,22 @@ def aggregate_review_verdict(review_err, scope_result, triad_block_reason, triad "responded", "skipped_low_context_mode", "not_dispatched", }] technical_scope = bool(failed_scope) and all(review_failure_is_technical(row) for row in failed_scope) - if (get_review_enforcement() == "advisory" + cyber = not review_enforcement_blocks("blocking") + if cyber or (get_review_enforcement() == "advisory" and (not review_err or triad_block_reason == "fixed_overflow") and (scope_result is None or not scope_result.blocked or technical_scope)): from ouroboros.tools.review import _record_advisory_override disclosure = ( - "Review enforcement=advisory: technical review failure permits continuing " - "on the independently bound candidate; failed or missing review is not a PASS.\n" + ("Cyber Pro: independent review does not prohibit action " if cyber else + "Review enforcement=advisory: technical review failure permits continuing ") + + "on the independently bound candidate; failed or missing review is not a PASS.\n" + combined_msg ) ctx._last_review_block_reason = block_reason _record_advisory_override(ctx, disclosure) ctx._review_advisory.append(disclosure) - ctx._review_degraded_reasons = list(getattr(ctx, "_review_degraded_reasons", []) or []) + ["review_technical_failure_advisory"] + ctx._review_degraded_reasons = list(getattr(ctx, "_review_degraded_reasons", []) or []) + ["review_cyber_authority" if cyber else "review_technical_failure_advisory"] return False, combined_msg, block_reason, _combined_findings, _scope_advisory_items return True, combined_msg, block_reason, _combined_findings, _scope_advisory_items diff --git a/ouroboros/tools/plan_render.py b/ouroboros/tools/plan_render.py index dab96fac4..cf144099d 100644 --- a/ouroboros/tools/plan_render.py +++ b/ouroboros/tools/plan_render.py @@ -10,6 +10,7 @@ from typing import Any, Dict, List, Optional from ouroboros.task_results import plan_review_notes_are_annotatable from ouroboros.tools.review_synthesis import PLAN_REVIEW_CONTROL_PREFIX from ouroboros.tools.plan_spec import MAX_FINDINGS_PER_SLOT +from ouroboros.tools.review_helpers import review_enforcement_blocks # B2 (honest DEGRADED): every aggregate reaches the control line as itself — the @@ -128,10 +129,20 @@ def _next_step(wave: dict, *, enforcement: str, cap: Optional[int], cycles_paid: f"Author finish recorded as {author.get('disposition')} against this exact " "review fingerprint; raw reviewer findings remain evidence. " ) - if enforcement == "blocking": + if review_enforcement_blocks(enforcement): author_note += "Blocking enforcement still holds the open plan gate. " + elif not review_enforcement_blocks("blocking"): + author_note += "Cyber Pro preserves final judgment with Ouroboros. " else: author_note += "Advisory enforcement permits proceeding with the review open. " + if not review_enforcement_blocks("blocking"): + return ( + author_note + "Cyber Pro: Ouroboros decides whether and how to continue. " + "The recorded verdict, open findings and any unresolved physical reviewers remain " + "independent facts; continuation does not close the wave or create a PASS. " + f"The existing $0 plan_task(review_disposition={{review_fingerprint: '{fp}', items: [...]}}) " + "can collect results or record a disposition without a new panel." + ) if bool(wave.get("closed")): if plan_review_notes_are_annotatable(wave): return ( diff --git a/ouroboros/tools/plan_review.py b/ouroboros/tools/plan_review.py index af793f39a..85aa4f6a5 100644 --- a/ouroboros/tools/plan_review.py +++ b/ouroboros/tools/plan_review.py @@ -96,6 +96,7 @@ from ouroboros.tools.plan_review_references import ( from ouroboros.tools.registry import ToolContext, ToolEntry from ouroboros.tools.review_helpers import review_wave_binding_fence, review_wave_budget_gate from ouroboros.review_records import build_author_disposition_from_mapping +from ouroboros.tools.review_helpers import review_enforcement_blocks from ouroboros.tools.review_synthesis import ( PLAN_REVIEW_CONTROL_PREFIX, ) @@ -556,7 +557,7 @@ async def _run_plan_review_async(ctx: ToolContext, request: _PlanRequest, *, col elif not resume_in_flight: # stale ⇒ identical envelope re-dispatches fresh stale, replay_snapshot = _plan_wave_replay_decision(slots_fn, existing) if not stale: - if enforcement == "advisory": + if not review_enforcement_blocks(enforcement): # Still-OPEN wave: re-invoke the emitter so a durable append that FAILED # at record time retries on replay (memo only on success ⇒ landed dedups). _emit_plan_review_advisory_open(ctx, state_root, task_id=task_id, @@ -730,7 +731,7 @@ async def _run_plan_review_async(ctx: ToolContext, request: _PlanRequest, *, col except (OSError, TimeoutError, ValueError) as exc: return _typed_refusal(ctx, "TOOL_ERROR", f"ERROR: PLAN_REVIEW_STATE_INVALID: {exc}") _emit_plan_review_reference(ctx, task_id, state_root=state_root) - if enforcement == "advisory" and not stored.get("closed"): + if not review_enforcement_blocks(enforcement) and not stored.get("closed"): # B2: loud at the moment — ONE typed owner-visible event per recorded open wave. _emit_plan_review_advisory_open(ctx, state_root, task_id=task_id, wave=stored, cycles_paid=paid_now, cap=cap) @@ -835,7 +836,7 @@ def _cycles_exhausted( f"⚠️ PLAN_REVIEW_CYCLES_EXHAUSTED: {cycles_paid} of {cap} paid plan-review cycles are spent " "for this task; no reviewer was called and no cycle was consumed. " ) - if enforcement == "blocking": + if review_enforcement_blocks(enforcement): head += ( "Blocking enforcement: the plan review stays OPEN, so implementation stays held — but " "finalization is RELEASED so the task can end honestly instead of waiting for a panel it " @@ -843,6 +844,8 @@ def _cycles_exhausted( "revised spec once the owner raises OUROBOROS_REVIEW_MAX_CYCLES, or finalizing now with " "outcome_tier=blocked_with_evidence. Do not start the work under an open blocking review." ) + elif not review_enforcement_blocks("blocking"): + head += "Cyber Pro permits proceeding by Ouroboros's judgment; the open review and spent cycles remain recorded facts." else: head += ( "Advisory enforcement: you may proceed with the review open; the host records and " @@ -907,7 +910,7 @@ def _apply_disposition(ctx: ToolContext, disposition: dict) -> str: text, state, wave = _collect.collect_wave_sync(ctx, state_root=root, task_id=task_id, wave=wave) except PlanReviewSourceUnavailable as exc: return _plan_unavailable(ctx, str(exc), "plan_review_exact_artifact_unavailable") - if not disposition.get("items"): # a pure $0 peek; items are applied even while slots run + if not disposition.get("items") and not disposition.get("author_disposition"): return text cycles_paid = int(state.get("cycles_paid") or 0) if wave.get("closed") and not plan_review_notes_are_annotatable(wave): @@ -934,9 +937,9 @@ def _apply_disposition(ctx: ToolContext, disposition: dict) -> str: if review_retry_cancelled(ctx): return _bad("ERROR: PLAN_REVIEW_DISPOSITION_INVALID: cancellation prevents author finish") - if not wave.get("paid"): + if not wave.get("paid") and review_enforcement_blocks("blocking"): return _bad("ERROR: PLAN_REVIEW_DISPOSITION_INVALID: author finish requires an actual first review dispatch") - if enforcement != "advisory": + if review_enforcement_blocks(enforcement): return _bad( "ERROR: PLAN_REVIEW_DISPOSITION_INVALID: author_disposition is advisory-only; " "the selected blocking enforcement remains authoritative" diff --git a/ouroboros/tools/plan_review_runtime.py b/ouroboros/tools/plan_review_runtime.py index 0c76b49e8..94b3b6a79 100644 --- a/ouroboros/tools/plan_review_runtime.py +++ b/ouroboros/tools/plan_review_runtime.py @@ -803,6 +803,9 @@ def emit_plan_review_advisory_open( json.dumps(wave.get("health_epoch") or [], sort_keys=True, default=str)) if key in _ADVISORY_OPEN_SEEN: return + from ouroboros.config import get_review_enforcement + from ouroboros.tools.review_helpers import review_enforcement_blocks + row = { "type": "plan_review_advisory_open", "surface": "plan_review", @@ -813,7 +816,8 @@ def emit_plan_review_advisory_open( "paid": bool(wave.get("paid")), "cycles_paid": int(cycles_paid), "cap": cap, - "enforcement": "advisory", + "enforcement": get_review_enforcement(), + "decision_authority": "cyber_pro" if not review_enforcement_blocks("blocking") else "advisory", # Bounded per-slot typed facts: who failed, with what code, until when. "slots": [ {"slot_id": a.get("slot_id"), "ok": bool(a.get("ok")), diff --git a/ouroboros/tools/review.py b/ouroboros/tools/review.py index dd667c0f6..8883912e0 100644 --- a/ouroboros/tools/review.py +++ b/ouroboros/tools/review.py @@ -55,6 +55,7 @@ from ouroboros.tools.review_helpers import ( format_name_status_for_preflight, format_review_history_entry as _format_review_entry, REVIEW_PROMPT_TOKEN_BUDGET, # noqa: F401 — patchable seam (see note above) + review_enforcement_blocks, single_line as _single_line, ) @@ -702,9 +703,12 @@ def _handle_review_block_or_warning( blocked_msg: str, advisory_prefix: str, ) -> Optional[str]: - """Either block immediately or downgrade to advisory warning.""" - if blocking_review: + """Apply action authority while preserving the independent review signal.""" + cyber = not review_enforcement_blocks("blocking") + if blocking_review and not cyber: return blocked_msg + if cyber: + advisory_prefix = "Cyber Pro: review does not prohibit action; original signal follows. " _record_advisory_override(ctx, blocked_msg) _append_review_warning(ctx, advisory_prefix + blocked_msg) ctx._review_iteration_count = 0 @@ -725,6 +729,8 @@ def _record_advisory_override(ctx: ToolContext, blocked_msg: str) -> None: append_jsonl(ctx.drive_logs() / "events.jsonl", { "ts": utc_now_iso(), "type": "review_advisory_override", + "review_enforcement": _cfg.get_review_enforcement(), + "decision_authority": "cyber_pro" if not review_enforcement_blocks("blocking") else "advisory", "block_reason": reason, "message_head": str(blocked_msg or "")[:600], "task_id": str(getattr(ctx, "task_id", "") or ""), @@ -955,7 +961,7 @@ def _prepare_unified_review(ctx: ToolContext, commit_message: str, ctx._triad_withheld_seat_records = [] # reset Q28-dropped seat records ctx._review_degraded_reasons = [] # reset degraded participation markers review_enforcement = _cfg.get_review_enforcement() - blocking_review = review_enforcement == "blocking" + blocking_review = review_enforcement_blocks(review_enforcement) diff_text, subject, capture_block = _capture_triad_staged_diff(ctx, target_repo, blocking_review) if diff_text is None: # capture failed: block (blocking) or advisory-skip (None) @@ -1207,7 +1213,7 @@ def _review_actor_label(row: dict) -> str: def _dispatch_unified_review(ctx: ToolContext, commit_message: str, prepared: dict) -> Optional[str]: """Dispatch an assembled triad packet and post-process the panel verdict.""" - blocking_review = prepared["blocking_review"] + blocking_review = prepared["blocking_review"] and review_enforcement_blocks("blocking") try: result_json = _handle_multi_model_review( ctx, @@ -1331,7 +1337,9 @@ def _dispatch_unified_review(ctx: ToolContext, commit_message: str, prepared: di _record_advisory_override(ctx, "; ".join(critical_fails[:5])) _append_review_warning( ctx, - "Review enforcement=Advisory: critical review findings did not block commit.", + ("Cyber Pro: critical review findings do not prohibit action." + if not review_enforcement_blocks("blocking") else + "Review enforcement=Advisory: critical review findings did not block commit."), ) for finding in getattr(ctx, "_last_review_critical_findings", []) or []: _append_review_warning(ctx, finding) diff --git a/ouroboros/tools/review_helpers.py b/ouroboros/tools/review_helpers.py index 4bc5ee72f..dc738f52d 100644 --- a/ouroboros/tools/review_helpers.py +++ b/ouroboros/tools/review_helpers.py @@ -36,6 +36,15 @@ REPO_ROOT = Path(__file__).resolve().parent.parent.parent # non-blocking skip gate leaves headroom for default 1M-context reviewer models. REVIEW_PROMPT_TOKEN_BUDGET = 920_000 + +def review_enforcement_blocks(enforcement: str | None = None) -> bool: + """Project action authority without changing configured policy or review facts.""" + from ouroboros.config import get_review_enforcement, get_runtime_mode + from ouroboros.runtime_mode_policy import runtime_mode_at_least + + selected = get_review_enforcement() if enforcement is None else enforcement + return selected == "blocking" and not runtime_mode_at_least(get_runtime_mode(), "cyber_pro") + # Tokenizer-density calibration shared by every review surface (triad, scope, plan, # deep self-review). estimate_tokens (chars/4) tracks GPT-style tokenizers, but a # real Claude scope pack estimated at 739,508 tokens measured 1,166,914 REAL tokens diff --git a/ouroboros/tools/scope_review.py b/ouroboros/tools/scope_review.py index 7a542acdf..e03a10e73 100644 --- a/ouroboros/tools/scope_review.py +++ b/ouroboros/tools/scope_review.py @@ -1,9 +1,8 @@ """Enforcement-aware Atlas-backed scope reviewer for the commit pipeline. Runs beside triad review and sees touched context plus a generated repo atlas. Critical findings follow -``OUROBOROS_REVIEW_ENFORCEMENT``: blocking enforcement blocks, advisory -enforcement reports them without blocking. Failed rows retain their original -status and typed origin. The commit aggregate applies advisory permission to +the selected enforcement outside Cyber; Cyber findings are advisory to action. +Failed rows retain their original status and typed origin. The commit aggregate applies permission to technical failures independently of candidate, custody and owner admission. In owner-selected ``low`` context mode no reviewer runs and a typed skip is recorded. """ @@ -48,6 +47,7 @@ from ouroboros.tools.review_helpers import ( build_touched_file_pack, # noqa: F401 -- facade import surface; leaves read it through the call-time handle load_checklist_section, # noqa: F401 -- facade import surface; leaves read it through the call-time handle review_drive_root, + review_enforcement_blocks, CRITICAL_FINDING_CALIBRATION, # noqa: F401 -- facade import surface; leaves read it through the call-time handle BINARY_EXTENSIONS, # noqa: F401 -- facade import surface; leaves read it through the call-time handle _SENSITIVE_EXTENSIONS, # noqa: F401 -- facade import surface; leaves read it through the call-time handle @@ -868,7 +868,7 @@ def run_scope_review( if critical_findings: from ouroboros import config as _cfg - if _cfg.get_review_enforcement() == "blocking": + if review_enforcement_blocks(_cfg.get_review_enforcement()): return ScopeReviewResult( blocked=True, block_message=_build_block_message(critical_findings, advisory_findings), diff --git a/ouroboros/tools/skill_exec.py b/ouroboros/tools/skill_exec.py index 2a2a4eee6..ab71fe047 100644 --- a/ouroboros/tools/skill_exec.py +++ b/ouroboros/tools/skill_exec.py @@ -575,19 +575,21 @@ def _author_finish_existing_skill_review( disposition: str, rationale: str, ) -> Optional[Dict[str, Any]]: - """Apply an explicit advisory author finish without buying a new panel. + """Record author finish in Advisory or Cyber without buying a new panel. - The first reviewer panel remains the source of findings. A later finish - call accepts the current payload after deterministic preflight, including - after a local fix. Reviewer hash, findings and status stay intact; - only the author record binds the newly accepted bytes. + Advisory needs prior feedback and a passing current preflight; Cyber may + continue with either missing or failed. Reviewer hash, findings and status + stay intact; only the author record binds the newly accepted bytes. """ from ouroboros.config import get_review_enforcement from ouroboros.review_records import build_author_disposition from ouroboros.skill_loader import compute_content_hash, load_review_state, save_review_state from ouroboros.skill_review import _run_deterministic_preflight + from ouroboros.tools.review_helpers import review_enforcement_blocks - if str(get_review_enforcement() or "").strip().lower() != "advisory": + enforcement = str(get_review_enforcement() or "").strip().lower() + cyber = not review_enforcement_blocks("blocking") + if review_enforcement_blocks(enforcement): return {"error": "SKILL_REVIEW_ERROR: explicit author finish requires advisory enforcement."} loaded = load_bound_skill(binding) if loaded is None: @@ -599,9 +601,9 @@ def _author_finish_existing_skill_review( ) drive_root = binding.state_drive_root review_state = load_review_state(drive_root, skill_name, skill_type=loaded.manifest.type, skill_dir=loaded.skill_dir) - if review_state.status == "pending": + if not cyber and review_state.status == "pending": return {"error": "SKILL_REVIEW_ERROR: existing review is pending or has no reviewer verdict."} - if not (review_state.findings or review_state.raw_actor_records or review_state.raw_result): + if not cyber and not (review_state.findings or review_state.raw_actor_records or review_state.raw_result): return {"error": "SKILL_REVIEW_ERROR: no prior reviewer evidence is available for author finish."} try: author_record = build_author_disposition( @@ -609,20 +611,29 @@ def _author_finish_existing_skill_review( rationale=rationale, subject_hash=current_hash, reviewer_signal=review_state.status, - enforcement="advisory", + enforcement=enforcement, ) except ValueError as exc: return {"error": f"SKILL_REVIEW_ERROR: {exc}"} previous_hash = str(review_state.reviewed_content_hash or review_state.content_hash or "") + preflight_facts = None if previous_hash != current_hash: - # A changed payload is accepted only after the existing deterministic - # gate checks the complete current payload. This is not a reviewer - # PASS: the prior findings remain attached as historical evidence. + # Current preflight is independent evidence; Cyber may continue with its + # failure, while ordinary Advisory still requires it to pass. preflight = _run_deterministic_preflight( ctx, drive_root, loaded, current_hash, persist=False, binding=binding, ) - if preflight is not None: + if preflight is not None and not cyber: return {"error": "SKILL_REVIEW_ERROR: deterministic preflight did not pass for the current payload."} + if preflight is not None: + from ouroboros.utils import append_jsonl, utc_now_iso + + preflight_facts = {"content_hash": current_hash, "status": preflight.status, + "findings": list(preflight.findings or []), "error": preflight.error} + append_jsonl(ctx.drive_logs() / "events.jsonl", { + "ts": utc_now_iso(), "type": "skill_review_author_preflight", + "skill_name": skill_name, "decision_authority": "cyber_pro", **preflight_facts, + }) review_state.author_disposition = author_record save_review_state(drive_root, skill_name, review_state) from ouroboros.skill_loader import auto_grant_if_enabled @@ -648,6 +659,7 @@ def _author_finish_existing_skill_review( "deps_status": deps_status, "deps_error": deps_error, "extension": extension, "review_stale": review_state.is_stale_for(current_hash), "review_gate": review_state.gate_for(current_hash), + **({"author_preflight": preflight_facts} if preflight_facts is not None else {}), } diff --git a/tests/test_commit_late_invocation_checkpoint.py b/tests/test_commit_late_invocation_checkpoint.py new file mode 100644 index 000000000..fb3e2038b --- /dev/null +++ b/tests/test_commit_late_invocation_checkpoint.py @@ -0,0 +1,180 @@ +"""A released commit still retains its exact pre-reserved reviewer invocation.""" +from __future__ import annotations + +import copy +import threading +import time +from concurrent.futures import ThreadPoolExecutor +from dataclasses import asdict +from types import SimpleNamespace + +import pytest + +from ouroboros import review_state as state_store +from ouroboros.review_execution import ReviewRouteKind +from ouroboros.review_records import ReviewSlot +from ouroboros.tools.commit_gate import _record_commit_attempt +from ouroboros.tools.git import _install_paid_dispatch_stamp +from ouroboros.tools.parallel_review import _reserve_parallel_review_roster + + +@pytest.fixture +def reserved(tmp_path): + repo, drive = tmp_path / "repo", tmp_path / "data" + repo.mkdir() + ctx = SimpleNamespace( + repo_dir=repo, drive_root=drive, task_id="late-commit-task", task_metadata={}, + current_task_type="", parent_task_id="", _review_advisory=[], + _current_review_tool_name="commit_reviewed", + _current_review_retry_key="commit_review:reserved-cycle", + _current_review_contract_fingerprint="original-contract", + _current_review_rebuttal_sha256="", _review_reconcile_only=False, + ) + _install_paid_dispatch_stamp(ctx, "fixture commit", time.time(), {"fingerprint": "original-subject"}) + _reserve_parallel_review_roster( + ctx, {"row_plan": {"models": ["fake/triad"], "routes": [ReviewRouteKind.AGENT_SESSION], + "efforts": ["high"], "slot_ids": ["triad-a"]}}, + [{"slot": ReviewSlot("scope-a", "fake/scope", route=ReviewRouteKind.AGENT_SESSION), + "prepared": object(), "final": None}], + ) + attempt = load(ctx) + assert attempt.status == "reviewing" and attempt.paid + return ctx + + +def load(ctx, *, number=None): + return state_store.load_state(ctx.drive_root).latest_attempt_for( + repo_key=state_store.make_repo_key(ctx.repo_dir), tool_name="commit_reviewed", + task_id=ctx.task_id, attempt=number or ctx._current_review_attempt_number, + ) + + +def args_for(ctx, surface="multi_model_review"): + item = load(ctx) + row = item.triad_raw_results[0] if surface == "multi_model_review" else item.scope_raw_result["raw_results"][0] + return dict(repo_key=item.repo_key, tool_name=item.tool_name, task_id=item.task_id, + attempt=item.attempt, review_retry_key=item.review_retry_key, surface=surface, + slot_id=row["slot_id"], operation_id=row["operation_id"], invocation_id="fixture-invocation-" + surface) + + +def immutable_evidence(item): + value = asdict(item) + value.pop("updated_ts") + value.pop("late_result_pending") + for row in value["triad_raw_results"] + value["scope_raw_result"]["raw_results"]: + row.pop("pending_invocation_id", None) + row.pop("late_result_pending", None) + return value + + +@pytest.mark.parametrize("status", ["reviewing", "reviewed", "succeeded", "failed", "blocked"]) +@pytest.mark.parametrize("surface", ["multi_model_review", "scope_review"]) +def test_exact_paid_checkpoint_survives_author_continuation(reserved, status, surface): + ctx = reserved + arguments = args_for(ctx, surface) + _record_commit_attempt(ctx, "fixture commit", status, _strict=True) + before = load(ctx) + state_store.checkpoint_pending_review_invocation(ctx.drive_root, **arguments) + after = load(ctx) + assert after.status == status + assert immutable_evidence(after) == immutable_evidence(before) + assert after.paid and after.late_result_pending + rows = after.triad_raw_results if surface == "multi_model_review" else after.scope_raw_result["raw_results"] + assert rows[0]["pending_invocation_id"] == arguments["invocation_id"] + assert len(state_store.load_state(ctx.drive_root).attempts) == 1 + + +@pytest.mark.parametrize("axis", ["repo_key", "tool_name", "task_id", "attempt", "review_retry_key", "surface", "slot_id", "operation_id"]) +def test_wrong_identity_cannot_checkpoint_another_attempt(reserved, axis): + ctx = reserved + _record_commit_attempt(ctx, "fixture commit", "succeeded", _strict=True) + arguments = args_for(ctx) + arguments[axis] = 99 if axis == "attempt" else "other-" + arguments[axis] + before = asdict(state_store.load_state(ctx.drive_root)) + with pytest.raises(ValueError): + state_store.checkpoint_pending_review_invocation(ctx.drive_root, **arguments) + assert asdict(state_store.load_state(ctx.drive_root)) == before + + +def test_replacement_cycle_does_not_receive_old_invocation(reserved): + ctx = reserved + arguments = args_for(ctx) + def replace(state): + current = state.latest_attempt_for(repo_key=arguments["repo_key"], task_id=ctx.task_id, + tool_name="commit_reviewed", attempt=arguments["attempt"]) + current.review_retry_key = "commit_review:new-cycle" + current.triad_raw_results[0]["operation_id"] = "new-operation" + current.status = "succeeded" + state_store.update_state(ctx.drive_root, replace) + before = asdict(state_store.load_state(ctx.drive_root)) + with pytest.raises(ValueError, match="unavailable"): + state_store.checkpoint_pending_review_invocation(ctx.drive_root, **arguments) + assert asdict(state_store.load_state(ctx.drive_root)) == before + + +@pytest.mark.parametrize("defect", ["unpaid", "settled_row", "other_invocation", "duplicate_slot"]) +def test_exact_reservation_requirements_remain(reserved, defect): + ctx = reserved + arguments = args_for(ctx) + def change(state): + current = state.latest_attempt_for(repo_key=arguments["repo_key"], task_id=ctx.task_id, + tool_name="commit_reviewed", attempt=arguments["attempt"]) + current.status = "succeeded" + if defect == "unpaid": + current.paid = False + elif defect == "settled_row": + current.triad_raw_results[0]["operation_state"] = "settled" + elif defect == "other_invocation": + current.triad_raw_results[0]["pending_invocation_id"] = "different-invocation" + else: + current.triad_raw_results.append(copy.deepcopy(current.triad_raw_results[0])) + state_store.update_state(ctx.drive_root, change) + before = asdict(state_store.load_state(ctx.drive_root)) + with pytest.raises(ValueError): + state_store.checkpoint_pending_review_invocation(ctx.drive_root, **arguments) + assert asdict(state_store.load_state(ctx.drive_root)) == before + + +def test_delayed_checkpoint_closure_and_parallel_slot_keep_terminal_status(reserved): + ctx = reserved + checkpoint = ctx._review_pending_invocation_checkpoint + arguments = [args_for(ctx, surface) for surface in ("multi_model_review", "scope_review")] + entered, release = threading.Barrier(3), threading.Event() + def late(arguments): + entered.wait(timeout=5) + assert release.wait(timeout=5) + checkpoint(**{key: arguments[key] for key in ("surface", "slot_id", "operation_id", "invocation_id")}) + with ThreadPoolExecutor(max_workers=2) as pool: + pending = [pool.submit(late, item) for item in arguments] + try: + entered.wait(timeout=5) + _record_commit_attempt(ctx, "fixture commit", "succeeded", _strict=True) + before = load(ctx) + finally: + release.set() + for item in pending: + item.result(timeout=5) + after = load(ctx) + assert after.status == "succeeded" and immutable_evidence(after) == immutable_evidence(before) + assert after.triad_raw_results[0]["pending_invocation_id"] == arguments[0]["invocation_id"] + assert after.scope_raw_result["raw_results"][0]["pending_invocation_id"] == arguments[1]["invocation_id"] + state_store.checkpoint_pending_review_invocation(ctx.drive_root, **arguments[0]) + assert immutable_evidence(load(ctx)) == immutable_evidence(after) + assert len(state_store.load_state(ctx.drive_root).attempts) == 1 + + +def test_newer_other_attempt_is_untouched_by_exact_older_checkpoint(reserved): + ctx = reserved + arguments = args_for(ctx) + _record_commit_attempt(ctx, "fixture commit", "succeeded", _strict=True) + older = load(ctx) + newer = copy.deepcopy(older) + newer.attempt += 1 + newer.review_retry_key = "commit_review:newer" + newer.triad_raw_results[0]["operation_id"] = "newer-operation" + state_store.update_state(ctx.drive_root, lambda state: state.record_attempt(newer)) + before = asdict(load(ctx, number=newer.attempt)) + state_store.checkpoint_pending_review_invocation(ctx.drive_root, **arguments) + assert load(ctx).status == "succeeded" + assert asdict(load(ctx, number=newer.attempt)) == before + assert load(ctx, number=arguments["attempt"]).triad_raw_results[0]["pending_invocation_id"] == arguments["invocation_id"] diff --git a/tests/test_delivery_candidate.py b/tests/test_delivery_candidate.py index 8f3cde111..1df57c18d 100644 --- a/tests/test_delivery_candidate.py +++ b/tests/test_delivery_candidate.py @@ -359,7 +359,7 @@ def test_service_outputs_finalize_before_acceptance_and_require_replacement(tmp_ controls = [str(row.get("content") or "") for row in model_calls[1] if "[DELIVERY_FINALIZATION_CONTROL]" in str(row.get("content") or "")] assert any("keep is NOT allowed" in text for text in controls) - # The current source selector may follow the unchanged control instruction. + # A fresh source observation follows the candidate-control instruction. assert trace["delivery_candidate"]["revision"] == 2 assert trace["delivery_candidate"]["finalization_control"] == "replace" assert trace["verification_events"][0]["kind"] == "services_stopped" diff --git a/tests/test_delivery_forced_finalization.py b/tests/test_delivery_forced_finalization.py index adcae1ad8..b55deb833 100644 --- a/tests/test_delivery_forced_finalization.py +++ b/tests/test_delivery_forced_finalization.py @@ -453,9 +453,9 @@ def test_budget_latch_preserves_stale_candidate_with_resume_disclosure( loop._publish_delivery_candidate(registry, old, trace) loop._latch_final_answer_marker(trace, f"FINAL ANSWER: {answer}") - # Owner evidence invalidates the candidate without adding a tool call. The - # unconditional FINAL ANSWER latch remains useful, but its unchanged text - # must retain its old evidence provenance and carry a loud resume disclosure. + # Unprocessed owner input prevents final acceptance, but does not itself + # change the semantic criteria. Preserve the answer and evidence provenance + # with a loud resume disclosure until Main processes the new source. registry._ctx._owner_directives = [{"content": "Late answer constraint"}] monkeypatch.setattr( accounting, @@ -496,14 +496,14 @@ def test_budget_latch_preserves_stale_candidate_with_resume_disclosure( assert rebound.acceptance_binding["acceptance_status"] == "unaccepted" assert rebound.acceptance_binding["authoritative"] is False assert rebound.acceptance_binding["stale_evidence"] is True - assert returned_trace["delivery_candidate"]["evidence_current"] is False + assert returned_trace["delivery_candidate"]["evidence_current"] is True forced = returned_trace["forced_finalization"] assert forced["source"] == ( "budget_latched_fallback_stale_evidence_resume_required" ) - assert forced["evidence_current"] is False + assert forced["evidence_current"] is True assert forced["evidence_revision"] == old.evidence_revision - assert forced["current_evidence_revision"] > old.evidence_revision + assert forced["current_evidence_revision"] == old.evidence_revision assert usage["_best_effort_extracted"] is True assert trace["tool_calls"] == [] @@ -531,6 +531,8 @@ def test_provider_unavailable_preserves_stale_candidate_with_resume_disclosure( "binding_hash": "binding-old", } loop._publish_delivery_candidate(registry, old, trace) + # Source acknowledgement and semantic evidence have separate generations. + # Provider failure cannot acknowledge this source or infer new criteria. registry._ctx._owner_directives = [{"content": "Late answer constraint"}] forced_calls = 0 @@ -559,12 +561,12 @@ def test_provider_unavailable_preserves_stale_candidate_with_resume_disclosure( assert rebound.acceptance_binding["acceptance_status"] == "unaccepted" assert rebound.acceptance_binding["authoritative"] is False assert rebound.acceptance_binding["stale_evidence"] is True - assert returned_trace["delivery_candidate"]["evidence_current"] is False + assert returned_trace["delivery_candidate"]["evidence_current"] is True forced = returned_trace["forced_finalization"] assert forced["source"] == "host_fallback_stale_evidence_resume_required" - assert forced["evidence_current"] is False + assert forced["evidence_current"] is True assert forced["evidence_revision"] == old.evidence_revision - assert forced["current_evidence_revision"] > old.evidence_revision + assert forced["current_evidence_revision"] == old.evidence_revision assert usage["_best_effort_extracted"] is True assert usage["terminal_origin"] == "host_salvage" assert trace["tool_calls"] == [] @@ -1328,8 +1330,9 @@ def test_child_result_change_during_host_panel_supersedes_pass(tmp_path, monkeyp assert another_round is True assert registry._ctx._task_acceptance_reviewed is False assert trace["review_runs"][0]["superseded_by_revision"] is True + # The unified subject includes material child-result evidence. assert trace["review_runs"][0]["superseded_reason"] == ( - "host_acceptance_evidence_revision_changed" + "host_acceptance_subject_changed" ) binding = loop._delivery_acceptance_binding( registry, trace, hashlib.sha256(answer.encode("utf-8")).hexdigest(), diff --git a/tests/test_forced_acceptance_subject.py b/tests/test_forced_acceptance_subject.py new file mode 100644 index 000000000..938b992e4 --- /dev/null +++ b/tests/test_forced_acceptance_subject.py @@ -0,0 +1,205 @@ +"""Forced delivery applies the same source-addressed Main subject decision.""" +from __future__ import annotations + +import copy +import json +import queue +import time + +import pytest + +from ouroboros import loop_forced_finalization as forced +from ouroboros.loop_acceptance import capture_acceptance_observation +from ouroboros.loop_delivery import delivery_subject_hash +from ouroboros.loop_messages import _record_owner_directive, owner_source_sha256 +from tests.test_delivery_forced_finalization import _bind_host_pass, _forced_test_context + + +def _bound(tmp_path): + loop, registry, ctx, trace = _forced_test_context(tmp_path, incoming=queue.Queue()) + _record_owner_directive(registry._ctx, source="initial_user", content="Prepare the full answer.", msg_id="initial") + capture_acceptance_observation(registry._ctx, trace, ctx.incoming_messages) + candidate = loop._replace_delivery_candidate( + registry, ctx, trace, "The complete verified answer.", control="awaiting_control", + ) + registry._ctx._delivery_control_required = True + _bind_host_pass(loop, registry, trace, candidate) + return loop, registry, ctx, trace, candidate + + +@pytest.mark.parametrize("followup", ["Как дела?", "Нужен тот же полный ответ, как договаривались."]) +def test_forced_status_ack_keeps_old_subject_and_verdict(tmp_path, monkeypatch, followup): + loop, registry, ctx, trace, candidate = _bound(tmp_path) + old = delivery_subject_hash(registry._ctx, trace) + old_source = owner_source_sha256(registry._ctx) + ctx.incoming_messages.put(followup) + calls = [] + + def answer(_ctx, **_kwargs): + observed = dict(registry._ctx._acceptance_observation) + calls.append(observed) + assert observed["owner_source_sha256"] != old_source + assert observed["owner_source_sha256"] in json.dumps(ctx.messages) + assert followup in json.dumps(ctx.messages, ensure_ascii=False) + return json.dumps({"delivery_control": "keep", "acceptance_subject": { + "owner_source_sha256": observed["owner_source_sha256"], + }}) + + monkeypatch.setattr(loop, "_call_forced_model_once", answer) + text, _usage, result = forced._forced_final_answer( + ctx, prompt="Finish now", fallback_text=candidate.full_text, reason_code="round_limit", + ) + assert len(calls) == 1 + assert text == candidate.full_text + assert delivery_subject_hash(registry._ctx, trace) == old + assert candidate.acceptance_binding["authoritative"] is True + assert result["forced_finalization"]["acceptance_authoritative"] is True + assert result["forced_acceptance_subject"] == {"applied": True, "reason": ""} + assert registry._ctx._owner_directives[-1]["content"] == followup + + +def test_forced_new_criterion_with_same_answer_is_new_unaccepted_subject(tmp_path, monkeypatch): + loop, registry, ctx, trace, candidate = _bound(tmp_path) + old = delivery_subject_hash(registry._ctx, trace) + old_answer_hash = candidate.content_sha256 + ctx.incoming_messages.put("Also check the budget.") + calls = [] + + def answer(_ctx, **_kwargs): + calls.append(1) + return json.dumps({"delivery_control": "keep", "acceptance_subject": { + "owner_source_sha256": registry._ctx._acceptance_observation["owner_source_sha256"], + "effective_criteria": "Full verified answer including the checked budget.", + }}) + + monkeypatch.setattr(loop, "_call_forced_model_once", answer) + monkeypatch.setattr(loop, "_run_task_acceptance_review_once", lambda **_kwargs: pytest.fail("no forced paid panel")) + text, _usage, result = forced._forced_final_answer( + ctx, prompt="Finish now", fallback_text=candidate.full_text, reason_code="round_limit", + ) + assert calls == [1] and text == candidate.full_text + assert registry._ctx._delivery_candidate.content_sha256 == old_answer_hash + assert delivery_subject_hash(registry._ctx, trace) != old + assert registry._ctx._delivery_candidate.effective_criteria.endswith("checked budget.") + assert result["forced_finalization"]["acceptance_authoritative"] is False + assert result["acceptance_decision"]["status"] == "finalized_unaccepted" + assert trace["review_runs"][0]["aggregate_signal"] == "PASS" + assert trace["review_runs"][0]["superseded_by_revision"] is True + + +@pytest.mark.parametrize("invalid", [ + {"owner_source_sha256": "not-the-observed-source"}, + {"effective_criteria": "Changed without an acknowledgement"}, + {"material_tool_indices": [999]}, +]) +def test_invalid_forced_subject_never_borrows_old_review(tmp_path, monkeypatch, invalid): + loop, registry, ctx, trace, candidate = _bound(tmp_path) + + def answer(_ctx, **_kwargs): + subject = {"owner_source_sha256": registry._ctx._acceptance_observation["owner_source_sha256"]} + subject.update(invalid) + if "owner_source_sha256" not in invalid and "effective_criteria" in invalid: + subject.pop("owner_source_sha256") + return json.dumps({"delivery_control": "keep", "acceptance_subject": subject}) + + monkeypatch.setattr(loop, "_call_forced_model_once", answer) + text, _usage, result = forced._forced_final_answer( + ctx, prompt="Finish now", fallback_text=candidate.full_text, reason_code="round_limit", + ) + assert text == candidate.full_text + assert result["forced_acceptance_subject"]["applied"] is False + assert result["forced_finalization"]["acceptance_authoritative"] is False + assert result["acceptance_decision"]["status"] == "finalized_unaccepted" + assert trace["review_runs"][0]["aggregate_signal"] == "PASS" + + +def test_arrival_during_single_forced_send_does_not_ack_unseen_source_or_resend(tmp_path, monkeypatch): + loop, registry, ctx, trace, candidate = _bound(tmp_path) + observed = [] + + def answer(_ctx, **_kwargs): + snapshot = dict(registry._ctx._acceptance_observation) + observed.append(snapshot) + ctx.incoming_messages.put("Late new requirement.") + return json.dumps({"delivery_control": "keep", "acceptance_subject": { + "owner_source_sha256": snapshot["owner_source_sha256"], + }}) + + monkeypatch.setattr(loop, "_call_forced_model_once", answer) + _text, _usage, result = forced._forced_final_answer( + ctx, prompt="Finish now", fallback_text=candidate.full_text, + reason_code="owner_requested_finalization", single_semantic_turn=True, + ) + assert len(observed) == 1 + assert registry._ctx._acceptance_observation == observed[0] + assert registry._ctx._acceptance_ack_source_sha256 != owner_source_sha256(registry._ctx) + assert result["forced_finalization"]["acceptance_authoritative"] is False + assert registry._ctx._owner_directives[-1]["content"] == "Late new requirement." + + +def test_existing_refresh_reobserves_before_second_send_not_after_first_reply(tmp_path, monkeypatch): + loop, registry, ctx, trace, candidate = _bound(tmp_path) + snapshots = [] + + def answer(_ctx, **_kwargs): + snapshot = dict(registry._ctx._acceptance_observation) + snapshots.append(snapshot) + if len(snapshots) == 1: + ctx.incoming_messages.put("Как дела?") + return json.dumps({"delivery_control": "keep", "acceptance_subject": { + "owner_source_sha256": snapshot["owner_source_sha256"], + }}) + + monkeypatch.setattr(loop, "_call_forced_model_once", answer) + text, _usage, result = forced._forced_final_answer( + ctx, prompt="Finish now", fallback_text=candidate.full_text, reason_code="round_limit", + ) + assert len(snapshots) == 2 + assert snapshots[0]["owner_source_sha256"] != snapshots[1]["owner_source_sha256"] + assert registry._ctx._acceptance_ack_source_sha256 == snapshots[1]["owner_source_sha256"] + assert text == candidate.full_text + assert result["forced_finalization"]["acceptance_authoritative"] is True + + +def test_prepared_budget_request_keeps_its_first_observation_and_exact_messages(tmp_path, monkeypatch): + loop, registry, ctx, trace, candidate = _bound(tmp_path) + ctx.incoming_messages.put("Status?") + prompt = forced._prepare_forced_prompt(ctx, "Budget final", trace) + first_observation = copy.deepcopy(registry._ctx._acceptance_observation) + send_messages = copy.deepcopy(ctx.messages) + loop._append_or_merge_user_message(send_messages, prompt) + prepared = object() + calls = [] + + def answer(_ctx, **kwargs): + calls.append(kwargs) + assert registry._ctx._acceptance_observation == first_observation + assert ctx.messages == send_messages + return json.dumps({"delivery_control": "keep", "acceptance_subject": { + "owner_source_sha256": first_observation["owner_source_sha256"], + }}) + + monkeypatch.setattr(loop, "_call_forced_model_once", answer) + text, _usage, result = forced._forced_final_answer( + ctx, prompt=prompt, fallback_text=candidate.full_text, reason_code="budget_exhausted", + _prompt_prepared=True, _initial_messages=send_messages, _admitted_request=prepared, + ) + assert len(calls) == 1 + assert calls[0]["admitted_request"] is prepared + assert calls[0]["initial_messages"] is send_messages + assert text == candidate.full_text and result["forced_finalization"]["acceptance_authoritative"] + + +def test_elapsed_deadline_preserves_input_without_claiming_it_was_processed(tmp_path, monkeypatch): + loop, registry, ctx, trace, candidate = _bound(tmp_path) + old_ack = registry._ctx._acceptance_ack_source_sha256 + ctx.incoming_messages.put("New budget criterion.") + ctx.deadline_ts = time.time() - 1 + monkeypatch.setattr(loop, "_call_forced_model_once", lambda *_a, **_k: pytest.fail("deadline already elapsed")) + text, _usage, result = forced._forced_final_answer( + ctx, prompt="Finish now", fallback_text=candidate.full_text, reason_code="deadline", + ) + assert candidate.full_text in text + assert registry._ctx._acceptance_ack_source_sha256 == old_ack + assert old_ack != owner_source_sha256(registry._ctx) + assert result["forced_finalization"]["acceptance_authoritative"] is False diff --git a/tests/test_owner_settings_write_seam.py b/tests/test_owner_settings_write_seam.py index 2c99fa496..cb961598e 100644 --- a/tests/test_owner_settings_write_seam.py +++ b/tests/test_owner_settings_write_seam.py @@ -113,7 +113,7 @@ def test_cyber_save_settings_can_configure_supervisor_and_keys(cyber_settings): @pytest.mark.parametrize("key,value", [ - ("OUROBOROS_SAFETY_MODE", "off"), ("OUROBOROS_CONTEXT_MODE", "low"), + ("OUROBOROS_SAFETY_MODE", "off"), ("OUROBOROS_CONTEXT_MODE", "low"), ("OUROBOROS_CONTEXT_MODE", "nano"), ]) def test_pro_lowering_ratchets_use_effective_boot_mode(cyber_settings, monkeypatch, key, value): from ouroboros import config as cfg @@ -165,7 +165,8 @@ def test_cyber_generic_post_saves_controls_and_preserves_fact_provenance(cyber_s @pytest.mark.parametrize("mode,status", [("cyber_pro", 200), ("pro", 409)]) -def test_context_owner_endpoint_allows_cyber_during_work(cyber_settings, monkeypatch, mode, status): +@pytest.mark.parametrize("context_mode", ["low", "nano"]) +def test_context_owner_endpoint_allows_cyber_during_work(cyber_settings, monkeypatch, mode, status, context_mode): from ouroboros import config as cfg from ouroboros.gateway import settings as settings_mod from supervisor.active_activity import get_direct_activity_registry @@ -179,15 +180,16 @@ def test_context_owner_endpoint_allows_cyber_during_work(cyber_settings, monkeyp registry.register("settings-author", 1) try: assert settings_mod._has_running_agent_tasks() - response = TestClient(app).post("/api/owner/context-mode", json={"mode": "low"}) + response = TestClient(app).post("/api/owner/context-mode", json={"mode": context_mode}) finally: registry.unregister("settings-author") assert response.status_code == status, response.text - assert json.loads(cyber_settings.read_text())["OUROBOROS_CONTEXT_MODE"] == ("low" if status == 200 else "max") + assert json.loads(cyber_settings.read_text())["OUROBOROS_CONTEXT_MODE"] == (context_mode if status == 200 else "max") @pytest.mark.parametrize("has_context", [True, False]) -def test_cyber_can_self_lower_context_and_author_its_marker(cyber_settings, has_context): +@pytest.mark.parametrize("context_mode", ["low", "nano"]) +def test_cyber_can_self_lower_context_and_author_its_marker(cyber_settings, has_context, context_mode): from ouroboros import config as cfg if not has_context: @@ -195,11 +197,11 @@ def test_cyber_can_self_lower_context_and_author_its_marker(cyber_settings, has_ raw.pop("OUROBOROS_CONTEXT_MODE") raw.pop("OUROBOROS_CONTEXT_MODE_AUTO_LOW") cyber_settings.write_text(json.dumps(raw)) - cfg.save_settings({**cfg.load_settings(), "OUROBOROS_CONTEXT_MODE": "low"}) + cfg.save_settings({**cfg.load_settings(), "OUROBOROS_CONTEXT_MODE": context_mode}) stored = json.loads(cyber_settings.read_text()) - assert stored["OUROBOROS_CONTEXT_MODE"] == "low" - assert stored["OUROBOROS_CONTEXT_MODE_AUTO_LOW"] == "false" - assert cfg.load_settings()["OUROBOROS_CONTEXT_MODE"] == "low" + assert stored["OUROBOROS_CONTEXT_MODE"] == context_mode + assert cfg.load_settings()["OUROBOROS_CONTEXT_MODE_AUTO_LOW"] == "false" + assert cfg.load_settings()["OUROBOROS_CONTEXT_MODE"] == context_mode def test_cyber_context_save_retains_current_task_snapshot(cyber_settings, monkeypatch): diff --git a/tests/test_review_cyber_authority.py b/tests/test_review_cyber_authority.py new file mode 100644 index 000000000..19109018c --- /dev/null +++ b/tests/test_review_cyber_authority.py @@ -0,0 +1,226 @@ +"""Cyber action authority is separate from reviewer and physical-operation facts.""" + +import copy +import json + +import pytest + +from ouroboros import config +from ouroboros.tools import git, plan_review +from ouroboros.tools.parallel_review import aggregate_review_verdict +from ouroboros.tools.review_helpers import build_scope_actor_record, review_enforcement_blocks +from ouroboros.tools.scope_review import ScopeReviewResult +from tests.test_advisory_inline_freshness import candidate # noqa: F401 +from tests.test_plan_review_engine import harness, _call, _state # noqa: F401 + + +@pytest.fixture(params=["pro", "cyber_pro"]) +def access(request, monkeypatch): + config.reset_runtime_mode_baseline_for_tests() + config.initialize_runtime_mode_baseline(request.param) + monkeypatch.setenv("OUROBOROS_REVIEW_ENFORCEMENT", "blocking") + yield request.param + config.reset_runtime_mode_baseline_for_tests() + + +def test_effective_authority_keeps_configured_enforcement(access): + assert config.get_review_enforcement() == "blocking" + assert review_enforcement_blocks() == (access == "pro") + assert not review_enforcement_blocks("advisory") + + +@pytest.mark.parametrize("status,phase", [ + ("responded", ""), ("error", "context"), ("error", "delivery"), + ("parse_failure", "format"), ("sub_floor", "window_authority"), + ("not_dispatched", "admission"), ("pending", "delivery"), +]) +def test_scope_review_facts_survive_action_authority(candidate, access, status, phase): # noqa: F811 + finding = {"item": "contract", "severity": "critical", "verdict": "FAIL", "reason": "Original criticism"} + result = ScopeReviewResult( + blocked=True, status=status, failure_phase=phase, raw_text="Exact original review", + block_message="Original failure", critical_findings=[finding], + operation_state="in_flight" if status == "pending" else "settled", + ) + candidate._last_scope_raw_results = [build_scope_actor_record(result)] + before = copy.deepcopy(result.__dict__) + blocked, _, _, findings, _ = aggregate_review_verdict( + None, result, "", [], candidate, "candidate", 0, candidate.repo_dir, + ) + assert blocked == (access == "pro") + assert result.__dict__ == before + assert findings[0]["verdict"] == "FAIL" + if access == "cyber_pro": + event = json.loads((candidate.drive_logs() / "events.jsonl").read_text().splitlines()[-1]) + assert event["review_enforcement"] == "blocking" + assert event["decision_authority"] == "cyber_pro" + + +def test_missing_preflight_does_not_become_a_review(candidate, access): # noqa: F811 + from ouroboros.review_state import load_state + + outcome = git._check_advisory_freshness(candidate, "candidate", paths=["change.py"]) + assert (outcome is None) == (access == "cyber_pro") + assert load_state(candidate.drive_root).advisory_runs == [] + + +def test_review_status_readiness_matches_actual_cyber_gate(candidate, access): # noqa: F811 + from ouroboros.tools.claude_advisory_review import _handle_review_status + from ouroboros.review_state import load_state + + projection = json.loads(_handle_review_status(candidate)) + assert projection["repo_commit_ready"] == (access == "cyber_pro") + assert not projection["advisory_runs"] + assert projection["latest_advisory_status"] != "fresh" + assert not load_state(candidate.drive_root).advisory_runs + + +def test_actual_staged_candidate_can_continue_after_failed_review(candidate, access, monkeypatch): # noqa: F811 + from ouroboros.tools import git_review_cycle + + result = ScopeReviewResult(blocked=True, status="error", failure_phase="context", block_message="Missing required source") + monkeypatch.setattr(git, "_advisory_and_tests_gate", lambda *a, **k: None) + monkeypatch.setattr(git, "_install_paid_dispatch_stamp", lambda *a, **k: None) + monkeypatch.setattr(git, "_reconcile_and_clear_review_roster", lambda *a, **k: None) + monkeypatch.setattr(git, "_run_parallel_review", lambda *a, **k: (None, result, "", [])) + git._reset_commit_review_state(candidate) + outcome = git_review_cycle._run_reviewed_stage_cycle( + candidate, "candidate", 0, paths=["change.py"], require_release_tag=False, + ) + assert outcome["status"] == ("passed" if access == "cyber_pro" else "blocked") + assert result.status == "error" and result.blocked + if access == "cyber_pro": + assert outcome["pre_fingerprint"]["fingerprint"] == outcome["post_fingerprint"]["fingerprint"] + assert "value = 2" in git.run_cmd(["git", "show", ":change.py"], cwd=candidate.repo_dir) + + +def test_pending_review_retains_custody_when_author_continues(candidate, access): # noqa: F811 + from ouroboros.review_state import load_state + + candidate._last_triad_raw_results = [{"slot_id": "s1", "status": "error", "operation_state": "in_flight", "operation_id": "op-1"}] + candidate._current_review_retry_key = "same-paid-work" + result = git._finalize_pending_review(candidate, "candidate", 0, + pre_fingerprint={"fingerprint": "fp"}, post_fingerprint={"fingerprint": "fp"}) + assert (result is None) == (access == "cyber_pro") + saved = load_state(candidate.drive_root).attempts[-1] + assert saved.status == "reviewing" and saved.late_result_pending + assert saved.triad_raw_results[0]["operation_id"] == "op-1" + if access == "cyber_pro": + git._record_commit_attempt(candidate, "candidate", "succeeded") + saved = load_state(candidate.drive_root).attempts[-1] + assert saved.late_result_pending + assert saved.triad_raw_results[0]["operation_state"] == "in_flight" + + +def test_pending_cyber_commit_uses_no_second_dispatch(candidate, access, monkeypatch): # noqa: F811 + from ouroboros.review_state import load_state + + candidate._current_review_retry_key = "old-review" + candidate._last_triad_raw_results = [{"slot_id": "critic", "operation_id": "original-op", "operation_state": "in_flight"}] + git._finalize_pending_review(candidate, "old candidate", 0, + pre_fingerprint={"fingerprint": "old"}, post_fingerprint={"fingerprint": "old"}) + before = load_state(candidate.drive_root).attempts[-1].triad_raw_results + git._reset_commit_review_state(candidate) + outcome = git._check_overlapping_review_attempt(candidate) + if access == "cyber_pro": + assert outcome is None + monkeypatch.setattr(git, "check_review_cycles_ceiling", lambda *a, **k: pytest.fail("new review admission")) + free = git._free_cycle_gate(candidate, "new candidate", 0, + pre_fingerprint={"fingerprint": "new"}, review_rebuttal="") + assert free["replay_reason"] == "review_pending" + assert load_state(candidate.drive_root).attempts[-1].triad_raw_results == before + monkeypatch.setattr(git, "_advisory_and_tests_gate", lambda *a, **k: None) + monkeypatch.setattr(git, "_run_parallel_review", lambda *a, **k: pytest.fail("duplicate paid panel")) + cycle = git._run_reviewed_stage_cycle(candidate, "new candidate", 0, + paths=["change.py"], require_release_tag=False) + assert cycle["status"] == "passed" + git._record_commit_attempt(candidate, "new candidate", "succeeded") + attempts = load_state(candidate.drive_root).attempts + assert attempts[-2].late_result_pending and attempts[-2].triad_raw_results == before + assert not attempts[-1].late_result_pending and not attempts[-1].triad_raw_results + else: + assert candidate._review_resume_pending or outcome is not None + + +@pytest.mark.parametrize("status,stale", [("blockers", False), ("pending", False), ("clean", True)]) +def test_skill_gate_does_not_relabel_the_verdict(access, status, stale): + from ouroboros.skill_review_status import skill_review_gate + + result = skill_review_gate(status, stale=stale, findings=[{"item": "skill_preflight", "verdict": "FAIL"}]) + assert result["executable_review"] == (access == "cyber_pro") + assert result["status"] == status and result["stale"] == stale + assert result["review_enforcement"] == "blocking" + assert result["preflight_failed"] == (not stale) + + +def test_skill_author_can_finish_without_fabricating_first_feedback(tmp_path, access, monkeypatch): + from ouroboros.skill_loader import load_review_state, save_enabled + from ouroboros.tool_access_types import ResolvedResourceBinding + from ouroboros.tools import skill_exec + from tests.test_skill_exec import _build_skill, _make_ctx + + ctx = _make_ctx(tmp_path) + directory = _build_skill(ctx.drive_root / "skills" / "external", "demo") + binding = ResolvedResourceBinding(profile="self_modification", root="skill_payload", operation="review", + base_path=directory, target_path=directory, source="test", skill_name="demo", state_drive_root=ctx.drive_root) + monkeypatch.setattr(skill_exec, "run_skill_review_lifecycle_blocking", lambda *a, **k: pytest.fail("Unexpected panel"), raising=False) + result = skill_exec._author_finish_existing_skill_review(ctx, binding, "demo", + disposition="accepted", rationale="Run this local greeter with the available evidence.") + if access == "pro": + assert "requires advisory" in result["error"] + return + assert "error" not in result, result + saved = load_review_state(ctx.drive_root, "demo") + assert saved.status == "pending" and not saved.raw_actor_records and not saved.raw_result + assert saved.author_disposition["reviewer_signal"] == "pending" + assert saved.author_disposition["enforcement"] == "blocking" + save_enabled(ctx.drive_root, "demo", True) + actual = json.loads(skill_exec._handle_skill_exec(ctx, skill="demo", script="hello.py")) + assert actual["exit_code"] == 0 and "hello from skill" in actual["stdout"] + + +def test_plan_author_finish_preserves_degraded_wave(harness, access, monkeypatch): # noqa: F811 + from ouroboros.tools.plan_review_artifacts import read_wave + + sub = harness.install({"s1": "", "s2": "", "s3": ""}) + ctx = harness.make_ctx() + _call(ctx) + before = _state(harness)["waves"][-1] + fingerprint = before["request_fingerprint"] + result = plan_review._apply_disposition(ctx, { + "review_fingerprint": fingerprint, "items": [], + "author_disposition": {"disposition": "deferred", "rationale": "Proceed with available evidence."}, + }) + after = _state(harness)["waves"][-1] + assert len(sub.calls) == 1 + assert after["aggregate"] == before["aggregate"] == "DEGRADED" + assert after["closed"] is False + if access == "cyber_pro": + exact = read_wave(harness.drive, ctx.task_id, after["wave_artifact"]) + assert exact["author_disposition"]["enforcement"] == "blocking" + assert "Cyber Pro" in result + else: + assert "DISPOSITION_INVALID" in result + + +@pytest.mark.parametrize("state", [None, {}, {"schema_version": 2, "current_attempt": {"fingerprint": "fp", "status": "open"}}]) +def test_plan_projection_preserves_unknown_evidence(access, state): + from ouroboros.task_results import plan_review_gate_projection + + before = copy.deepcopy(state) + result = plan_review_gate_projection(state, "blocking") + assert result["allow"] == (access == "cyber_pro") + assert result["enforcement"] == "blocking" + assert result["outcome"] == "" and not result["closed"] + assert state == before + if access == "cyber_pro": + assert result["decision_authority"] == "cyber_pro" + assert result["review_status"] in {"invalid", "absent", "open"} + + +def test_real_force_plan_decision_uses_cyber_authority(harness, access): # noqa: F811 + from ouroboros.owner_hurry import force_plan_decision + + ctx = harness.make_ctx(force_plan=True) + result = force_plan_decision(ctx, {}, enforcement="blocking") + assert result["allow"] == (access == "cyber_pro") + assert result["enforcement"] == "blocking" and not result["closed"] diff --git a/tests/test_skill_cyber_payload.py b/tests/test_skill_cyber_payload.py new file mode 100644 index 000000000..d29fe78f4 --- /dev/null +++ b/tests/test_skill_cyber_payload.py @@ -0,0 +1,117 @@ +"""Review projection does not deny Cyber the bytes of an inert skill resource.""" + +import hashlib +import importlib.util +import json +import pathlib + +import pytest + +from ouroboros import config +from ouroboros.skill_loader import SkillPayloadUnreadable, compute_content_hash, load_skill +from ouroboros.skill_review_packs import _SkillBinaryPayload, _SkillFileUnreadable, _build_skill_file_packs, _read_skill_file + + +@pytest.fixture(params=["pro", "cyber_pro"]) +def access(request, monkeypatch): + config.reset_runtime_mode_baseline_for_tests() + config.initialize_runtime_mode_baseline(request.param) + # Caller-controlled env cannot elevate the effective boot mode. + monkeypatch.setenv("OUROBOROS_RUNTIME_MODE", "cyber_pro") + yield request.param + config.reset_runtime_mode_baseline_for_tests() + + +@pytest.mark.parametrize("name", [".env", "prod.env", "credentials.json", "id_rsa"]) +def test_sensitive_named_inert_resource_is_hashed_and_read_in_cyber(tmp_path, access, name): + skill = tmp_path / "demo" + skill.mkdir() + manifest = b'{"name":"demo","description":"Resource fixture","type":"instruction"}' + payload = b"INERT_TEST_RESOURCE=first\n" + (skill / "skill.json").write_bytes(manifest) + resource = skill / name + resource.write_bytes(payload) + if access == "pro": + with pytest.raises(SkillPayloadUnreadable, match="credential filename"): + compute_content_hash(skill) + return + expected = hashlib.sha256() + for path in sorted(skill.iterdir()): + expected.update(path.name.encode() + b"\0" + hashlib.sha256(path.read_bytes()).digest()) + content_hash = compute_content_hash(skill) + assert content_hash == expected.hexdigest() + pack = "\n".join(_build_skill_file_packs(skill, expected_content_hash=content_hash)) + assert f"### {name}\n" in pack and payload.decode() in pack + assert not load_skill(skill, tmp_path / "data").load_error + resource.write_bytes(b"INERT_TEST_RESOURCE=second\n") + assert compute_content_hash(skill) != content_hash + with pytest.raises(_SkillFileUnreadable, match="changed after hashing"): + _build_skill_file_packs(skill, expected_content_hash=content_hash) + + +@pytest.mark.parametrize("payload", [ + b"\x7fELF" + b"\0" * 16, # Valid UTF-8 bytes still represent a binary-format fixture. + b"MZ\x90\0" + b"\xff" * 16, + b"\xcf\xfa\xed\xfe" + b"\0" * 16, + importlib.util.MAGIC_NUMBER + b"\0" * 16, +]) +def test_native_magic_resource_has_exact_descriptor_not_decoded_text(tmp_path, access, payload): + resource = tmp_path / "notes.txt" # The suffix does not determine binary content. + resource.write_bytes(payload) + content_hash = compute_content_hash(tmp_path) # Native bytes were already hashable. + if access == "pro": + with pytest.raises(_SkillBinaryPayload): + _read_skill_file(resource) + return + text, digest, descriptor = _read_skill_file(resource) + assert text is None and digest == hashlib.sha256(payload).digest() + assert descriptor["path"] == "notes.txt" and descriptor["size"] == len(payload) + assert descriptor["sha256"] == hashlib.sha256(payload).hexdigest() + assert descriptor["format_from_magic"] + pack = "\n".join(_build_skill_file_packs(tmp_path, expected_content_hash=content_hash)) + assert "descriptor only, content not inlined" in pack + assert descriptor["sha256"] in pack and "\0" not in pack + resource.write_bytes(payload + b"\0") + assert compute_content_hash(tmp_path) != content_hash + + +def test_binary_content_and_sensitive_name_share_one_cyber_hash_surface(tmp_path, access): + resource = tmp_path / "credentials.json" + payload = b"\x7fELF" + b"\0" * 16 + resource.write_bytes(payload) + if access == "pro": + with pytest.raises(SkillPayloadUnreadable): + compute_content_hash(tmp_path) + return + content_hash = compute_content_hash(tmp_path) + pack = "\n".join(_build_skill_file_packs(tmp_path, expected_content_hash=content_hash)) + assert hashlib.sha256(payload).hexdigest() in pack + assert "credentials.json (binary file" in pack + + +def test_real_unreadable_source_stays_an_error(tmp_path, access, monkeypatch): + resource = tmp_path / "resource.dat" + resource.write_bytes(b"real bytes") + original = pathlib.Path.read_bytes + + def unavailable(path): + if path == resource: + raise OSError("fixture read failed") + return original(path) + + monkeypatch.setattr(pathlib.Path, "read_bytes", unavailable) + with pytest.raises(_SkillFileUnreadable, match="fixture read failed"): + _read_skill_file(resource) + + +def test_binary_manifest_is_not_fabricated_as_supported_text(tmp_path, access): + (tmp_path / "SKILL.md").write_bytes(b"\xff\xfe\x80opaque manifest") + loaded = load_skill(tmp_path, tmp_path / "state-root") + assert loaded is not None and loaded.load_error + assert "UnicodeDecodeError" in loaded.load_error + assert not loaded.content_hash + assert not loaded.available_for_execution + # Unsupported manifest parsing does not prevent raw resource inspection. + text, digest, descriptor = _read_skill_file(tmp_path / "SKILL.md") + assert text is None and descriptor["sha256"] == digest.hex() + assert json.loads(json.dumps(descriptor))["path"] == "SKILL.md"