Log scope context preparation failures after retrieval migration

Adapt PR #752 to the existing context_unavailable return after the packet-fit path retired. Append one best-effort row through the existing durable and live log pipeline after model-control propagation, preserving admission and the original result. Port the contributor tests to both retrieving transports, including forwarding and logging failure, successful paging and final model binding.
This commit is contained in:
Ouroboros 2026-09-18 12:51:25 +03:00
parent 54108aeb96
commit a62bf0d8f2
3 changed files with 142 additions and 134 deletions

View file

@ -421,6 +421,8 @@ Every scope row retrieves in every context mode. A native API row preserves its
`review_admission.prepare_scope_review` builds each brief through `scope_review_session.build_scope_session_task`, without dispatch. It carries intent, rebuttal, history and open obligations; the touched-path dispositions and candidate sizes; the repository index; shared governance tiers and scope checklist; calibration/output contract; task evidence; and the required-source manifest. Stable governance precedes change-relative context for cache reuse. The full staged diff stays inline while the first send fits, measured by `native_first_send_chars` including schemas and wrapper. Otherwise it becomes one exact paged source in the task artifact store, with a byte-exact `.review-drive` view for a delegated session. An unavailable store is disclosed and keeps the diff inline. Native inline sizing uses the landing threshold within `scope_first_send_bound`; sessions use `SESSION_INLINE_DIFF_CEILING_CHARS` because the harness owns their context selection. A managed-update subject is the authoritative M0→S resolution delta, never a substitute `git diff --cached` over the two-parent candidate. Invalid roots, missing subjects and failed assembly remain typed pre-dispatch failures; only an irreducible first send reaches `native_bound_below_first_send`.
When context construction fails before dispatch, preparation best-effort appends one `scope_review_preparation_failed` row to the existing events log, whose normal sink forwards it live. It names the task, slot, final model and original `context_unavailable` reason. This is row-local preparation evidence, not a panel verdict; logging failure preserves the same typed result, and model-control exceptions still propagate before logging.
`scope_required_sources.py` declares the change-relative reading minimum: touched protected runtime, prompts and frozen contracts; protection-owned families (`GIT_OPS_FAMILY_PATHS` and tool-dispatch leaves); and declared Python/browser contract twins, with both sides included. An ordinary touched file needs no whole-body obligation because the diff already carries its complete change. Each readable row binds raw-byte `source_revision`, normalized-text `complete_sha256`/`complete_chars`, range basis and candidate tree. Deleted or renamed sources retain their exact baseline preimages; paged diffs join the same manifest; unavailable sources remain explicit rows. Byte-identical governance already inline is satisfied by delivery without another read. The final rows are shared by brief and policy, and `required_sources_ref` plus `SCOPE_REQUIRED_SOURCES_POLICY` bind the actual manifest and review-contract fingerprint, so old replay authority cannot survive a changed contract. This minimum never restricts the reviewer's further reading.
Coverage records what could be observed: native receipts match source identity, opened root/path and delivered character intervals; sessions fold weaker journal evidence over the same rows. States are `complete`, `incomplete` with missing ranges, `declared_empty` for an explicitly empty manifest, and `unobserved` when measurement is unavailable. A changed source is a `source_gap`; availability or inline delivery never proves understanding. `review_context_atlas.repository_index` supplies orientation only: compact tracked-path dispositions, collapsed excluded classes, touched-file and direct-importer facts (size, digest, language, symbols and imports, with bounded totals disclosed). It renders no bodies, spends nothing, writes nothing and never limits what the reviewer may open.

View file

@ -514,6 +514,17 @@ def prepare_scope_review(
except (RuntimeError, StagedDiffUnavailable, OSError, ValueError) as exc:
from ouroboros.llm_claudexor import propagate_model_error
propagate_model_error(exc)
# Row-local preparation evidence, before any reviewer is dispatched.
try:
sr.append_jsonl(ctx.drive_logs() / "events.jsonl", {
"ts": sr.utc_now_iso(), "type": "scope_review_preparation_failed",
"task_id": getattr(ctx, "task_id", "") or "", "slot_id": slot_id,
"model": scope_model_id, "status": "error",
"failure_phase": "context", "failure_code": "context_unavailable",
"reason": str(exc),
})
except Exception:
pass
return None, sr.ScopeReviewResult(
blocked=True,
block_message=(

View file

@ -1,18 +1,20 @@
"""Scope packet preparation failures remain visible without changing admission."""
"""Scope preparation diagnostics survive the move from packets to retrieval."""
from __future__ import annotations
from dataclasses import asdict
import json
import logging
import queue
import subprocess
from types import SimpleNamespace
from unittest.mock import Mock
import pytest
from ouroboros import utils
from ouroboros.reviewer_window import ReviewerWindow
from ouroboros.tools import review_admission as admission
from ouroboros.tools import scope_review as sr
from ouroboros.tools import scope_review_session as session
from supervisor.log_addressing import make_server_log_sink
@ -28,8 +30,7 @@ def scope_env(tmp_path, monkeypatch):
"## Intent / Scope Review Checklist\n\nplaceholder\n", encoding="utf-8")
(repo / "docs" / "DEVELOPMENT.md").write_text("development\n", encoding="utf-8")
(repo / "BIBLE.md").write_text("constitution\n", encoding="utf-8")
(repo / "prompts").mkdir()
(repo / "prompts" / "required.md").write_text("x" * 935_000, encoding="utf-8")
(repo / ".gitignore").write_text(".review-drive/\n", encoding="utf-8")
(repo / "example.py").write_text("value = 1\n", encoding="utf-8")
for args in (("init",), ("add", "."), ("commit", "-m", "initial")):
subprocess.run(
@ -43,116 +44,65 @@ def scope_env(tmp_path, monkeypatch):
repo_dir=repo, drive_root=drive, task_id="scope-task",
drive_logs=lambda: drive / "logs", pending_events=[],
)
monkeypatch.setattr(sr, "_scope_review_skipped_in_low_context", lambda: False)
monkeypatch.setattr(sr, "_scope_window", lambda *_a, **_k: ReviewerWindow(
window_tokens=1_000_000, status="confirmed"))
monkeypatch.setattr(sr, "_effective_scope_input_limit", lambda **_k: 200_000)
monkeypatch.setattr(admission, "density_probe_before_size_refusal", lambda *_a, **_k: "warm")
monkeypatch.setattr(session, "scope_first_send_bound", lambda _brief: 900_000)
dispatch = Mock(side_effect=AssertionError("preparation must not dispatch a reviewer"))
monkeypatch.setattr(sr, "_call_scope_llm", dispatch)
forwarded = []
bridge = SimpleNamespace(push_log=forwarded.append)
monkeypatch.setattr(utils, "_log_sink", make_server_log_sink(
bridge, drive, running={"scope-task": {"task": dict(_TASK_ADDRESS)}},
))
return ctx, forwarded
token = sr._SCOPE_CONTEXT_MANIFEST.set({})
yield ctx, forwarded, dispatch
sr._SCOPE_CONTEXT_MANIFEST.reset(token)
dispatch.assert_not_called()
def _events(ctx):
path = ctx.drive_logs() / "events.jsonl"
path = ctx.drive_root / "logs" / "events.jsonl"
return [json.loads(line) for line in path.read_text(encoding="utf-8").splitlines()] if path.exists() else []
@pytest.mark.parametrize("window,status", [(1_000_000, "fixed_overflow"), (200_000, "sub_floor")])
@pytest.mark.parametrize("limit,mixed", [(200_000, False), (6_000, True)])
def test_unassembled_packet_persists_and_forwards_one_row(scope_env, monkeypatch, window, status, limit, mixed):
ctx, forwarded = scope_env
monkeypatch.setattr(sr, "_scope_window", lambda *_a, **_k: ReviewerWindow(
window_tokens=window, status="confirmed"))
monkeypatch.setattr(sr, "_effective_scope_input_limit", lambda **_k: limit)
@pytest.mark.parametrize("route", ["api_chat", "agent_session"])
@pytest.mark.parametrize("model_override", [False, True])
def test_context_failure_persists_and_forwards_one_row(scope_env, monkeypatch, route, model_override):
ctx, forwarded, _ = scope_env
# Exercise the actual builder's missing-checklist refusal and direct caller.
monkeypatch.setattr(session, "load_checklist_section", lambda _section: "")
waiter = SimpleNamespace(overrides={"reviewer:scope-slot": {
"model": "test/final", "model_account_override": "profile-final", "use_local": False,
}} if model_override else {})
monkeypatch.setattr("ouroboros.model_wait.current_model_wait", lambda: waiter)
prepared, final = admission.prepare_scope_review(ctx, "Update value", scope_model="test/model", slot_id="scope-api")
final = sr.run_scope_review(ctx, "Update value", scope_model="test/model", slot_id="scope-slot", route=route)
assert prepared is None and final.blocked and final.status == status
assert "prompts/required.md" in final.block_message
events = _events(ctx)
assert len(events) == 1
event = events[0]
expected_model = "test/final" if model_override and route == "api_chat" else "test/model"
assert final.blocked and final.status == "error"
assert final.failure_phase == "context" and final.failure_code == "context_unavailable"
assert final.model_id == expected_model and "could not be loaded" in final.block_message
event, = _events(ctx)
assert forwarded == [{**event, **_TASK_ADDRESS}]
assert not _TASK_ADDRESS.keys() & event.keys(), "addressing must not mutate the durable row"
assert event["ts"]
assert {key: value for key, value in event.items() if key != "ts"} == {
"type": "scope_review_pack_unassembled", "task_id": "scope-task",
"slot_id": "scope-api", "model": "test/model", "status": status,
"prompt_tokens": final.prompt_chars // 4, "prompt_tokens_source": "estimated",
"prompt_tokens_budget": limit, "headroom_tokens": limit - final.prompt_chars // 4,
"unassembled_required": ["prompts/required.md"], "atlas_overflowed": mixed,
"type": "scope_review_preparation_failed", "task_id": "scope-task",
"slot_id": "scope-slot", "model": expected_model, "status": "error",
"failure_phase": "context", "failure_code": "context_unavailable",
"reason": (
"Intent / Scope Review Checklist could not be loaded from docs/CHECKLISTS.md — "
"scope review cannot run without its checklist (fail-closed)."
),
}
if not mixed:
assert event["headroom_tokens"] > 0
def test_fixed_overflow_warning_names_unassembled_artifact(scope_env, caplog):
ctx, _ = scope_env
with caplog.at_level(logging.WARNING, logger=sr.__name__):
admission.prepare_scope_review(ctx, "Update value", scope_model="test/model")
warnings = [record.getMessage() for record in caplog.records if record.name == sr.__name__]
assert any("Scope review pack did not assemble:" in text and "prompts/required.md" in text for text in warnings)
assert not any("irreducible scope prompt" in text for text in warnings)
@pytest.mark.parametrize("rebuilt_status", [None, "fixed_overflow"])
@pytest.mark.parametrize("model_override", [False, True])
def test_only_final_density_rebuild_state_is_logged(scope_env, monkeypatch, rebuilt_status, model_override):
ctx, forwarded = scope_env
calls = []
limits = []
waiter = SimpleNamespace(overrides={})
monkeypatch.setattr("ouroboros.model_wait.current_model_wait", lambda: waiter)
def probe(*_args, **_kwargs):
if model_override:
waiter.overrides["reviewer:scope-api"] = {
"model": "test/final", "model_account_override": "profile-final", "use_local": False,
}
return "warm" if model_override else "measured"
def build(*_args, **_kwargs):
calls.append(1)
sr._SCOPE_CONTEXT_MANIFEST.set({"selected": [], "build": len(calls)})
if len(calls) == 1:
return None, sr._TouchedContextStatus(status="fixed_overflow", token_count=900)
return ("assembled", None) if rebuilt_status is None else (
None, sr._TouchedContextStatus(status=rebuilt_status, token_count=350, atlas_overflowed=True))
def limit(**kwargs):
limits.append(kwargs)
return 300 if len(limits) == 1 else 999
monkeypatch.setattr(sr, "_build_scope_prompt", build)
monkeypatch.setattr(sr, "_effective_scope_input_limit", limit)
monkeypatch.setattr(admission, "density_probe_before_size_refusal", probe)
prepared, final = admission.prepare_scope_review(ctx, "Update value", scope_model="test/model", slot_id="scope-api")
assert len(calls) == 2
expected_model = "test/final" if model_override else "test/model"
binding = {"model_role": "reviewer:scope-api", "credential_profile_id": "profile-final" if model_override else ""}
if model_override:
binding["use_local"] = False
assert limits == [{"scope_model": expected_model, "window_binding": binding}]
if rebuilt_status is None:
assert prepared is not None and final is None and _events(ctx) == forwarded == []
else:
assert prepared is None and final.context_manifest["build"] == 2
event, = _events(ctx)
assert forwarded == [{**event, **_TASK_ADDRESS}]
assert event["model"] == expected_model
assert event["prompt_tokens"] == 350 and event["prompt_tokens_budget"] == 300
assert event["headroom_tokens"] == -50 and event["atlas_overflowed"] is True
assert event["unassembled_required"] == []
assert ctx.pending_events == []
@pytest.mark.parametrize("failure", ["false", "append_error", "missing_logs", "sink_error"])
def test_diagnostic_failure_preserves_preparation_result(scope_env, monkeypatch, failure):
ctx, _ = scope_env
def test_diagnostic_failure_preserves_exact_preparation_result(scope_env, monkeypatch, failure):
ctx, _, _ = scope_env
monkeypatch.setattr(session, "load_checklist_section", lambda _section: "")
with monkeypatch.context() as baseline:
baseline.setattr(sr, "append_jsonl", lambda *_args: False)
_, expected = admission.prepare_scope_review(ctx, "Update value", scope_model="test/model")
calls = []
def broken_append(*args):
@ -173,49 +123,94 @@ def test_diagnostic_failure_preserves_preparation_result(scope_env, monkeypatch,
monkeypatch.setattr(utils, "_log_sink", broken_sink)
prepared, final = admission.prepare_scope_review(ctx, "Update value", scope_model="test/model")
assert prepared is None and final.blocked and final.status == "fixed_overflow"
assert final.model_id == "test/model" and "prompts/required.md" in final.block_message
assert final.context_manifest["ladder_steps"]
if failure != "missing_logs":
assert len(calls) == 1
if failure == "sink_error":
assert len(_events(ctx)) == 1
assert prepared is None and asdict(final) == asdict(expected)
assert len(calls) == (0 if failure == "missing_logs" else 1)
assert len(_events(ctx)) == (1 if failure == "sink_error" else 0)
@pytest.mark.parametrize("status", ["empty", "omitted"])
def test_other_preparation_refusals_do_not_emit_fit_event(scope_env, monkeypatch, status):
ctx, forwarded = scope_env
monkeypatch.setattr(sr, "_build_scope_prompt", lambda *_a, **_k: (
None, sr._TouchedContextStatus(status=status, omitted_paths=["unreadable.py"])))
prepared, final = admission.prepare_scope_review(ctx, "Update value", scope_model="test/model")
assert prepared is None and final.status == status
assert _events(ctx) == forwarded == []
@pytest.mark.parametrize("route", ["api_chat", "agent_session"])
@pytest.mark.parametrize("delivery", ["inline", "paged", "retrieved_by_reviewer"])
def test_successful_retrieval_preparation_emits_no_failure(scope_env, monkeypatch, route, delivery):
from ouroboros.tools import review_binary_context
def test_fit_event_stays_row_local_when_retrieving_quorum_yields_seat(scope_env, monkeypatch):
from ouroboros.review_execution import ReviewRouteKind
from ouroboros.tools import parallel_review
from ouroboros.tools import scope_review_session
ctx, forwarded = scope_env
slots = [SimpleNamespace(
slot_id=name, model="test/model", route=route, effort="", session_target="",
session_profile="", subagent_id="",
) for name, route in [
("scope-api", ReviewRouteKind.API_CHAT),
("scope-session-one", ReviewRouteKind.AGENT_SESSION),
("scope-session-two", ReviewRouteKind.AGENT_SESSION),
]]
monkeypatch.setattr(parallel_review, "scope_reviewer_slots", lambda: slots)
monkeypatch.setattr(scope_review_session, "build_scope_session_task", lambda *_a, **_k: ("retrieve evidence", {}))
rows = parallel_review._prepare_scope_rows(
ctx, "Update value", goal="", scope="", review_rebuttal="",
history_snapshot=[], scope_history=[],
ctx, forwarded, _ = scope_env
if delivery == "paged":
# Use the real source store and reader addresses, with a small test ceiling.
monkeypatch.setattr(session, "scope_first_send_bound", lambda _brief: 1)
monkeypatch.setattr(session, "SESSION_INLINE_DIFF_CEILING_CHARS", 1)
elif delivery == "retrieved_by_reviewer":
monkeypatch.setattr(review_binary_context, "capture_staged_diff", Mock(
side_effect=review_binary_context.StagedDiffUnavailable("fixture diff unavailable")))
prepared, final = admission.prepare_scope_review(
ctx, "Update value", scope_model="test/model", slot_id="scope-slot", route=route,
subagent_id="fixture-actor",
)
assert rows[0]["final"].blocked is False and rows[0]["final"].block_message == ""
assert all(row["prepared"] and row["final"] is None for row in rows[1:])
assert final is None
assert prepared["context_manifest"]["diff_delivery"] == delivery, prepared["context_manifest"]
assert _events(ctx) == forwarded == []
if delivery == "paged":
from ouroboros.artifacts import read_actor_source_bytes
source = prepared["context_manifest"]["diff_source"]
assert read_actor_source_bytes(ctx.drive_root, ctx.task_id, source).decode("utf-8") == (
review_binary_context.capture_staged_diff(ctx.repo_dir))
def test_model_control_error_propagates_before_diagnostic(scope_env, monkeypatch):
from ouroboros.llm_claudexor import ClaudexorModelError
ctx, forwarded, _ = scope_env
error = ClaudexorModelError({"code": "model_outcome_unknown", "message": "original custody"})
monkeypatch.setattr(session, "build_scope_session_task", Mock(side_effect=error))
with pytest.raises(ClaudexorModelError) as raised:
admission.prepare_scope_review(ctx, "Update value", scope_model="test/model")
assert raised.value is error and _events(ctx) == forwarded == []
def test_worker_log_envelope_forwards_without_a_new_dispatch_kind(scope_env, monkeypatch):
from supervisor import events
from supervisor.worker_process import WORKER_LOG_SINK_SUPPRESSED_TYPES
ctx, forwarded, _ = scope_env
monkeypatch.setattr(session, "load_checklist_section", lambda _section: "")
outgoing = queue.Queue()
monkeypatch.setattr(utils, "_log_sink", lambda row: utils.emit_log_event(outgoing, row))
sr.run_scope_review(ctx, "Update value", scope_model="test/model")
durable, = _events(ctx)
assert durable["type"] not in WORKER_LOG_SINK_SUPPRESSED_TYPES
envelope = outgoing.get_nowait()
assert outgoing.empty() and envelope == {"type": "log_event", "data": durable}
events.dispatch_event(envelope, SimpleNamespace(
DRIVE_ROOT=ctx.drive_root, RUNNING={ctx.task_id: {"task": dict(_TASK_ADDRESS)}},
bridge=SimpleNamespace(push_log=forwarded.append), append_jsonl=utils.append_jsonl,
))
assert forwarded == [{**durable, **_TASK_ADDRESS}] and _events(ctx) == [durable]
def test_parallel_preparation_failure_stays_row_local(scope_env, monkeypatch):
from ouroboros.review_execution import ReviewRouteKind
from ouroboros.tools import parallel_review
ctx, forwarded, _ = scope_env
slots = [SimpleNamespace(
slot_id=name, model="test/model", route=route, effort="", session_target="",
session_profile="", subagent_id="fixture-actor",
) for name, route in [("scope-native", ReviewRouteKind.API_CHAT), ("scope-session", ReviewRouteKind.AGENT_SESSION)]]
monkeypatch.setattr(parallel_review, "scope_reviewer_slots", lambda: slots)
original = session.build_scope_session_task
def build(repo, brief):
if brief.slot_id == "scope-native":
raise RuntimeError("fixture context unavailable")
return original(repo, brief)
monkeypatch.setattr(session, "build_scope_session_task", build)
rows = parallel_review._prepare_scope_rows(
ctx, "Update value", goal="", scope="", review_rebuttal="", history_snapshot=[], scope_history=[],
)
assert rows[0]["prepared"] is None and rows[0]["final"].failure_code == "context_unavailable"
assert rows[1]["prepared"] and rows[1]["final"] is None
event, = _events(ctx)
assert forwarded == [{**event, **_TASK_ADDRESS}] and event["slot_id"] == "scope-api"
assert event["status"] == "fixed_overflow"
assert not {"blocked", "verdict", "block_message", "advisory_findings"} & event.keys()
assert forwarded == [{**event, **_TASK_ADDRESS}] and event["slot_id"] == "scope-native"
assert not {"blocked", "verdict", "block_message", "context_manifest"} & event.keys()