Merge recovered tool and review source repairs into private synthesis

Keep runtime terminal custody and UI model history alongside the source-complete tool review path. Generated inventories derive from this combined tree.

Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
This commit is contained in:
Anton 2026-09-11 03:13:56 +03:00
commit 0ce757d15b
63 changed files with 1930 additions and 190 deletions

View file

@ -253,6 +253,7 @@ def _zero_diff_message(repo: pathlib.Path, prefix: dict[str, str], paths: list[s
review_history_section="", review_history_section="",
diff_text="", diff_text="",
changed_files="\n".join(paths), changed_files="\n".join(paths),
task_evidence_section="", # No task trace is part of this zero-diff baseline.
) )
return { return {
"constitutional_head_preamble_plus_BIBLE": _constitutional_head(repo), "constitutional_head_preamble_plus_BIBLE": _constitutional_head(repo),

File diff suppressed because one or more lines are too long

View file

@ -837,10 +837,10 @@ declared keys whose values are `None`, `""` or `[]`) is ignored, never mistaken
for a second operation, while any non-empty list, undeclared key or non-blank for a second operation, while any non-empty list, undeclared key or non-blank
string is meaning and makes the call mixed. Never replay the plan envelope with string is meaning and makes the call mixed. Never replay the plan envelope with
the disposition. the disposition.
Any real blocking finding, including one below quorum, stays open for a changed Any real blocking finding, including one below quorum, stays open pending a paid
spec or a justified rejection evaluated in the next paid delta cycle; advice delta review of a changed spec or justified rejection, when capacity remains; advice
does not become a blocker through repetition. Blocking `REVISE_PLAN` likewise does not become a blocker through repetition. Blocking `REVISE_PLAN` likewise
requires another panel; advisory requires another panel when another paid cycle is available; advisory
may proceed only under loud host disclosure and the agent's rationale. may proceed only under loud host disclosure and the agent's rationale.
Reviewers return findings, including optional alternatives, not a required Reviewers return findings, including optional alternatives, not a required
competing plan. A blocking finding must name the spec id it breaks, and there competing plan. A blocking finding must name the spec id it breaks, and there
@ -982,6 +982,8 @@ That required presence test is the enforcing surface; CHECKLISTS item 11
## Review & Commit Protocol ## Review & Commit Protocol
Keep optional task evidence outside the stable governance prefix; shrink its excerpt before reducing existing review material. A source pointer gives a packet-only model no retrieval capability. Rejoin preserves the original hash and project-local view while any physical reviewer may still read it. Removing an ignored view never deletes the canonical source; no separate notes corpus, blanket ToolResult metadata or mandatory whole-history read belongs to this evidence.
Reviewed commits separate improvement evidence from candidate-bound authority. Reviewed commits separate improvement evidence from candidate-bound authority.
Finish the edits and focused tests, then call `commit_reviewed`; standalone Finish the edits and focused tests, then call `commit_reviewed`; standalone
`preflight_review` remains available when an earlier critique is useful. `preflight_review` remains available when an earlier critique is useful.
@ -1674,9 +1676,7 @@ schedule retain their separate existing CI owners.
`outputs` are copied when the service stops. Directory outputs become a `outputs` are copied when the service stops. Directory outputs become a
complete manifest plus streamed zip. Policy-rejected members are skipped complete manifest plus streamed zip. Policy-rejected members are skipped
with explicit notes; missing or unreadable members fail the directory copy. with explicit notes; missing or unreadable members fail the directory copy.
`run_script` stages its temporary script under the active workspace (`.ouroboros/tmp_scripts`) for a `run_script` stages workspace scripts in a unique owned directory under `.ouroboros/tmp_scripts`, with a local `.gitignore` written before execution. Raw Git status and patches exclude that scratch without hiding neighbouring user files or changing Git configuration. Each call cleans only its own directory and then empty shared parents; cleanup failure preserves the process result with a warning. Non-workspace scripts keep the task-drive layout and garbage collection, so relative imports, generated files and toolchain
workspace-bound script and under the task drive otherwise — never the
system-repo temp path — so relative imports, generated files and toolchain
discovery observe the requested cwd (`ouroboros/tools/shell.py`; discovery observe the requested cwd (`ouroboros/tools/shell.py`;
`tests/test_shell_run_shell.py`). `tests/test_shell_run_shell.py`).
- Policy denials stay separate from execution failures: - Policy denials stay separate from execution failures:
@ -3114,6 +3114,8 @@ all source modes and its selected target independently of browser navigation.
### Declarative widgets ### Declarative widgets
A module handler that calls `OuroborosWidget.openExternal(url)` or `window.open(url)` from an anchor click also calls `event.preventDefault()`. Invoke the helper directly during the gesture, before awaiting other work; automatic relaying respects an already-handled click.
`web/modules/widgets.js` is the host for reviewed widget declarations: `web/modules/widgets.js` is the host for reviewed widget declarations:
forms/actions, text/data/media, tabs/charts, async jobs, files, forms/actions, text/data/media, tabs/charts, async jobs, files,
map/calendar/kanban, and composition through `group`, `metric`, and map/calendar/kanban, and composition through `group`, `metric`, and

View file

@ -2,7 +2,7 @@
AST-derived inventory of compatibility facades, regenerated by `python scripts/regenerate_inventories.py`. Do not edit. A facade row is any runtime module whose top-level `from <population module> import ...` statements carry the `noqa: F401` re-export marker — the codebase's declared "this binding exists for its binding, not for this module's own use" convention (reference FACADE_CONSUMERS method). Leaf domains come from `ouroboros/domains.toml`; a leaf outside the facade's domain is marked ✗ (that edge also appears in the manifest's pinned direction matrix). `tests/test_generated_inventories.py` pins byte-identity, so any re-export surface change must regenerate this file. AST-derived inventory of compatibility facades, regenerated by `python scripts/regenerate_inventories.py`. Do not edit. A facade row is any runtime module whose top-level `from <population module> import ...` statements carry the `noqa: F401` re-export marker — the codebase's declared "this binding exists for its binding, not for this module's own use" convention (reference FACADE_CONSUMERS method). Leaf domains come from `ouroboros/domains.toml`; a leaf outside the facade's domain is marked ✗ (that edge also appears in the manifest's pinned direction matrix). `tests/test_generated_inventories.py` pins byte-identity, so any re-export surface change must regenerate this file.
- facade modules: **57**; marked re-export bindings: **2339**; cross-domain facade→leaf pairs: **131** - facade modules: **57**; marked re-export bindings: **2340**; cross-domain facade→leaf pairs: **131**
| facade | domain | bindings | leaves | | facade | domain | bindings | leaves |
|---|---|---:|---| |---|---|---:|---|
@ -49,7 +49,7 @@ AST-derived inventory of compatibility facades, regenerated by `python scripts/r
| `ouroboros/tools/review.py` | D06 | 46 | `ouroboros/llm.py` (1 ✗D02)<br>`ouroboros/review_substrate.py` (3)<br>`ouroboros/reviewer_window.py` (2)<br>`ouroboros/tools/review_helpers.py` (20)<br>`ouroboros/tools/review_multi_model.py` (8)<br>`ouroboros/tools/review_response.py` (2)<br>`ouroboros/tools/review_synthesis.py` (1)<br>`ouroboros/triad_review.py` (4)<br>`ouroboros/utils.py` (5 ✗D18) | | `ouroboros/tools/review.py` | D06 | 46 | `ouroboros/llm.py` (1 ✗D02)<br>`ouroboros/review_substrate.py` (3)<br>`ouroboros/reviewer_window.py` (2)<br>`ouroboros/tools/review_helpers.py` (20)<br>`ouroboros/tools/review_multi_model.py` (8)<br>`ouroboros/tools/review_response.py` (2)<br>`ouroboros/tools/review_synthesis.py` (1)<br>`ouroboros/triad_review.py` (4)<br>`ouroboros/utils.py` (5 ✗D18) |
| `ouroboros/tools/review_helpers.py` | D06 | 60 | `ouroboros/tools/release_sync.py` (1 ✗D10)<br>`ouroboros/tools/review_file_pack.py` (29)<br>`ouroboros/tools/review_prompt_text.py` (27)<br>`ouroboros/utils.py` (3 ✗D18) | | `ouroboros/tools/review_helpers.py` | D06 | 60 | `ouroboros/tools/release_sync.py` (1 ✗D10)<br>`ouroboros/tools/review_file_pack.py` (29)<br>`ouroboros/tools/review_prompt_text.py` (27)<br>`ouroboros/utils.py` (3 ✗D18) |
| `ouroboros/tools/scope_review.py` | D06 | 85 | `ouroboros/review_substrate.py` (2)<br>`ouroboros/tools/review_binary_context.py` (3)<br>`ouroboros/tools/review_context_atlas.py` (7)<br>`ouroboros/tools/review_helpers.py` (17)<br>`ouroboros/tools/review_synthesis.py` (1)<br>`ouroboros/tools/scope_review_budget.py` (15)<br>`ouroboros/tools/scope_review_contract.py` (7)<br>`ouroboros/tools/scope_review_pack.py` (20)<br>`ouroboros/tools/scope_window.py` (8)<br>`ouroboros/utils.py` (5 ✗D18) | | `ouroboros/tools/scope_review.py` | D06 | 85 | `ouroboros/review_substrate.py` (2)<br>`ouroboros/tools/review_binary_context.py` (3)<br>`ouroboros/tools/review_context_atlas.py` (7)<br>`ouroboros/tools/review_helpers.py` (17)<br>`ouroboros/tools/review_synthesis.py` (1)<br>`ouroboros/tools/scope_review_budget.py` (15)<br>`ouroboros/tools/scope_review_contract.py` (7)<br>`ouroboros/tools/scope_review_pack.py` (20)<br>`ouroboros/tools/scope_window.py` (8)<br>`ouroboros/utils.py` (5 ✗D18) |
| `ouroboros/tools/shell.py` | D05 | 72 | `ouroboros/artifacts.py` (3)<br>`ouroboros/config.py` (2 ✗D12)<br>`ouroboros/deadline_utils.py` (1 ✗D01)<br>`ouroboros/platform_layer.py` (4 ✗D18)<br>`ouroboros/runtime_mode_policy.py` (1 ✗D13)<br>`ouroboros/shell_parse.py` (2 ✗D13)<br>`ouroboros/tool_access.py` (10 ✗D04)<br>`ouroboros/tools/deliverables_shell.py` (1 ✗D13)<br>`ouroboros/tools/process_facts.py` (2 ✗D04)<br>`ouroboros/tools/shell_audit.py` (5)<br>`ouroboros/tools/shell_effects.py` (12)<br>`ouroboros/tools/shell_outputs.py` (14)<br>`ouroboros/tools/shell_process.py` (11)<br>`ouroboros/tools/verify.py` (1)<br>`ouroboros/utils.py` (1 ✗D18)<br>`ouroboros/workspace_executor.py` (2 ✗D17) | | `ouroboros/tools/shell.py` | D05 | 73 | `ouroboros/artifacts.py` (3)<br>`ouroboros/config.py` (2 ✗D12)<br>`ouroboros/deadline_utils.py` (1 ✗D01)<br>`ouroboros/platform_layer.py` (4 ✗D18)<br>`ouroboros/runtime_mode_policy.py` (1 ✗D13)<br>`ouroboros/shell_parse.py` (3 ✗D13)<br>`ouroboros/tool_access.py` (10 ✗D04)<br>`ouroboros/tools/deliverables_shell.py` (1 ✗D13)<br>`ouroboros/tools/process_facts.py` (2 ✗D04)<br>`ouroboros/tools/shell_audit.py` (5)<br>`ouroboros/tools/shell_effects.py` (12)<br>`ouroboros/tools/shell_outputs.py` (14)<br>`ouroboros/tools/shell_process.py` (11)<br>`ouroboros/tools/verify.py` (1)<br>`ouroboros/utils.py` (1 ✗D18)<br>`ouroboros/workspace_executor.py` (2 ✗D17) |
| `ouroboros/tools/shell_guards.py` | D13 | 14 | `ouroboros/tools/write_shape.py` (14) | | `ouroboros/tools/shell_guards.py` | D13 | 14 | `ouroboros/tools/write_shape.py` (14) |
| `ouroboros/tools/shell_process.py` | D05 | 1 | `ouroboros/tools/process_facts.py` (1 ✗D04) | | `ouroboros/tools/shell_process.py` | D05 | 1 | `ouroboros/tools/process_facts.py` (1 ✗D04) |
| `ouroboros/tools/subagent_integration.py` | D07 | 13 | `ouroboros/headless.py` (2 ✗D17)<br>`ouroboros/tools/subagent_integration_delegated.py` (11) | | `ouroboros/tools/subagent_integration.py` | D07 | 13 | `ouroboros/headless.py` (2 ✗D17)<br>`ouroboros/tools/subagent_integration_delegated.py` (11) |

View file

@ -451,8 +451,8 @@ _SEED_COMPLETE_MARKER = ".bootstrap-seed-complete"
_POST_BOOTSTRAP_NEW_NATIVE_SEEDS = frozenset({"telegram", "unix_computer_use"}) _POST_BOOTSTRAP_NEW_NATIVE_SEEDS = frozenset({"telegram", "unix_computer_use"})
def _read_skill_manifest_version(skill_dir: pathlib.Path) -> str: def _read_skill_manifest(skill_dir: pathlib.Path):
"""Return a seed skill manifest version via the shared parser, or ``""``.""" """Read a seed manifest with the shared parser, or return None."""
for candidate in ("SKILL.md", "skill.json"): for candidate in ("SKILL.md", "skill.json"):
path = skill_dir / candidate path = skill_dir / candidate
if not path.is_file(): if not path.is_file():
@ -465,9 +465,15 @@ def _read_skill_manifest_version(skill_dir: pathlib.Path) -> str:
from ouroboros.contracts.skill_manifest import parse_skill_manifest_text from ouroboros.contracts.skill_manifest import parse_skill_manifest_text
manifest = parse_skill_manifest_text(text) manifest = parse_skill_manifest_text(text)
except Exception: except Exception:
return "" return None
return str(manifest.version or "").strip() return manifest
return "" return None
def _read_skill_manifest_version(skill_dir: pathlib.Path) -> str:
"""Return the parsed seed version, or an empty string when unavailable."""
manifest = _read_skill_manifest(skill_dir)
return str(manifest.version or "").strip() if manifest is not None else ""
def _reseed_native_skill_in_place( def _reseed_native_skill_in_place(
@ -532,11 +538,37 @@ def _per_skill_version_resync(
if not (target / ".seed-origin").is_file(): if not (target / ".seed-origin").is_file():
# User-managed skill in native/: never touch. # User-managed skill in native/: never touch.
continue continue
seed_version = _read_skill_manifest_version(entry) seed_manifest = _read_skill_manifest(entry)
target_version = _read_skill_manifest_version(target) target_manifest = _read_skill_manifest(target)
if seed_manifest is None or target_manifest is None:
continue
seed_version = str(seed_manifest.version or "").strip()
target_version = str(target_manifest.version or "").strip()
if not seed_version or not target_version: if not seed_version or not target_version:
continue continue
if seed_version == target_version: if seed_version == target_version:
try:
from ouroboros.skill_loader import compute_content_hash
seed_hash = compute_content_hash(
entry, manifest_entry=seed_manifest.entry,
manifest_scripts=seed_manifest.scripts,
)
target_hash = compute_content_hash(
target, manifest_entry=target_manifest.entry,
manifest_scripts=target_manifest.scripts,
)
except Exception as exc:
log_obj.warning(
"Native skill %s version %s payload comparison unavailable (%s)",
entry.name, seed_version, type(exc).__name__,
)
else:
if seed_hash != target_hash:
log_obj.warning(
"Native skill %s version %s payload differs (seed=%s, installed=%s); "
"installed files retained because the manifest version is unchanged",
entry.name, seed_version, seed_hash, target_hash,
)
continue continue
log_obj.info( log_obj.info(
"Native skill %s version drift (seed=%s, installed=%s) — re-seeding", "Native skill %s version drift (seed=%s, installed=%s) — re-seeding",

View file

@ -408,13 +408,14 @@ def run_llm_loop(
stateful_executor = StatefulToolExecutor() stateful_executor = StatefulToolExecutor()
exit_ctx = _LoopExitContext( exit_ctx = _LoopExitContext(
tools, drive_root, task_id, event_queue, drive_logs, accumulated_usage, llm_trace, tools, drive_root, task_id, event_queue, drive_logs, accumulated_usage, llm_trace,
ctx, getattr(ctx, "_execution_trace", None),
) )
_owner_msg_seen: set = set() _owner_msg_seen: set = set()
MAX_ROUNDS = _resolve_loop_max_rounds(ctx) MAX_ROUNDS = _resolve_loop_max_rounds(ctx)
round_idx = 0 round_idx, free_redial = 0, False
free_redial = False
transport_wait = None transport_wait = None
limit_ctx: Optional[_RoundLimitContext] = None limit_ctx: Optional[_RoundLimitContext] = None
ctx._execution_trace = llm_trace
try: try:
if saved: if saved:
active_model, active_effort, active_use_local, active_context_mode, round_idx, context_fit_plan = resume_native_loop( active_model, active_effort, active_use_local, active_context_mode, round_idx, context_fit_plan = resume_native_loop(

View file

@ -351,6 +351,8 @@ class _LoopExitContext:
drive_logs: pathlib.Path drive_logs: pathlib.Path
accumulated_usage: Dict[str, Any] accumulated_usage: Dict[str, Any]
llm_trace: Dict[str, Any] llm_trace: Dict[str, Any]
trace_ctx: Any = None
previous_execution_trace: Any = None
def _handle_budget_exceeded( def _handle_budget_exceeded(
@ -512,6 +514,8 @@ def _cleanup_loop_resources(
ctx: _LoopExitContext, ctx: _LoopExitContext,
) -> None: ) -> None:
"""Release attempt-scoped executors, services, and delegated runs.""" """Release attempt-scoped executors, services, and delegated runs."""
if ctx.trace_ctx is not None:
ctx.trace_ctx._execution_trace = ctx.previous_execution_trace
if stateful_executor: if stateful_executor:
try: try:
from ouroboros.tools.browser import cleanup_browser from ouroboros.tools.browser import cleanup_browser

View file

@ -1233,7 +1233,7 @@ def handle_tool_calls(
def _maybe_auto_attach_image( def _maybe_auto_attach_image(
exec_result: Dict[str, Any], exec_result: Dict[str, Any],
tools: Optional[ToolRegistry], tools: Optional[ToolRegistry],
) -> None: ) -> Optional[Dict[str, str]]:
"""Same-round image attachment for tool results that explicitly offer one. """Same-round image attachment for tool results that explicitly offer one.
A successful result whose JSON carries ``auto_attach_image: <local path>`` — a A successful result whose JSON carries ``auto_attach_image: <local path>`` — a
@ -1274,6 +1274,7 @@ def _maybe_auto_attach_image(
raw = exec_result.get("result") raw = exec_result.get("result")
if not isinstance(raw, str) or '"auto_attach_image"' not in raw: if not isinstance(raw, str) or '"auto_attach_image"' not in raw:
return return
observation = None
try: try:
parsed = json.loads(raw) parsed = json.loads(raw)
path = parsed.get("auto_attach_image") if isinstance(parsed, dict) else None path = parsed.get("auto_attach_image") if isinstance(parsed, dict) else None
@ -1282,13 +1283,16 @@ def _maybe_auto_attach_image(
ctx = getattr(tools, "_ctx", None) ctx = getattr(tools, "_ctx", None)
if ctx is None: if ctx is None:
return return
observation = {"status": "unavailable"}
from ouroboros.tools.vision import attach_local_image_to_context from ouroboros.tools.vision import attach_local_image_to_context
ok, note = attach_local_image_to_context(ctx, path) ok, note = attach_local_image_to_context(ctx, path)
observation["status"] = "attached" if ok else "unavailable"
if not ok: if not ok:
log.debug("auto-attach skipped for %s: %s", path, note) log.debug("auto-attach skipped for %s: %s", path, note)
except Exception: # noqa: BLE001 - attachment is an enhancement, never a failure except Exception: # noqa: BLE001 - attachment is an enhancement, never a failure
log.debug("auto-attach image failed", exc_info=True) log.debug("auto-attach image failed", exc_info=True)
return observation
def reclaim_trace_refs(tool_ctx: Any) -> Dict[str, Any]: def reclaim_trace_refs(tool_ctx: Any) -> Dict[str, Any]:
@ -1480,8 +1484,11 @@ def process_tool_results(
# tool messages answering the same assistant turn: the user(image) injection # tool messages answering the same assistant turn: the user(image) injection
# then preserves tool-result contiguity BY CONSTRUCTION instead of relying on # then preserves tool-result contiguity BY CONSTRUCTION instead of relying on
# the transport's adjacency repair to fix an interleaving we created ourselves. # the transport's adjacency repair to fix an interleaving we created ourselves.
by_call = {row["tool_call_id"]: row for row in llm_trace["tool_calls"][-len(results):]}
for exec_result in results: for exec_result in results:
_maybe_auto_attach_image(exec_result, tools) observation = _maybe_auto_attach_image(exec_result, tools)
if observation:
by_call[exec_result["tool_call_id"]]["image_attachment"] = observation
return error_count return error_count

View file

@ -532,6 +532,16 @@ def plan_review_reminder(decision: Dict[str, Any]) -> str:
f"{tag} An open plan review from a previous schema cannot be honored. Re-call " f"{tag} An open plan review from a previous schema cannot be honored. Re-call "
"plan_task with your goal, plan and spec to start a fresh review before finalizing." "plan_task with your goal, plan and spec to start a fresh review before finalizing."
) )
from ouroboros.review_cycles import review_max_cycles
cap = review_max_cycles()
if cap is not None and int(decision.get("cycles_paid") or 0) >= cap:
return (
f"{tag} The paid plan-review cycle cap is reached; the task cannot dispatch another paid panel "
"for either an unchanged or revised request. The recorded findings and lawful free "
"dispositions remain available; a disposition does not close blocking findings "
"or a degraded wave. Existing in-flight custody can still settle."
)
if decision.get("reviewer_slots_degraded"): if decision.get("reviewer_slots_degraded"):
# B2: facts, never a retry coach (P5). The replay promise is CONDITIONAL — # B2: facts, never a retry coach (P5). The replay promise is CONDITIONAL —
# wording SSOT: plan_render._degraded_replay_note (a free replay exists only # wording SSOT: plan_render._degraded_replay_note (a free replay exists only

View file

@ -3,6 +3,7 @@
from __future__ import annotations from __future__ import annotations
import json import json
import copy
import hashlib import hashlib
import logging import logging
import pathlib import pathlib
@ -426,6 +427,234 @@ def build_task_acceptance_evidence(
return _accept_enforce_budget(ev, budget=budget_chars) return _accept_enforce_budget(ev, budget=budget_chars)
def _commit_source_payload(ctx: Any, ref: dict, *, manifest: bool = False) -> dict:
"""Read an exact selected redacted source, never scan the observability store."""
from ouroboros.observability import read_blob_ref
from ouroboros.tool_access import canonical_data_root
errors = []
for root in dict.fromkeys((str(ctx.drive_root), str(canonical_data_root(ctx)))):
try:
source = ref
if manifest:
path = pathlib.Path(str(ref.get("path") or "")).resolve(strict=True)
path.relative_to(pathlib.Path(root).resolve() / "observability" / "calls")
raw = path.read_bytes()
if hashlib.sha256(raw).hexdigest() != ref.get("sha256"):
raise ValueError("call manifest digest mismatch")
record = json.loads(raw)
if record.get("task_id") != ctx.task_id or record.get("call_id") != ref.get("call_id"):
raise ValueError("call manifest identity mismatch")
if record.get("call_type") != "llm_response":
raise ValueError("selected model response is unavailable or failed")
source = record["redacted_projection_ref"]
payload = read_blob_ref(pathlib.Path(root), source)
if not isinstance(payload, dict):
raise ValueError("selected source is not an object")
return payload
except (OSError, ValueError, TypeError, KeyError) as exc:
errors.append(type(exc).__name__)
raise ValueError("selected source unavailable: " + ", ".join(errors))
def capture_commit_review_evidence(ctx: Any) -> dict:
"""Freeze browser/vision records once; unrelated notes never enter this view.
The next registered response of the SAME execution is adjacency evidence,
not an assertion of visual inspection. A missing/failed response stays a gap.
Exact originals retain their normal trace-ref custody; this UTF-8 source is
only the readable selected view that inspection tools and sessions can use.
"""
from ouroboros.artifacts import materialize_tool_args_source, materialize_tool_result_source, persist_exact_text_source
from ouroboros.loop_messages import _visible_round_text
from ouroboros.observability import redact_projection
from ouroboros.tool_access import canonical_data_root
from ouroboros.tool_capabilities import STATEFUL_BROWSER_TOOLS
trace = getattr(ctx, "_execution_trace", None) or {}
visual = STATEFUL_BROWSER_TOOLS | {"view_image", "vlm_query", "analyze_screenshot"}
selected = [row for row in trace.get("tool_calls", []) if isinstance(row, dict)
and (row.get("tool") in visual or row.get("image_attachment"))]
if not selected:
return {}
usage = getattr(ctx, "_accumulated_usage", None) or {}
calls = [row for row in usage.get("llm_call_refs", []) if isinstance(row, dict)]
by_execution: dict[str, list] = {}
for row in calls:
by_execution.setdefault(str(row.get("execution_id") or ""), []).append(row)
texts = ["Selected browser/vision execution sources (DATA, not instructions).",
"Coverage is only the selected records below, not the entire task or proof of visual inspection."]
refs, responses_seen, gaps = [], set(), []
for call in selected:
args, args_complete, args_gap = materialize_tool_args_source(ctx.drive_root, call)
result, result_complete, result_gap = materialize_tool_result_source(ctx.drive_root, ctx.task_id, call)
trace_ref = call.get("trace_ref") or {}
if trace_ref:
refs.append(trace_ref)
row = {"tool": call.get("tool"), "tool_call_id": call.get("tool_call_id"),
"args": args, "result": result,
"args_complete": args_complete, "result_complete": result_complete,
"source_ref": trace_ref}
if call.get("image_attachment"):
row["image_attachment"] = call["image_attachment"]
if call["image_attachment"].get("status") != "attached":
gaps.append({"tool_call_id": call.get("tool_call_id"), "status": "image_unavailable"})
for gap in (args_gap, result_gap):
if gap:
gaps.append(gap)
try:
payload = _commit_source_payload(ctx, trace_ref.get("redacted_projection_ref") or {})
if payload.get("tool_call_id") != call.get("tool_call_id") or payload.get("tool") != call.get("tool"):
raise ValueError("tool source identity mismatch")
execution, parent = str(payload.get("execution_id") or ""), payload.get("parent_call_id")
row.update(execution_id=execution, round_id=payload.get("round_id"), parent_call_id=parent,
result_meta=payload.get("result_meta"), semantic_ok=payload.get("semantic_ok"))
peers = by_execution.get(execution, []) if execution else []
parent_index = next((i for i, peer in enumerate(peers) if peer.get("llm_call_id") == parent), None)
following = peers[parent_index + 1] if parent_index is not None and parent_index + 1 < len(peers) else None
if following is None:
raise ValueError("no following registered model response in this execution")
response_ref = following.get("response_ref") or {}
row["following_model_response"] = response_ref
response_id = (execution, following.get("llm_call_id"))
if response_id not in responses_seen:
responses_seen.add(response_id)
response = _commit_source_payload(ctx, response_ref, manifest=True)
message = response.get("message")
visible = _visible_round_text(message.get("content")) if isinstance(message, dict) else ""
row["following_visible_text"] = visible
row["following_visible_text_status"] = "recorded" if visible else "no_visible_text"
refs.append(response_ref)
except (OSError, ValueError, TypeError, KeyError) as exc:
gap = {"tool_call_id": call.get("tool_call_id"), "status": "source_unavailable", "reason": str(exc)}
row["following_response_gap"] = gap
gaps.append(gap)
texts.append(json.dumps(redact_projection(row).value, ensure_ascii=False, indent=2, default=str))
text = "\n\n".join(texts)
canonical = canonical_data_root(ctx)
exact, source_ref, issue = persist_exact_text_source(canonical, ctx.task_id, source_id="commit_review", text=text)
if issue:
gaps.append(issue)
# Only this bounded display is carried in memory/requests; the full selected
# source lives in the existing source handle, not another notes corpus.
return {"source_ref": source_ref if not issue else {}, "task_id": ctx.task_id,
"data_root": str(canonical), "selected_count": len(selected),
"unselected_count": len(trace.get("tool_calls", [])) - len(selected),
"source_chars": len(text), "source_complete": not gaps,
"gap_count": len(gaps), "source_status": "unavailable" if issue else "ready",
"preview": truncate_within_limit(exact or text, _ACCEPT_NOTES_CAP),
"original_refs": copy.deepcopy(refs)}
def restore_commit_review_evidence(ctx: Any, source_ref: dict) -> dict:
"""Recover one preflight view from its recorded canonical source identity."""
from ouroboros.artifacts import read_actor_source_bytes
from ouroboros.tool_access import canonical_data_root
root = canonical_data_root(ctx)
raw = read_actor_source_bytes(root, ctx.task_id, source_ref)
text = raw.decode("utf-8")
return {"source_ref": dict(source_ref), "task_id": ctx.task_id, "data_root": str(root),
"source_chars": len(text), "source_status": "ready", "source_complete": None,
"selected_count": None, "gap_count": None,
"preview": truncate_within_limit(text, _ACCEPT_NOTES_CAP), "original_refs": []}
def pending_commit_review_evidence(ctx: Any) -> dict:
"""Read the frozen request's evidence on reconciliation, never current trace."""
attempt = getattr(ctx, "_pending_review_attempt", None)
scope = getattr(attempt, "scope_raw_result", {}) or {}
rows = [*(getattr(attempt, "triad_raw_results", []) or []), *(scope.get("raw_results") or [scope])]
for row in rows:
try:
payload = _commit_source_payload(ctx, (row.get("prompt_ref") or {}).get("redacted_projection_ref") or {})
request = payload.get("request") or {}
if request.get("task_id") != ctx.task_id:
continue
evidence = (request.get("evidence") or {}).get("task_execution")
if isinstance(evidence, dict) and evidence:
return evidence
except (OSError, ValueError, TypeError, KeyError, AttributeError):
continue
return {} # Legacy/pending-without-source is not permission to make a new one.
def materialize_commit_review_session_view(evidence: dict, repo_dir: Any) -> dict:
"""Put exact canonical bytes inside this review's already-ignored project."""
if not evidence or not evidence.get("source_ref"):
return evidence or {}
from ouroboros.artifacts import read_actor_source_bytes
from ouroboros.task_results import validate_task_id
from ouroboros.utils import run_cmd, write_bytes_atomic
result = dict(evidence)
try:
task_id = validate_task_id(evidence["task_id"])
source_ref = evidence["source_ref"]
digest = str(source_ref["sha256"])
raw = read_actor_source_bytes(evidence["data_root"], task_id, source_ref)
relative = pathlib.Path(".review-drive") / task_id / (digest + ".txt")
repo = pathlib.Path(repo_dir).resolve()
run_cmd(["git", "check-ignore", "--quiet", "--", relative.as_posix()], cwd=repo)
path = repo / relative
if not path.is_file() or path.read_bytes() != raw:
write_bytes_atomic(path, raw)
result.update(session_path=str(path), session_relative_path=relative.as_posix(), session_source_status="ready")
except (OSError, ValueError, TypeError, KeyError, RuntimeError) as exc:
result.update(session_source_status="unavailable", session_source_error=type(exc).__name__)
return result
def release_commit_review_session_view(evidence: dict) -> None:
"""Remove only this settled attempt's exact disposable copy, never its source."""
if not evidence or evidence.get("session_source_status") != "ready":
return
path = pathlib.Path(evidence["session_path"])
try:
if hashlib.sha256(path.read_bytes()).hexdigest() != evidence["source_ref"]["sha256"]:
return
path.unlink()
path.parent.rmdir()
except OSError:
pass # Concurrent consumers or retained neighbouring sources keep their directory.
def commit_review_evidence_refs(evidence: dict) -> list:
return ([evidence["source_ref"]] if evidence.get("source_ref") else []) + list(evidence.get("original_refs") or [])
def commit_review_evidence_section(evidence: dict, *, delivery: str, compact: bool = False) -> str:
"""One strictly bounded optional exhibit; provenance is never claimed reading."""
if not evidence:
return ""
ref = evidence.get("source_ref") or {}
source = evidence.get("session_relative_path") if delivery == "session" else ref.get("path")
status = evidence.get("session_source_status", "unavailable") if delivery == "session" else evidence.get("source_status")
header = ("## Recorded task browser/vision evidence\n"
f"task_id={evidence.get('task_id')}; selected_records={evidence.get('selected_count')}; "
f"unselected_records={evidence.get('unselected_count')}; "
f"source_complete={evidence.get('source_complete')}; gaps={evidence.get('gap_count')}; "
f"source_chars={evidence.get('source_chars')}.\n"
f"Retained source ({status}): {source or 'unavailable'}; sha256={ref.get('sha256', 'unavailable')}.\n")
if delivery == "packet":
header += "This packet-only reviewer has no file tools; the ref is host-retained provenance, not unseen evidence you read. Judge only the excerpt.\n"
elif status == "ready":
header += ("Read needed ranges from the project-relative file. " if delivery == "session" else
"Read needed ranges with read_file(root='artifact_store', path=the source path above, start_line/max_lines/start_char). ")
header += "Source availability and the following model response do not prove visual inspection or full reviewer coverage.\n"
else:
header += "evidence_delivery=partial: full source retrieval is unavailable; judge only the excerpt.\n"
if compact:
return truncate_within_limit(header + "evidence_delivery=partial: excerpt omitted to fit; no added verification claim.\n", _ACCEPT_NOTES_CAP)
preview = str(evidence.get("preview") or "")
complete = (not compact and status == "ready" and len(preview) == evidence.get("source_chars")
and len(header) + len(preview) + 80 <= _ACCEPT_NOTES_CAP)
header += "evidence_delivery=" + ("complete_selected" if complete else "partial") + "; other task history not selected.\n"
room = max(0, _ACCEPT_NOTES_CAP - len(header) - len("\nRecorded excerpt:\n"))
return header + "\nRecorded excerpt:\n" + truncate_within_limit(preview, room)
def collect_review_evidence( def collect_review_evidence(
drive_root: Any, drive_root: Any,
*, *,

View file

@ -64,6 +64,13 @@ def make_index_page(web_dir: pathlib.Path):
async def index_page(_request) -> FileResponse | HTMLResponse: async def index_page(_request) -> FileResponse | HTMLResponse:
index = web_dir / "index.html" index = web_dir / "index.html"
if index.exists(): if index.exists():
if _request.headers.get("X-Ouroboros-Telegram-MiniApp") == "1":
# The authenticated sidecar proxy already stamps this presentation
# hint. Bootstrap navigates to this document, so it needs its own SDK.
html = index.read_text(encoding="utf-8")
html = html.replace("<html", '<html data-ouroboros-host="telegram"', 1)
html = html.replace("<head>", '<head><script async src="https://telegram.org/js/telegram-web-app.js"></script>', 1)
return HTMLResponse(html)
return FileResponse(str(index), media_type="text/html") return FileResponse(str(index), media_type="text/html")
return HTMLResponse("<html><body><h1>Ouroboros — web/ not found</h1></body></html>", status_code=404) return HTMLResponse("<html><body><h1>Ouroboros — web/ not found</h1></body></html>", status_code=404)

View file

@ -32,7 +32,7 @@ EMBEDDED_WINDOWS_ABSOLUTE_PATH_RE = re.compile(
r"|\\\\[^\s'\"),;\]\\/]+[\\/][^\s'\"),;\]]+" r"|\\\\[^\s'\"),;\]\\/]+[\\/][^\s'\"),;\]]+"
r")" r")"
) )
_SHELLS = {"sh", "bash", "zsh"} POSIX_SHELL_HEADS = frozenset({"sh", "bash", "zsh", "dash", "ash"})
def recover_stringified_argv(text: Any) -> List[str] | None: def recover_stringified_argv(text: Any) -> List[str] | None:
@ -577,7 +577,7 @@ def local_shell_subject(raw_cmd: Any, _depth: int = 0) -> Any:
result.append(leading or ";") result.append(leading or ";")
argv = shell_argv(segment) argv = shell_argv(segment)
head = pathlib.PurePath(argv[0]).name.lower().removesuffix(".exe") if argv else "" head = pathlib.PurePath(argv[0]).name.lower().removesuffix(".exe") if argv else ""
if head in _SHELLS: if head in POSIX_SHELL_HEADS:
body = shell_command_string(argv) body = shell_command_string(argv)
local = local_shell_subject(body, _depth + 1) if body else body local = local_shell_subject(body, _depth + 1) if body else body
if local != body: if local != body:
@ -680,7 +680,7 @@ def sudo_noninteractive_violation(raw_cmd: Any) -> bool:
_env, command = collect_leading_env(segment) _env, command = collect_leading_env(segment)
while command: while command:
head = pathlib.PurePath(str(command[0])).name.lower() head = pathlib.PurePath(str(command[0])).name.lower()
if head in _SHELLS: if head in POSIX_SHELL_HEADS:
inline = shell_command_string(command) inline = shell_command_string(command)
if inline and sudo_noninteractive_violation(inline): if inline and sudo_noninteractive_violation(inline):
return True return True
@ -719,7 +719,7 @@ def shell_command_string(argv: List[str]) -> str:
def shell_argv_with_inline(raw_cmd: Any) -> List[str]: def shell_argv_with_inline(raw_cmd: Any) -> List[str]:
argv = shell_argv(raw_cmd) argv = shell_argv(raw_cmd)
if argv and pathlib.PurePath(argv[0]).name.lower() in _SHELLS: if argv and pathlib.PurePath(argv[0]).name.lower() in POSIX_SHELL_HEADS:
inline = shell_command_string(argv) inline = shell_command_string(argv)
if inline: if inline:
return argv + shell_argv(inline) return argv + shell_argv(inline)

View file

@ -128,10 +128,10 @@ BAND_PATHS = {
"ouroboros/gateway/host_service.py": "The one loopback callback boundary for reviewed skills: token auth, the chat/decision/presence/WS-relay routes and, with #667, the operation read/cancel that joins existing chat, routing, turn and task records; one trust boundary, one module.", "ouroboros/gateway/host_service.py": "The one loopback callback boundary for reviewed skills: token auth, the chat/decision/presence/WS-relay routes and, with #667, the operation read/cancel that joins existing chat, routing, turn and task records; one trust boundary, one module.",
"ouroboros/gateway/settings.py": "Retiring persistent auto-Low removed the former giant debt; the remaining owner and reviewer settings endpoints stay centralized while tracked in the shrinking band.", "ouroboros/gateway/settings.py": "Retiring persistent auto-Low removed the former giant debt; the remaining owner and reviewer settings endpoints stay centralized while tracked in the shrinking band.",
"ouroboros/gateways/claudexor.py": "The existing owned-engine gateway also owns typed model operations and exact-byte resource transfer; no second control client.", "ouroboros/gateways/claudexor.py": "The existing owned-engine gateway also owns typed model operations and exact-byte resource transfer; no second control client.",
"ouroboros/launcher_bootstrap.py": "Native seed version resync keeps manifest parsing and equal-version payload diagnostics with the existing bootstrap owner; no separate loader or overwrite policy.",
"ouroboros/loop_acceptance_review.py": "F6 upstream sync: the A-material acceptance family (paid identity, free replay, identical-refusal terminal, dialogue history) folded into the campaign review leaf per the sync principle (upstream leaf acceptance_dialogue.py retired)", "ouroboros/loop_acceptance_review.py": "F6 upstream sync: the A-material acceptance family (paid identity, free replay, identical-refusal terminal, dialogue history) folded into the campaign review leaf per the sync principle (upstream leaf acceptance_dialogue.py retired)",
"ouroboros/loop_delivery.py": "F6 upstream sync: the delivery-protocol upstream deltas (hold-control literals, trailing-object/fence-aware protocol parsers) folded into the campaign delivery leaf (upstream leaf delivery_protocol.py retired)", "ouroboros/loop_delivery.py": "F6 upstream sync: the delivery-protocol upstream deltas (hold-control literals, trailing-object/fence-aware protocol parsers) folded into the campaign delivery leaf (upstream leaf delivery_protocol.py retired)",
"ouroboros/loop_forced_finalization.py": "Forced-finalization rail of the v7 L-B loop split: one cohesive owner for the forced/orphan/absorption path, moved byte-preserving from loop.py (D01 lane).", "ouroboros/loop_forced_finalization.py": "Forced-finalization rail of the v7 L-B loop split: one cohesive owner for the forced/orphan/absorption path, moved byte-preserving from loop.py (D01 lane).",
"ouroboros/loop_tool_execution.py": None,
"ouroboros/marketplace/ouroboroshub.py": "Entered the band from 373 lines: the hubflow sprint added the adopt transaction (eligibility prelude, CAS re-verification, move-aside + state-quintet snapshot, verified rollback with per-step error collection, retention finalize) beside the existing install/update flows (hubflow sprint, adopt-in-ouroboroshub owner decision D4).", "ouroboros/marketplace/ouroboroshub.py": "Entered the band from 373 lines: the hubflow sprint added the adopt transaction (eligibility prelude, CAS re-verification, move-aside + state-quintet snapshot, verified rollback with per-step error collection, retention finalize) beside the existing install/update flows (hubflow sprint, adopt-in-ouroboroshub owner decision D4).",
"ouroboros/mcp_client.py": "F3.1 typed-organ producer cutover (D05 entry 7b) plus E5+s2r2 (#447): the MCP transport keeps the SDK-owned error bit as a typed ToolResult, follows nextCursor pagination with injective 12-hex slugs, and discloses collision/pagination omissions; grew into the band from 984 lines, shrink-only otherwise.", "ouroboros/mcp_client.py": "F3.1 typed-organ producer cutover (D05 entry 7b) plus E5+s2r2 (#447): the MCP transport keeps the SDK-owned error bit as a typed ToolResult, follows nextCursor pagination with injective 12-hex slugs, and discloses collision/pagination omissions; grew into the band from 984 lines, shrink-only otherwise.",
"ouroboros/memory.py": "ibl-2b09abdadd25: scratchpad content-size cap added alongside the existing block-count cap in append_scratchpad_block's eviction loop", "ouroboros/memory.py": "ibl-2b09abdadd25: scratchpad content-size cap added alongside the existing block-count cap in append_scratchpad_block's eviction loop",
@ -142,6 +142,7 @@ BAND_PATHS = {
"ouroboros/request_wire_recovery.py": "E4 (#447): typed CustomToolProjectionError fallback keeps the wire-recovery ladder alive; includes the one-site-sufficient decision record at both retry catch sites", "ouroboros/request_wire_recovery.py": "E4 (#447): typed CustomToolProjectionError fallback keeps the wire-recovery ladder alive; includes the one-site-sufficient decision record at both retry catch sites",
"ouroboros/review.py": "Entered the band from 952 lines: re-anchoring the size ratchet on the official line added the candidate and pairwise base-vs-tip transition validators (validate_size_ratchet_candidate/validate_size_ratchet_transition_against_base) with merge-aware previous-manifest resolution, replacing the retired first-parent history audit (update-flow-redesign sprint, Q7-C/Q18-A/Q19-A owner decisions).", "ouroboros/review.py": "Entered the band from 952 lines: re-anchoring the size ratchet on the official line added the candidate and pairwise base-vs-tip transition validators (validate_size_ratchet_candidate/validate_size_ratchet_transition_against_base) with merge-aware previous-manifest resolution, replacing the retired first-parent history audit (update-flow-redesign sprint, Q7-C/Q18-A/Q19-A owner decisions).",
"ouroboros/review_custody.py": "Review custody now owns the shared typed retry-rail history and frozen actor reconstruction so physical outcomes cannot be lost between the substrate and reconciliation.", "ouroboros/review_custody.py": "Review custody now owns the shared typed retry-rail history and frozen actor reconstruction so physical outcomes cannot be lost between the substrate and reconciliation.",
"ouroboros/review_evidence.py": "Commit-review visual source selection and exact source reuse belong with existing task evidence provenance and artifact materialization, including actual same-round attachments; no separate evidence store or routing policy.",
"ouroboros/review_evidence_sections.py": "The acceptance evidence assembler keeps individually materialized trajectory completeness and non-dict evidence disclosure beside final packet budgeting; the P3 and landed #736 paths share the same provenance decisions instead of a second evidence formatter.", "ouroboros/review_evidence_sections.py": "The acceptance evidence assembler keeps individually materialized trajectory completeness and non-dict evidence disclosure beside final packet budgeting; the P3 and landed #736 paths share the same provenance decisions instead of a second evidence formatter.",
"ouroboros/review_native_episode.py": "v7 follow-up F3 (owner Q6=A): a surface-declared mandatory reading became a typed FLOOR on the native episode bound (native_mandatory_read_bound / native_mandatory_read_disclosure / native_episode_transcript_bound \u2014 the one computation the advisory previews in its prompt's MANDATORY READ budget); the bound arithmetic stays with the episode that enforces it, 989->1062, no new subsystem, same seam; shrink-only direction.", "ouroboros/review_native_episode.py": "v7 follow-up F3 (owner Q6=A): a surface-declared mandatory reading became a typed FLOOR on the native episode bound (native_mandatory_read_bound / native_mandatory_read_disclosure / native_episode_transcript_bound \u2014 the one computation the advisory previews in its prompt's MANDATORY READ budget); the bound arithmetic stays with the episode that enforces it, 989->1062, no new subsystem, same seam; shrink-only direction.",
"ouroboros/reviewer_slot_config.py": "Absorbed the reviewer_slots() builder from review_substrate (altitude) and configured-subagent row resolution for the generic reviewer-actor bridge.", "ouroboros/reviewer_slot_config.py": "Absorbed the reviewer_slots() builder from review_substrate (altitude) and configured-subagent row resolution for the generic reviewer-actor bridge.",
@ -152,7 +153,6 @@ BAND_PATHS = {
"ouroboros/task_pacing.py": "Cost-ceiling SSOT grew into the band while absorbing the cache-aware wrap-up reservation, the prepared/prospective wrap-up candidates, the deciding-spend basis vocabulary and the exhausted-ceiling texts; one owner for pacing decisions instead of a second pricing authority beside usage_accounting.py (at its ceiling).", "ouroboros/task_pacing.py": "Cost-ceiling SSOT grew into the band while absorbing the cache-aware wrap-up reservation, the prepared/prospective wrap-up candidates, the deciding-spend basis vocabulary and the exhausted-ceiling texts; one owner for pacing decisions instead of a second pricing authority beside usage_accounting.py (at its ceiling).",
"ouroboros/task_status.py": None, "ouroboros/task_status.py": None,
"ouroboros/tools/browser.py": None, "ouroboros/tools/browser.py": None,
"ouroboros/tools/claude_advisory_review.py": "F2.3b advisory re-derive: the 2279-line GIANT parent re-enters the band at 1434 after the preflight_review_prompt/preflight_review_run split (admission policy, native episode, size gates and the tool entries stay with the facade)",
"ouroboros/tools/commit_gate.py": "Grew INTO the band by the review-wave fix binding the actor reference (delivery class) into the commit review contract fingerprint \u2014 same-module contract identity, splitting it would separate the fingerprint from its gate.", "ouroboros/tools/commit_gate.py": "Grew INTO the band by the review-wave fix binding the actor reference (delivery class) into the commit review contract fingerprint \u2014 same-module contract identity, splitting it would separate the fingerprint from its gate.",
"ouroboros/tools/core.py": "D05 ledger split (rows 311-349): read/list and owner-chat delivery spans moved to core_file_tools/core_artifacts; facade re-enters the band from above (2283 -> 1373) and shrinks further when the residual catalog split lands", "ouroboros/tools/core.py": "D05 ledger split (rows 311-349): read/list and owner-chat delivery spans moved to core_file_tools/core_artifacts; facade re-enters the band from above (2283 -> 1373) and shrinks further when the residual catalog split lands",
"ouroboros/tools/delegate.py": "D07 finisher DEL1 split brought the nanny-verb monolith DOWN from the 1600 hard cap into the band (1600->1263); terminal-evidence family extracted to tools/delegate_terminal_evidence.py, shrink-only direction", "ouroboros/tools/delegate.py": "D07 finisher DEL1 split brought the nanny-verb monolith DOWN from the 1600 hard cap into the band (1600->1263); terminal-evidence family extracted to tools/delegate_terminal_evidence.py, shrink-only direction",

View file

@ -38,10 +38,14 @@ def _build_skill_review_history_section(
if not history: if not history:
return "" return ""
lines = ["\n## Previous skill review attempts (anti-thrashing context)\n"] lines = ["\n## Previous skill review attempts (anti-thrashing context)\n"]
for idx, entry in enumerate(history[-3:], start=1): for entry in history[-3:]:
content_hash = str(entry.get("content_hash") or "")[:12] content_hash = str(entry.get("content_hash") or "")[:12]
status = entry.get("status", "?") status = entry.get("status", "?")
lines.append(f"### Attempt {idx}: status={status}, content_hash={content_hash}") ordinal = entry.get("review_round", "unknown")
snapshot = entry.get("snapshot_attempt", "unknown")
revised = entry.get("snapshot_revised", "unknown")
lines.append(f"### Review round {ordinal}, snapshot attempt {snapshot}: "
f"status={status}, content_hash={content_hash}, snapshot_revised={revised}")
fail_findings = entry.get("fail_findings") or [] fail_findings = entry.get("fail_findings") or []
if fail_findings: if fail_findings:
lines.append("FAIL findings (concrete reasons):") lines.append("FAIL findings (concrete reasons):")

View file

@ -166,7 +166,7 @@ def _advisory_child_timeout(ctx: object) -> Optional[float]:
def _run_advisory_native( def _run_advisory_native(
prompt: str, repo_dir: pathlib.Path, ctx: ToolContext, slot, model: str, prompt: str, repo_dir: pathlib.Path, ctx: ToolContext, slot, model: str,
mandatory_read_corpus_chars: int = 0, mandatory_read_corpus_chars: int = 0, task_evidence: Optional[dict] = None,
): ):
"""The advisory as a bounded native inspection episode, rehydrated into the """The advisory as a bounded native inspection episode, rehydrated into the
same result structure the retired SDK path produced (only the transport same result structure the retired SDK path produced (only the transport
@ -193,8 +193,12 @@ def _run_advisory_native(
str(_task_metadata.get("deadline_at") or "") str(_task_metadata.get("deadline_at") or "")
if isinstance(_task_metadata, dict) else "" if isinstance(_task_metadata, dict) else ""
) )
from ouroboros.review_evidence import commit_review_evidence_refs
evidence = task_evidence or {}
request = ReviewRequest( request = ReviewRequest(
surface="advisory_review", surface="advisory_review",
evidence={"task_execution": evidence} if evidence else {},
evidence_refs=commit_review_evidence_refs(evidence),
goal="Advisory pre-review of the live worktree.", goal="Advisory pre-review of the live worktree.",
task_id=str(getattr(ctx, "task_id", "") or ""), task_id=str(getattr(ctx, "task_id", "") or ""),
session_root=str(repo_dir), session_root=str(repo_dir),
@ -207,6 +211,8 @@ def _run_advisory_native(
no_proxy=True, no_proxy=True,
deadline_at=deadline_at, deadline_at=deadline_at,
) )
if evidence:
request.policy["native_data_root"] = evidence["data_root"]
# The dispatch builder for api_chat rows (`use_local` off the resolved # The dispatch builder for api_chat rows (`use_local` off the resolved
# route): the bound previewed below and the episode's window are ONE route. # route): the bound previewed below and the episode's window are ONE route.
rslot = _dc_replace( rslot = _dc_replace(
@ -1181,6 +1187,9 @@ def _handle_advisory_pre_review(
return counters return counters
return f"{changed_files.count(chr(10)) + 1} file(s) changed" return f"{changed_files.count(chr(10)) + 1} file(s) changed"
from ouroboros.review_evidence import capture_commit_review_evidence
task_evidence = (dict(getattr(ctx, "_commit_review_evidence", None) or {}) if prepared
else capture_commit_review_evidence(ctx)) if not resuming else {}
import time as _time import time as _time
_advisory_start = _time.monotonic() _advisory_start = _time.monotonic()
items, raw_result, model_used, prompt_chars = _run_claude_advisory( items, raw_result, model_used, prompt_chars = _run_claude_advisory(
@ -1191,7 +1200,8 @@ def _handle_advisory_pre_review(
scope=scope, scope=scope,
paths=paths, paths=paths,
options={"drive_root": drive_root, "review_rebuttal": review_rebuttal, options={"drive_root": drive_root, "review_rebuttal": review_rebuttal,
"execution": execution, "snapshot_hash": snapshot_hash}, "execution": execution, "snapshot_hash": snapshot_hash,
"task_evidence": task_evidence, "owns_task_evidence": not prepared},
) )
_advisory_duration = _time.monotonic() - _advisory_start _advisory_duration = _time.monotonic() - _advisory_start
advisory_meta = dict(getattr(ctx, "_last_claude_advisory_meta", {}) or {}) advisory_meta = dict(getattr(ctx, "_last_claude_advisory_meta", {}) or {})

View file

@ -38,7 +38,8 @@ def get_tools() -> List[ToolEntry]:
"name": "schedule_followup", "name": "schedule_followup",
"description": ( "description": (
"Register a deferred follow-up task that the supervisor scheduler " "Register a deferred follow-up task that the supervisor scheduler "
"enqueues as an ordinary root task. Supply exactly one trigger: run_at " "enqueues as an ordinary root task. Root tasks only; subagents report "
"the proposed follow-up to their parent. Supply exactly one trigger: run_at "
"for a one-shot ISO 8601 instant (naive = UTC), or cron for a recurring " "for a one-shot ISO 8601 instant (naive = UTC), or cron for a recurring "
"5-field expression with an optional IANA timezone. Write the objective " "5-field expression with an optional IANA timezone. Write the objective "
"in your own words — it becomes each future task's text verbatim. The " "in your own words — it becomes each future task's text verbatim. The "

View file

@ -70,6 +70,14 @@ def _review_custody_pending(ctx: ToolContext) -> bool:
) )
def _release_review_evidence_if_settled(ctx: ToolContext) -> None:
"""The existing review-custody boundary owns temporary session-file lifetime."""
from ouroboros.review_evidence import release_commit_review_session_view
if not _review_custody_pending(ctx):
release_commit_review_session_view(getattr(ctx, "_commit_review_evidence", None) or {})
def _fingerprint_staged_diff(repo_dir: pathlib.Path) -> Dict[str, Any]: def _fingerprint_staged_diff(repo_dir: pathlib.Path) -> Dict[str, Any]:
"""Bind review to the exact commit material, not only a textual diff. """Bind review to the exact commit material, not only a textual diff.
@ -710,6 +718,12 @@ def _run_reviewed_stage_cycle(
from ouroboros.review_state import compute_snapshot_hash from ouroboros.review_state import compute_snapshot_hash
prepared_snapshot = compute_snapshot_hash(pathlib.Path(ctx.repo_dir), commit_message, paths=advisory_paths) prepared_snapshot = compute_snapshot_hash(pathlib.Path(ctx.repo_dir), commit_message, paths=advisory_paths)
from ouroboros.review_evidence import capture_commit_review_evidence, pending_commit_review_evidence
if not getattr(ctx, "_advisory_reconciled", False):
ctx._commit_review_evidence = (
pending_commit_review_evidence(ctx) if getattr(ctx, "_review_reconcile_only", False)
else capture_commit_review_evidence(ctx) if advisory_replay is None else {})
advisory_gate_outcome = None advisory_gate_outcome = None
if not bool(getattr(ctx, "_review_reconcile_only", False)): if not bool(getattr(ctx, "_review_reconcile_only", False)):
advisory_gate_outcome = _git()._advisory_and_tests_gate( advisory_gate_outcome = _git()._advisory_and_tests_gate(
@ -723,6 +737,7 @@ def _run_reviewed_stage_cycle(
goal=goal, scope=scope, goal=goal, scope=scope,
) )
if advisory_gate_outcome is not None: if advisory_gate_outcome is not None:
_release_review_evidence_if_settled(ctx)
return advisory_gate_outcome return advisory_gate_outcome
if not bool(getattr(ctx, "_review_reconcile_only", False)): if not bool(getattr(ctx, "_review_reconcile_only", False)):
after_preflight = _git()._fingerprint_staged_diff(pathlib.Path(ctx.repo_dir)) after_preflight = _git()._fingerprint_staged_diff(pathlib.Path(ctx.repo_dir))
@ -731,6 +746,7 @@ def _run_reviewed_stage_cycle(
ctx, commit_message, commit_start, pre_fingerprint, after_preflight, worktree_changed=changed, ctx, commit_message, commit_start, pre_fingerprint, after_preflight, worktree_changed=changed,
) )
if revalidation is not None: if revalidation is not None:
_release_review_evidence_if_settled(ctx)
return revalidation return revalidation
_git()._record_commit_attempt( _git()._record_commit_attempt(
ctx, ctx,
@ -794,6 +810,7 @@ def _run_reviewed_stage_cycle(
) )
finally: finally:
_git()._reconcile_and_clear_review_roster(ctx) _git()._reconcile_and_clear_review_roster(ctx)
_release_review_evidence_if_settled(ctx)
blocked, combined_msg, block_reason, combined_findings, scope_advisory = _git()._aggregate_review_verdict( blocked, combined_msg, block_reason, combined_findings, scope_advisory = _git()._aggregate_review_verdict(
review_err, review_err,
scope_result, scope_result,

View file

@ -94,12 +94,16 @@ def _actor_outcome(actor: dict) -> str:
+ ": " + str(actor.get("error"))) + ": " + str(actor.get("error")))
def _degraded_replay_note(wave: dict) -> str: def _degraded_replay_note(wave: dict, *, paid_available: bool = True) -> str:
"""The honest replay mechanics of one recorded DEGRADED wave (aligned with the """The honest replay mechanics of one recorded DEGRADED wave (aligned with the
engine's `plan_wave_replay_decision`): a wave with structural snapshot evidence engine's `plan_wave_replay_decision`): a wave with structural snapshot evidence
replays while its epoch and the reviewer roster stand; one without (its slots replays while its epoch and the reviewer roster stand; one without (its slots
died at dispatch time, invisible to the pre-fan-out snapshot) never replays — died at dispatch time, invisible to the pre-fan-out snapshot) never replays —
a transient death is never cached as structural.""" a transient death is never cached as structural."""
if not paid_available:
replay = ("an identical envelope can replay this result for free while its health epoch and roster stand; "
if wave.get("health_epoch") else "no structural lane evidence was recorded; ")
return replay + "the cycle cap is reached, so neither an identical nor revised request can start another paid panel, even after lane recovery"
if wave.get("health_epoch"): if wave.get("health_epoch"):
return ( return (
"an identical envelope replays this recorded result at no further cost while " "an identical envelope replays this recorded result at no further cost while "
@ -142,7 +146,8 @@ def _next_step(wave: dict, *, enforcement: str, cap: Optional[int], cycles_paid:
f"{counts.get('configured', 0)} configured slot(s) — below the review quorum " f"{counts.get('configured', 0)} configured slot(s) — below the review quorum "
f"({counts.get('quorum', '?')}). Per-slot typed states (code and reset time, when " f"({counts.get('quorum', '?')}). Per-slot typed states (code and reset time, when "
"known) are listed under Reviewer slots above. This wave is recorded and OPEN; " "known) are listed under Reviewer slots above. This wave is recorded and OPEN; "
f"{_degraded_replay_note(wave)}; a changed spec starts the next paid cycle. " f"{_degraded_replay_note(wave, paid_available=not at_cap)}. "
+ ("A changed spec may start another paid cycle. " if not at_cap else "")
) )
if wave.get("quorum_unreachable"): if wave.get("quorum_unreachable"):
# Naming asymmetry, on purpose: the wave fact is the bare # Naming asymmetry, on purpose: the wave fact is the bare
@ -169,16 +174,15 @@ def _next_step(wave: dict, *, enforcement: str, cap: Optional[int], cycles_paid:
ids = ", ".join(str(f.get("finding_id") or f.get("id")) for f in blocking[:4]) ids = ", ".join(str(f.get("finding_id") or f.get("id")) for f in blocking[:4])
text += ( text += (
f"NOTE: {len(blocking)} BLOCKING finding(s) below quorum ({ids}) stay OPEN whatever " f"NOTE: {len(blocking)} BLOCKING finding(s) below quorum ({ids}) stay OPEN whatever "
"you disposition — a blocking finding closes only through a changed spec " "you disposition. "
"(new fingerprint, next paid cycle) or a reject that the next paid delta cycle judges. " + ("A changed spec or a justified rejection may be judged in another paid cycle. "
if not at_cap else "The cycle cap is reached; no further paid panel is available. ")
) )
else: else:
text = ( text = (
"Blocking findings: accept ⇒ change the spec and re-call plan_task (new fingerprint, " "Blocking findings remain OPEN. A disposition records your rationale without closing REVISE_PLAN. "
f"{'the cap is reached — no further paid cycle' if at_cap else 'next paid cycle ' + str(cycles_paid + 1) + ('' if cap is None else f' of {cap}')}); " + ("The cycle cap is reached; no further paid panel is available. " if at_cap else
"reject ⇒ record reject + rationale via review_disposition naming this fingerprint — it " "You may change the spec or record a justified rejection for a subsequent paid delta review. ")
"rides into the next paid delta cycle where reviewers mark it resolved or still-open. "
"A disposition never closes REVISE_PLAN. "
) )
if enforcement == "blocking": if enforcement == "blocking":
text += ( text += (
@ -205,6 +209,17 @@ def _next_step(wave: dict, *, enforcement: str, cap: Optional[int], cycles_paid:
def _closure_note_view(note: str) -> str:
"""Legacy host notes describe state; the current renderer owns available steps."""
prefix = str(note).partition(":")[0]
meaning = {
"blocking_finding_below_quorum_stays_open": "blocking findings remain open after disposition",
"revise_plan_not_closable_by_disposition": "disposition does not close blocking findings",
"degraded_not_closable_by_disposition": "no parseable reviewer quorum; disposition does not close the wave",
}.get(prefix)
return f"{prefix}: {meaning}" if meaning else str(note)
def _render_wave( def _render_wave(
wave: dict, *, cap: Optional[int], cycles_paid: int, enforcement: str, wave: dict, *, cap: Optional[int], cycles_paid: int, enforcement: str,
cached: bool = False, notes: Optional[List[str]] = None, reminder: str = "", cached: bool = False, notes: Optional[List[str]] = None, reminder: str = "",
@ -244,7 +259,7 @@ def _render_wave(
# Banner aligned with _next_step: the replay promise depends on whether the # Banner aligned with _next_step: the replay promise depends on whether the
# wave carries structural snapshot evidence (see _degraded_replay_note). # wave carries structural snapshot evidence (see _degraded_replay_note).
lines += ["", "⚠️ DEGRADED: no parseable reviewer quorum — recorded as an OPEN wave; " lines += ["", "⚠️ DEGRADED: no parseable reviewer quorum — recorded as an OPEN wave; "
+ _degraded_replay_note(wave) + "."] + _degraded_replay_note(wave, paid_available=cap is None or cycles_paid < cap) + "."]
actor_lines = [ actor_lines = [
f"- {a.get('slot_id')} · {a.get('model')} · {a.get('route')} · host_file_read: " f"- {a.get('slot_id')} · {a.get('model')} · {a.get('route')} · host_file_read: "
f"{a.get('host_file_read_attestation')} · {_actor_outcome(a)}" f"{a.get('host_file_read_attestation')} · {_actor_outcome(a)}"
@ -281,7 +296,7 @@ def _render_wave(
lines += ["", "### Dispositions", "", "```json", lines += ["", "### Dispositions", "", "```json",
json.dumps(wave.get("dispositions"), ensure_ascii=False, indent=2), "```"] json.dumps(wave.get("dispositions"), ensure_ascii=False, indent=2), "```"]
if wave.get("closure_notes") or notes: if wave.get("closure_notes") or notes:
lines += ["", "Closure notes: " + "; ".join([*(wave.get("closure_notes") or []), *(notes or [])])] lines += ["", "Closure notes: " + "; ".join(_closure_note_view(note) for note in [*(wave.get("closure_notes") or []), *(notes or [])])]
outcome, closed = wave_control_state(wave) outcome, closed = wave_control_state(wave)
lines += [ lines += [
"", "## Plan Review Contract", "", "", "## Plan Review Contract", "",

View file

@ -18,8 +18,8 @@ replays the recorded wave free (no panel, no cycle). Closure
(``plan_spec.closure_after_disposition``): GREEN closes; (``plan_spec.closure_after_disposition``): GREEN closes;
Note-only REVIEW_REQUIRED closes immediately; need_evidence closes by disposition Note-only REVIEW_REQUIRED closes immediately; need_evidence closes by disposition
at $0; a below-quorum blocking finding stays open. REVISE_PLAN never closes by at $0; a below-quorum blocking finding stays open. REVISE_PLAN never closes by
disposition — accept ⇒ changed spec (next paid cycle), reject ⇒ rationale rides disposition. A subsequent paid delta review may evaluate a changed spec or a
into the next delta cycle. Under blocking enforcement an open wave HOLDS justified rejection when another paid cycle is available. Under blocking enforcement an open wave HOLDS
finalization (``owner_hurry.force_plan_decision``); at the cap the typed finalization (``owner_hurry.force_plan_decision``); at the cap the typed
``plan_review_cycles_exhausted`` result + event leave the honest exits: owner ``plan_review_cycles_exhausted`` result + event leave the honest exits: owner
unstick or a ``blocked_with_evidence`` terminal. Advisory proceeds open under the unstick or a ``blocked_with_evidence`` terminal. Advisory proceeds open under the
@ -205,8 +205,9 @@ _DISPOSITION_SCHEMA = {
"description": ( "description": (
"Disposition mode only (send ONLY this field): answer the findings of the wave " "Disposition mode only (send ONLY this field): answer the findings of the wave "
"named by review_fingerprint. note/need_evidence findings close at $0; a blocking " "named by review_fingerprint. note/need_evidence findings close at $0; a blocking "
"finding you accept needs a changed spec (new call), one you reject rides with your " "finding stays open. A subsequent paid delta review may consider a changed spec or "
"rationale into the next paid cycle. Never closes REVISE_PLAN." "justified rejection when another paid cycle is available. Recording a disposition "
"consumes no cycle and never closes REVISE_PLAN."
), ),
"properties": { "properties": {
"review_fingerprint": {"type": "string"}, "review_fingerprint": {"type": "string"},
@ -240,8 +241,8 @@ def get_tools():
"reviewers return typed findings against the spec (blocking findings must name " "reviewers return typed findings against the spec (blocking findings must name "
"the spec element they break); the host aggregates: GREEN closes; " "the spec element they break); the host aggregates: GREEN closes; "
"Notes are optional; need_evidence closes by review_disposition at no cost; REVISE_PLAN needs " "Notes are optional; need_evidence closes by review_disposition at no cost; REVISE_PLAN needs "
"a changed spec (next paid cycle) or a reject-with-rationale judged in the next " "a changed spec or justified rejection judged by a subsequent paid delta review "
"cycle. Cycles are bounded by the owner's Max review cycles; an unchanged " "when another paid cycle is available. Cycles are bounded by the owner's Max review cycles; an unchanged "
"envelope replays the recorded result for free (a locator a reviewer asked for " "envelope replays the recorded result for free (a locator a reviewer asked for "
"with need_evidence is attached by the host next time and makes the envelope " "with need_evidence is attached by the host next time and makes the envelope "
"new). Under blocking enforcement an " "new). Under blocking enforcement an "

View file

@ -787,9 +787,9 @@ def closure_after_disposition(
GREEN → closed. Notes are optional advice, so a note-only REVIEW_REQUIRED GREEN → closed. Notes are optional advice, so a note-only REVIEW_REQUIRED
wave closes without dispositions. Need_evidence still requires a disposition wave closes without dispositions. Need_evidence still requires a disposition
(accept|reject|defer + rationale). REVISE_PLAN → NEVER closed by (accept|reject|defer + rationale). REVISE_PLAN → NEVER closed by
disposition (blocking needs a changed spec → new cycle, or reject-with- disposition. A subsequent paid delta review may consider a changed spec or
rationale → next paid delta cycle). DEGRADED → not closable by disposition justified rejection when another paid cycle is available. DEGRADED is not
(rerun the wave). Advisory enforcement never flips ``closed``: the caller closable by disposition. Advisory enforcement never flips ``closed``: the caller
may proceed with the wave open under loud disclosure — this function only may proceed with the wave open under loud disclosure — this function only
reports. Control-line invariants (``tools.plan_render reports. Control-line invariants (``tools.plan_render
._parse_plan_review_control``): GREEN ⇒ closed, REVISE_PLAN ⇒ not closed. ._parse_plan_review_control``): GREEN ⇒ closed, REVISE_PLAN ⇒ not closed.
@ -839,18 +839,16 @@ def closure_after_disposition(
notes.append("no_findings_recorded: REVIEW_REQUIRED without findings closes vacuously") notes.append("no_findings_recorded: REVIEW_REQUIRED without findings closes vacuously")
if any(f.get("class") == "blocking" for f in items): if any(f.get("class") == "blocking" for f in items):
notes.append( notes.append(
"blocking_finding_below_quorum_stays_open: revise the spec or let the next " "blocking_finding_below_quorum_stays_open: blocking findings remain open after disposition"
"paid delta cycle judge the rejection"
) )
elif verdict == "REVISE_PLAN": elif verdict == "REVISE_PLAN":
closed = False closed = False
notes.append( notes.append(
"revise_plan_not_closable_by_disposition: blocking findings need a changed spec " "revise_plan_not_closable_by_disposition: disposition does not close blocking findings"
"(new cycle) or reject-with-rationale judged in the next paid delta cycle"
) )
elif verdict == "DEGRADED": elif verdict == "DEGRADED":
closed = False closed = False
notes.append("degraded_not_closable_by_disposition: fewer parseable reviewer slots than quorum — rerun the wave") notes.append("degraded_not_closable_by_disposition: fewer parseable reviewer slots than quorum; disposition does not close the wave")
else: else:
closed = False closed = False
notes.append(f"unknown_aggregate:{verdict or '<empty>'}") notes.append(f"unknown_aggregate:{verdict or '<empty>'}")

View file

@ -235,9 +235,9 @@ def _build_advisory_prompt(
# precedent (plan_review_runtime's retrieving-session task and its # precedent (plan_review_runtime's retrieving-session task and its
# DEVELOPMENT.md "Core Governance Artifacts" row), NOT BIBLE P3 # DEVELOPMENT.md "Core Governance Artifacts" row), NOT BIBLE P3
# retrieving-scope. The advisory session pack deliberately contains # retrieving-scope. The advisory session pack deliberately contains
# only the staged diff, the changed-file pack, and PUBLIC repository # the staged diff, changed-file pack and public repository documents.
# documents — no redacted-class evidence — so the pointer form leaks # Selected task execution evidence is separately redacted and bound to
# nothing the api form redacts. # its canonical source before either retrieving delivery runs.
bible = _car()._mandatory_read_pointer(repo_dir, "BIBLE.md") bible = _car()._mandatory_read_pointer(repo_dir, "BIBLE.md")
checklists = _car()._mandatory_read_pointer(repo_dir, "docs/CHECKLISTS.md", section=checklist_name) checklists = _car()._mandatory_read_pointer(repo_dir, "docs/CHECKLISTS.md", section=checklist_name)
dev_guide = _car()._mandatory_read_pointer(repo_dir, "docs/DEVELOPMENT.md") dev_guide = _car()._mandatory_read_pointer(repo_dir, "docs/DEVELOPMENT.md")
@ -361,6 +361,7 @@ def _build_advisory_prompt(
"## ARCHITECTURE.md (System structure — critical for version sync and module checks)\n\n" "## ARCHITECTURE.md (System structure — critical for version sync and module checks)\n\n"
f"{arch_doc}\n\n{skill_host_context}\n\n{blocking_history}\n\n" f"{arch_doc}\n\n{skill_host_context}\n\n{blocking_history}\n\n"
f"{build_rebuttal_section(str(prompt_context.get('review_rebuttal') or ''))}\n" f"{build_rebuttal_section(str(prompt_context.get('review_rebuttal') or ''))}\n"
f"{prompt_context.get('task_evidence_section') or ''}\n"
f"## Commit message\n\n{commit_message}\n\n" f"## Commit message\n\n{commit_message}\n\n"
f"## Changed files (git status --porcelain)\n\n{changed_files}\n\n" f"## Changed files (git status --porcelain)\n\n{changed_files}\n\n"
"## Current touched files (full content — read these with read_file for deeper inspection)\n\n" "## Current touched files (full content — read these with read_file for deeper inspection)\n\n"

View file

@ -401,7 +401,7 @@ def _checkpoint_advisory_execution(ctx, repo_dir, commit_message, paths, options
) from exc ) from exc
def _run_advisory_delegated(prompt: str, repo_dir: pathlib.Path, ctx: ToolContext, *, execution=None, checkpoint=None): def _run_advisory_delegated(prompt: str, repo_dir: pathlib.Path, ctx: ToolContext, *, execution=None, checkpoint=None, task_evidence=None):
"""The advisory as a delegated agent session on the SHARED executor seam. """The advisory as a delegated agent session on the SHARED executor seam.
One substrate executor (``AgentSessionReviewExecutor``) owns the session: One substrate executor (``AgentSessionReviewExecutor``) owns the session:
@ -430,8 +430,12 @@ def _run_advisory_delegated(prompt: str, repo_dir: pathlib.Path, ctx: ToolContex
str(_task_metadata.get("deadline_at") or "") str(_task_metadata.get("deadline_at") or "")
if isinstance(_task_metadata, dict) else "" if isinstance(_task_metadata, dict) else ""
) )
from ouroboros.review_evidence import commit_review_evidence_refs
evidence = task_evidence or {}
request = ReviewRequest( request = ReviewRequest(
surface="advisory_review", surface="advisory_review",
evidence={"task_execution": evidence} if evidence else {},
evidence_refs=commit_review_evidence_refs(evidence),
goal="Advisory pre-review of the live worktree.", goal="Advisory pre-review of the live worktree.",
task_id=str(getattr(ctx, "task_id", "") or ""), task_id=str(getattr(ctx, "task_id", "") or ""),
session_root=str(repo_dir), session_root=str(repo_dir),
@ -504,6 +508,26 @@ def _note_meta_error(ctx: ToolContext, meta: dict, err_msg: str) -> None:
pass pass
def _prepare_advisory_task_evidence(ctx, repo_dir, options, execution, delegated_route):
"""Bind the original readable view before either advisory delivery sends.
Pending custody selects its recorded canonical source. Only hosted sessions
materialize that source inside the project; native reads keep the data root.
"""
from ouroboros.review_evidence import materialize_commit_review_session_view, restore_commit_review_evidence
evidence = dict(options.get("task_evidence") or {})
if execution.get("pending_invocation_id"):
evidence = restore_commit_review_evidence(ctx, execution["evidence_source_ref"]) if execution.get("evidence_source_ref") else {}
if delegated_route and evidence:
evidence = materialize_commit_review_session_view(evidence, repo_dir)
if options.get("owns_task_evidence") is False:
ctx._commit_review_evidence = evidence
if evidence.get("source_ref"):
execution["evidence_source_ref"] = evidence["source_ref"]
return evidence
def _run_claude_advisory( def _run_claude_advisory(
repo_dir: pathlib.Path, repo_dir: pathlib.Path,
commit_message: str, commit_message: str,
@ -558,7 +582,13 @@ def _run_claude_advisory(
_note_meta_error(ctx, {"execution": execution}, message) _note_meta_error(ctx, {"execution": execution}, message)
return [], message, "", 0 return [], message, "", 0
from ouroboros.review_evidence import commit_review_evidence_section
resuming = bool(execution.get("pending_invocation_id")) resuming = bool(execution.get("pending_invocation_id"))
try:
task_evidence = _prepare_advisory_task_evidence(ctx, repo_dir, options, execution, delegated_route)
except (OSError, ValueError, TypeError) as exc:
return assembly_failure(f"⚠️ ADVISORY_ERROR: frozen task evidence is unavailable: {type(exc).__name__}")
task_evidence_section = commit_review_evidence_section(task_evidence, delivery="session" if delegated_route else "native")
if resuming: if resuming:
from ouroboros.delegate_custody import custody_root, invocation_record from ouroboros.delegate_custody import custody_root, invocation_record
@ -615,6 +645,7 @@ def _run_claude_advisory(
"review_surface": review_surface, "review_surface": review_surface,
"review_rebuttal": str(options.get("review_rebuttal") or ""), "review_rebuttal": str(options.get("review_rebuttal") or ""),
"expected_items": expected_items, "expected_items": expected_items,
"task_evidence_section": task_evidence_section,
}, },
# Both deliveries RETRIEVE governance docs via mandatory-read # Both deliveries RETRIEVE governance docs via mandatory-read
# pointers (the session with its own tools, the native episode with # pointers (the session with its own tools, the native episode with
@ -625,6 +656,12 @@ def _run_claude_advisory(
except Exception as exc: except Exception as exc:
return assembly_failure(f"⚠️ ADVISORY_ERROR: failed to build advisory prompt: {exc}") return assembly_failure(f"⚠️ ADVISORY_ERROR: failed to build advisory prompt: {exc}")
if not resuming and task_evidence_section:
oversized = (len(prompt) > _ADVISORY_PROMPT_MAX_CHARS or
(not delegated_route and _car()._api_window_skip_warning(model, prompt, managed_subject_diff, slot=_slot)))
if oversized:
prompt = prompt.replace(task_evidence_section, commit_review_evidence_section(
task_evidence, delivery="session" if delegated_route else "native", compact=True), 1)
prompt_chars = len(prompt) prompt_chars = len(prompt)
diag = _car()._get_runtime_diagnostics(model, prompt_chars, resolved_paths) diag = _car()._get_runtime_diagnostics(model, prompt_chars, resolved_paths)
size_skip = None if resuming else _car()._predispatch_size_skip( size_skip = None if resuming else _car()._predispatch_size_skip(
@ -641,22 +678,17 @@ def _run_claude_advisory(
try: try:
if delegated_route: if delegated_route:
# 5.8: only the transport changes — the delegated session runs the # Same prompt/root/result contract; session timing and cost belong
# SAME advisory prompt in the same repo root and rehydrates the same # to the nanny and delegate_custody, not the retired SDK budget.
# result structure. The SDK budget kill is replaced by the runner's
# nanny-enforced time cap; cost settles through delegate_custody.
scope_effort = "" # the session route carries its own effort scope_effort = "" # the session route carries its own effort
custody_args = ({"execution": execution, "checkpoint": checkpoint} custody_args = ({"execution": execution, "checkpoint": checkpoint}
if options.get("snapshot_hash") else {}) if options.get("snapshot_hash") else {})
result, model = _car()._run_advisory_delegated(prompt, repo_dir, ctx, **custody_args) result, model = _car()._run_advisory_delegated(prompt, repo_dir, ctx, **custody_args,
**({"task_evidence": task_evidence} if task_evidence else {}))
else: else:
# The native inspection episode (the retired Claude-SDK # Same prompt/root/result contract; native inspection has no round
# transport's successor): same prompt, same repo root, same result # cap, only its reviewer-window transcript bound. Each provider call
# structure. The SDK budget kill is replaced by the episode's # rides the ordinary advisory_review usage ledger.
# transcript bound derived from THIS reviewer's own window
# (``review_native_episode.review_native_transcript_bound``) — no
# round cap; every provider call rides the ordinary usage ledger
# under category=advisory_review.
scope_effort = _slot.effort or "low" scope_effort = _slot.effort or "low"
if _car().owner_deadline_exhausted_for_context(ctx, reserve_sec=_car().get_finalization_grace_sec()): if _car().owner_deadline_exhausted_for_context(ctx, reserve_sec=_car().get_finalization_grace_sec()):
raise TimeoutError("owner deadline leaves no dispatch window for advisory review") raise TimeoutError("owner deadline leaves no dispatch window for advisory review")
@ -667,6 +699,7 @@ def _run_claude_advisory(
result, model = _car()._run_advisory_native( result, model = _car()._run_advisory_native(
prompt, repo_dir, ctx, _slot, model, prompt, repo_dir, ctx, _slot, model,
mandatory_read_corpus_chars=_car()._mandatory_read_corpus_chars(repo_dir, review_surface), mandatory_read_corpus_chars=_car()._mandatory_read_corpus_chars(repo_dir, review_surface),
**({"task_evidence": task_evidence} if task_evidence else {}),
) )
usage = dict(getattr(result, "usage", {}) or {}) usage = dict(getattr(result, "usage", {}) or {})
@ -719,6 +752,8 @@ def _run_claude_advisory(
return [], err_msg + "\n\nReviewer output:\n" + raw_received, model, prompt_chars return [], err_msg + "\n\nReviewer output:\n" + raw_received, model, prompt_chars
raw_text = str(result.result_text or "") raw_text = str(result.result_text or "")
event_facts = {key: meta[key] for key in (
"model", "session_id", "prompt_chars", "cost_usd", "review_surface")}
if raw_text.strip() in {"", "(no output)"}: if raw_text.strip() in {"", "(no output)"}:
err_msg = _car()._format_advisory_error( err_msg = _car()._format_advisory_error(
@ -730,12 +765,8 @@ def _run_claude_advisory(
) )
_car().emit_review_event(ctx, { _car().emit_review_event(ctx, {
"type": "advisory_suspect_result", "type": "advisory_suspect_result",
"model": model, **event_facts,
"session_id": meta.get("session_id", ""),
"prompt_chars": prompt_chars,
"cost_usd": float(result.cost_usd or 0),
"reason": "advisory result had empty output", "reason": "advisory result had empty output",
"review_surface": review_surface,
}) })
execution.update(failure_phase="format", failure_code="empty_response", source_text=original) execution.update(failure_phase="format", failure_code="empty_response", source_text=original)
_note_meta_error(ctx, meta, err_msg) _note_meta_error(ctx, meta, err_msg)
@ -759,12 +790,8 @@ def _run_claude_advisory(
) )
_car().emit_review_event(ctx, { _car().emit_review_event(ctx, {
"type": "advisory_suspect_result", "type": "advisory_suspect_result",
"model": model, **event_facts,
"session_id": meta.get("session_id", ""),
"prompt_chars": prompt_chars,
"cost_usd": float(result.cost_usd or 0),
"reason": contract_error, "reason": contract_error,
"review_surface": review_surface,
}) })
execution.update(failure_phase="format", failure_code="checklist_contract", source_text=original) execution.update(failure_phase="format", failure_code="checklist_contract", source_text=original)
_note_meta_error(ctx, meta, err_msg) _note_meta_error(ctx, meta, err_msg)
@ -773,12 +800,8 @@ def _run_claude_advisory(
if contract_warning: if contract_warning:
_car().emit_review_event(ctx, { _car().emit_review_event(ctx, {
"type": "advisory_contract_warning", "type": "advisory_contract_warning",
"model": model, **event_facts,
"session_id": meta.get("session_id", ""),
"prompt_chars": prompt_chars,
"cost_usd": float(result.cost_usd or 0),
"warning": contract_warning, "warning": contract_warning,
"review_surface": review_surface,
}) })
try: try:
meta["status"] = "completed_with_contract_warning" meta["status"] = "completed_with_contract_warning"

View file

@ -449,6 +449,7 @@ _REVIEW_PROMPT_TEMPLATE_DYNAMIC = """\
{changed_files} {changed_files}
{rebuttal_section}{review_history_section} {rebuttal_section}{review_history_section}
{task_evidence_section}
""" """
@ -1070,6 +1071,13 @@ def _prepare_unified_review(ctx: ToolContext, commit_message: str,
session_profile=row_plan["session_profiles"][i], use_local=row_plan["use_local"][i]) session_profile=row_plan["session_profiles"][i], use_local=row_plan["use_local"][i])
for i in api_indices] for i in api_indices]
from ouroboros.review_evidence import commit_review_evidence_section, materialize_commit_review_session_view
task_evidence = dict(getattr(ctx, "_commit_review_evidence", None) or {})
if any(route is ReviewRouteKind.AGENT_SESSION for route in row_routes):
task_evidence = materialize_commit_review_session_view(task_evidence, target_repo)
ctx._commit_review_evidence = task_evidence
task_evidence_compact = False
goal_section = build_goal_section(goal, scope, commit_message) goal_section = build_goal_section(goal, scope, commit_message)
scope_section = build_scope_section(scope) scope_section = build_scope_section(scope)
@ -1094,9 +1102,16 @@ def _prepare_unified_review(ctx: ToolContext, commit_message: str,
review_history_section=review_history_section, review_history_section=review_history_section,
diff_text=staged_diff, diff_text=staged_diff,
changed_files=review_changed, changed_files=review_changed,
task_evidence_section=commit_review_evidence_section(task_evidence, delivery="packet", compact=task_evidence_compact),
) )
return stable + "\n" + dynamic, len(stable) + 1 return stable + "\n" + dynamic, len(stable) + 1
def _compact_task_evidence():
nonlocal task_evidence_compact
task_evidence_compact = True
if task_evidence:
_assemble_prompt.compact_optional_evidence = _compact_task_evidence
# P3 stays one-pass. The api pack, its fit ladder and the fixed_overflow # P3 stays one-pass. The api pack, its fit ladder and the fixed_overflow
# gate exist ONLY for the api rows (5.2/5.7): a session row retrieves with # gate exist ONLY for the api rows (5.2/5.7): a session row retrieves with
# its own tools, so it neither constrains the fit limit nor is blocked by # its own tools, so it neither constrains the fit limit nor is blocked by
@ -1168,7 +1183,7 @@ def _prepare_unified_review(ctx: ToolContext, commit_message: str,
"prompt": prompt, "stable_prefix_len": stable_prefix_len, "prompt": prompt, "stable_prefix_len": stable_prefix_len,
"models": models, "routes": row_routes, "row_plan": row_plan, "models": models, "routes": row_routes, "row_plan": row_plan,
"session_task": session_task, "target_repo": target_repo, "session_task": session_task, "target_repo": target_repo,
"blocking_review": blocking_review, "blocking_review": blocking_review, "task_evidence": task_evidence,
}, None, False }, None, False
@ -1191,6 +1206,7 @@ def _dispatch_unified_review(ctx: ToolContext, commit_message: str, prepared: di
session_root=str(prepared["target_repo"]), session_root=str(prepared["target_repo"]),
row_plan=prepared["row_plan"], row_plan=prepared["row_plan"],
retry_key=str(prepared.get("retry_key") or ""), retry_key=str(prepared.get("retry_key") or ""),
task_evidence=prepared.get("task_evidence"),
) )
result = json.loads(result_json) result = json.loads(result_json)
except Exception as e: except Exception as e:

View file

@ -196,6 +196,9 @@ def fit_triad_prompt(api_models: list, assemble, current_files_section: str,
slot_limits = {key: _slot_input_limit(index) for index, key in enumerate(keys)} slot_limits = {key: _slot_input_limit(index) for index, key in enumerate(keys)}
input_limit = _rv._quorum_input_token_limit(keys, slot_limits) input_limit = _rv._quorum_input_token_limit(keys, slot_limits)
prompt, stable_prefix_len = assemble(current_files_section, diff_text) prompt, stable_prefix_len = assemble(current_files_section, diff_text)
if input_limit and estimate_tokens(prompt) > input_limit and getattr(assemble, "compact_optional_evidence", None):
assemble.compact_optional_evidence()
prompt, stable_prefix_len = assemble(current_files_section, diff_text)
if input_limit and estimate_tokens(prompt) > input_limit: if input_limit and estimate_tokens(prompt) > input_limit:
# Cold-start density rung: every api slot the full prompt overflows and # Cold-start density rung: every api slot the full prompt overflows and
# whose route has no fresh witness gets ONE bounded probe on a slice of # whose route has no fresh witness gets ONE bounded probe on a slice of
@ -403,6 +406,14 @@ def prepare_scope_review(
# RETRIEVES class: a session row and a configured-subagent api row deliver # RETRIEVES class: a session row and a configured-subagent api row deliver
# by retrieval — neither assembles the packet/atlas below. # by retrieval — neither assembles the packet/atlas below.
retrieves = delivery_retrieves(route, subagent_id) retrieves = delivery_retrieves(route, subagent_id)
from ouroboros.review_evidence import commit_review_evidence_section, materialize_commit_review_session_view
task_evidence = dict(getattr(ctx, "_commit_review_evidence", None) or {})
if delegated:
task_evidence = materialize_commit_review_session_view(task_evidence, repo_dir)
ctx._commit_review_evidence = task_evidence
task_evidence_section = commit_review_evidence_section(
task_evidence, delivery="session" if delegated else "native" if retrieves else "packet")
from ouroboros.tools.review_binary_context import StagedDiffUnavailable from ouroboros.tools.review_binary_context import StagedDiffUnavailable
from ouroboros.tools.review_subject import managed_review_subject from ouroboros.tools.review_subject import managed_review_subject
@ -430,6 +441,7 @@ def prepare_scope_review(
drive_root=pathlib.Path(ctx.drive_root) if getattr(ctx, "drive_root", None) else None, drive_root=pathlib.Path(ctx.drive_root) if getattr(ctx, "drive_root", None) else None,
governance_repo_dir=governance_repo, governance_repo_dir=governance_repo,
managed_subject=subject, managed_subject=subject,
task_evidence_section=task_evidence_section,
) )
sr._SCOPE_CONTEXT_MANIFEST.set(session_manifest) sr._SCOPE_CONTEXT_MANIFEST.set(session_manifest)
prompt, context_status = session_task, None prompt, context_status = session_task, None
@ -454,6 +466,7 @@ def prepare_scope_review(
represent_binary=subject is not None, represent_binary=subject is not None,
managed_subject=subject, managed_subject=subject,
window_binding=window_binding, window_binding=window_binding,
task_evidence=task_evidence,
), ),
) )
@ -541,7 +554,7 @@ def prepare_scope_review(
"session_target": session_target, "session_target": session_target,
"session_profile": session_profile, "session_profile": session_profile,
"subagent_id": subagent_id, "subagent_id": subagent_id,
"window_binding": window_binding, "window_binding": window_binding, "task_evidence": task_evidence,
"use_local": override.get("use_local"), "use_local": override.get("use_local"),
"context_manifest": sr._current_scope_context_manifest(), "context_manifest": sr._current_scope_context_manifest(),
"stable_prefix_len": int(sr._SCOPE_STABLE_PREFIX_LEN.get() or 0), "stable_prefix_len": int(sr._SCOPE_STABLE_PREFIX_LEN.get() or 0),
@ -568,6 +581,7 @@ def commit_gate_paid_seats(triad_prepared, triad_exited, scope_rows) -> list:
TRIAD_ROLE_HINT, TRIAD_USER_TURN, _review_output_budget, triad_api_messages, TRIAD_ROLE_HINT, TRIAD_USER_TURN, _review_output_budget, triad_api_messages,
) )
from ouroboros.triad_review import REVIEW_JSON_ARRAY_CONTRACT from ouroboros.triad_review import REVIEW_JSON_ARRAY_CONTRACT
from ouroboros.review_evidence import commit_review_evidence_section
sr = _scope() sr = _scope()
@ -618,7 +632,7 @@ def commit_gate_paid_seats(triad_prepared, triad_exited, scope_rows) -> list:
chars = native_first_send_chars( chars = native_first_send_chars(
str(triad_prepared.get("target_repo") or ""), surface="multi_model_review", str(triad_prepared.get("target_repo") or ""), surface="multi_model_review",
role_hint=TRIAD_ROLE_HINT, slot_id=slot_id, role_hint=TRIAD_ROLE_HINT, slot_id=slot_id,
session_task=str(triad_prepared.get("session_task") or ""), session_task=str(triad_prepared.get("session_task") or "") + ("\n\n" + commit_review_evidence_section(triad_prepared["task_evidence"], delivery="native") if triad_prepared.get("task_evidence") else ""),
output_contract=REVIEW_JSON_ARRAY_CONTRACT, output_contract=REVIEW_JSON_ARRAY_CONTRACT,
) )
else: else:

View file

@ -136,7 +136,7 @@ def _handle_multi_model_review(ctx: ToolContext, content: str = "",
surface: str = "multi_model_review", surface: str = "multi_model_review",
session_policy: dict = None, session_policy: dict = None,
usage_attribution: dict = None, usage_attribution: dict = None,
retry_key: str = "") -> str: retry_key: str = "", task_evidence: dict = None) -> str:
if models is None: if models is None:
models = [] models = []
try: try:
@ -153,13 +153,13 @@ def _handle_multi_model_review(ctx: ToolContext, content: str = "",
_multi_model_review_async(content, prompt, models, ctx, stable_prefix_len, _multi_model_review_async(content, prompt, models, ctx, stable_prefix_len,
routes, session_task, session_root, row_plan, routes, session_task, session_root, row_plan,
surface, session_policy, usage_attribution, surface, session_policy, usage_attribution,
retry_key), retry_key, task_evidence),
).result() ).result()
except RuntimeError: except RuntimeError:
result = asyncio.run(_multi_model_review_async(content, prompt, models, ctx, stable_prefix_len, result = asyncio.run(_multi_model_review_async(content, prompt, models, ctx, stable_prefix_len,
routes, session_task, session_root, row_plan, routes, session_task, session_root, row_plan,
surface, session_policy, usage_attribution, surface, session_policy, usage_attribution,
retry_key)) retry_key, task_evidence))
return json.dumps(result, ensure_ascii=False) return json.dumps(result, ensure_ascii=False)
except Exception as e: except Exception as e:
log.error("Multi-model review failed: %s", e, exc_info=True) log.error("Multi-model review failed: %s", e, exc_info=True)
@ -180,7 +180,7 @@ async def _query_model(
session_target: str = "", session_target: str = "",
session_profile: str = "", session_profile: str = "",
surface: str = "multi_model_review", session_policy: dict = None, usage_attribution: dict = None, surface: str = "multi_model_review", session_policy: dict = None, usage_attribution: dict = None,
retry_key: str = "", subagent_id: str = "", use_local: bool | None = None, retry_key: str = "", subagent_id: str = "", use_local: bool | None = None, task_evidence: dict = None,
): ):
async with semaphore: async with semaphore:
slot = None slot = None
@ -192,6 +192,13 @@ async def _query_model(
# RETRIEVES class (session row OR configured-subagent api row): the # RETRIEVES class (session row OR configured-subagent api row): the
# compact session task replaces the assembled pack for both. # compact session task replaces the assembled pack for both.
retrieves = delivery_retrieves(slot_route, subagent_id) retrieves = delivery_retrieves(slot_route, subagent_id)
from ouroboros.review_evidence import commit_review_evidence_refs, commit_review_evidence_section
evidence = task_evidence or {}
policy = dict(session_policy or {"output_contract": _rev().REVIEW_JSON_ARRAY_CONTRACT}) if retrieves else {}
if retrieves and evidence:
session_task += "\n\n" + commit_review_evidence_section(evidence, delivery="session" if delegated else "native")
if not delegated:
policy["native_data_root"] = evidence["data_root"]
_out_budget = _review_output_budget() _out_budget = _review_output_budget()
request = ReviewRequest( request = ReviewRequest(
surface=surface, surface=surface,
@ -205,7 +212,9 @@ async def _query_model(
no_proxy=True, no_proxy=True,
session_task=session_task if retrieves else "", session_task=session_task if retrieves else "",
session_root=session_root if retrieves else "", session_root=session_root if retrieves else "",
policy=(session_policy or {"output_contract": _rev().REVIEW_JSON_ARRAY_CONTRACT}) if retrieves else {}, policy=policy,
evidence={"task_execution": evidence} if evidence else {},
evidence_refs=commit_review_evidence_refs(evidence),
usage_attribution=usage_attribution or {}, usage_attribution=usage_attribution or {},
task_attempt=getattr(ctx, "task_attempt", None) if ctx is not None else None, task_attempt=getattr(ctx, "task_attempt", None) if ctx is not None else None,
retry_key=str(retry_key or ""), retry_key=str(retry_key or ""),
@ -281,7 +290,7 @@ async def _multi_model_review_async(content: str, prompt: str,
surface: str = "multi_model_review", surface: str = "multi_model_review",
session_policy: dict = None, session_policy: dict = None,
usage_attribution: dict = None, usage_attribution: dict = None,
retry_key: str = ""): retry_key: str = "", task_evidence: dict = None):
from ouroboros.review_execution import ReviewRouteKind, delivery_retrieves from ouroboros.review_execution import ReviewRouteKind, delivery_retrieves
row_routes = list(routes or []) + [ReviewRouteKind.API_CHAT] * max(0, len(models) - len(routes or [])) row_routes = list(routes or []) + [ReviewRouteKind.API_CHAT] * max(0, len(models) - len(routes or []))
@ -332,7 +341,7 @@ async def _multi_model_review_async(content: str, prompt: str,
effort=row_efforts[idx], session_target=row_targets[idx], effort=row_efforts[idx], session_target=row_targets[idx],
session_profile=row_profiles[idx], surface=surface, session_profile=row_profiles[idx], surface=surface,
session_policy=session_policy, usage_attribution=usage_attribution, session_policy=session_policy, usage_attribution=usage_attribution,
retry_key=retry_key, subagent_id=row_actors[idx], use_local=row_local[idx]) retry_key=retry_key, subagent_id=row_actors[idx], use_local=row_local[idx], task_evidence=task_evidence)
for idx, m in enumerate(models) for idx, m in enumerate(models)
] ]
results = await asyncio.gather(*tasks) results = await asyncio.gather(*tasks)

View file

@ -356,6 +356,7 @@ def build_scope_review_prompt(
diff_text: str, diff_text: str,
repo_pack_placeholder: str, repo_pack_placeholder: str,
critical_calibration: str, critical_calibration: str,
task_evidence_section: str = "",
) -> tuple: ) -> tuple:
# STABLE-FIRST for provider prompt caching: instructions, checklist and # STABLE-FIRST for provider prompt caching: instructions, checklist and
# canonical docs are byte-stable across commits and form the cache-marked # canonical docs are byte-stable across commits and form the cache-marked
@ -445,6 +446,8 @@ wider repository pack as omission.
{history_block} {history_block}
{task_evidence_section}
## Current touched files (post-change — what the file looks like NOW) ## Current touched files (post-change — what the file looks like NOW)
Files deleted by this diff appear here with an explicit `DELETED` marker and Files deleted by this diff appear here with an explicit `DELETED` marker and

View file

@ -301,7 +301,7 @@ def _call_scope_llm(
session_root: str = "", session_root: str = "",
slot_effort: str = "", slot_effort: str = "",
session_target: str = "", session_target: str = "",
session_profile: str = "", retry_key: str = "", subagent_id: str = "", use_local: bool | None = None, session_profile: str = "", retry_key: str = "", subagent_id: str = "", use_local: bool | None = None, task_evidence: dict = None,
) -> tuple: ) -> tuple:
"""Execute the scope review call synchronously — api pack or agent session. """Execute the scope review call synchronously — api pack or agent session.
@ -334,8 +334,15 @@ def _call_scope_llm(
try: try:
from ouroboros.review_substrate import ReviewRequest, run_review_request from ouroboros.review_substrate import ReviewRequest, run_review_request
from ouroboros.review_evidence import commit_review_evidence_refs
evidence = task_evidence or {}
policy = {"output_contract": SCOPE_RETRIEVING_OUTPUT_CONTRACT} if retrieves else {}
if retrieves and not delegated and evidence:
policy["native_data_root"] = evidence["data_root"]
request = ReviewRequest( request = ReviewRequest(
surface="scope_review", surface="scope_review",
evidence={"task_execution": evidence} if evidence else {},
evidence_refs=commit_review_evidence_refs(evidence),
goal=SCOPE_USER_TURN, goal=SCOPE_USER_TURN,
messages=messages, messages=messages,
task_id=str(getattr(ctx, "task_id", "") or "scope_review") if ctx is not None else "scope_review", retry_key=str(retry_key or ""), task_id=str(getattr(ctx, "task_id", "") or "scope_review") if ctx is not None else "scope_review", retry_key=str(retry_key or ""),
@ -347,7 +354,7 @@ def _call_scope_llm(
session_root=session_root if retrieves else "", session_root=session_root if retrieves else "",
reconcile_only=bool(getattr(ctx, "_review_reconcile_only", False)), reconcile_only=bool(getattr(ctx, "_review_reconcile_only", False)),
deadline_at=_owner_deadline_at(ctx), deadline_at=_owner_deadline_at(ctx),
policy={"output_contract": SCOPE_RETRIEVING_OUTPUT_CONTRACT} if retrieves else {}, policy=policy,
) )
row = scope_reviewer_slots([scope_model], effort=scope_effort)[0] row = scope_reviewer_slots([scope_model], effort=scope_effort)[0]
slot = replace( slot = replace(
@ -685,7 +692,7 @@ def run_scope_review(
route=route, session_task=session_task, session_root=str(repo_dir), route=route, session_task=session_task, session_root=str(repo_dir),
slot_effort=slot_effort, session_target=session_target, slot_effort=slot_effort, session_target=session_target,
session_profile=session_profile, retry_key=retry_key, subagent_id=subagent_id, session_profile=session_profile, retry_key=retry_key, subagent_id=subagent_id,
use_local=prepared.get("use_local"), use_local=prepared.get("use_local"), task_evidence=prepared.get("task_evidence"),
) # type: ignore[arg-type] ) # type: ignore[arg-type]
_usage = dict(usage or {}) _usage = dict(usage or {})
host_route = _usage.get("model_role_route") or {} host_route = _usage.get("model_role_route") or {}

View file

@ -491,6 +491,7 @@ class _ScopePromptContext:
# None for every ordinary commit — the pack then reads the staged diff. # None for every ordinary commit — the pack then reads the staged diff.
managed_subject: Optional[Any] = None managed_subject: Optional[Any] = None
window_binding: Optional[dict] = None window_binding: Optional[dict] = None
task_evidence: Optional[dict] = None
def _build_scope_prompt( def _build_scope_prompt(
@ -601,6 +602,8 @@ def _build_scope_prompt(
if touched_status is not None: if touched_status is not None:
return None, touched_status return None, touched_status
from ouroboros.review_evidence import commit_review_evidence_section
task_evidence_compact = False
repo_pack_placeholder = "__GENERATED_SCOPE_ATLAS_PENDING__" repo_pack_placeholder = "__GENERATED_SCOPE_ATLAS_PENDING__"
def _assemble_prompt(current_files_section: str) -> str: def _assemble_prompt(current_files_section: str) -> str:
@ -613,6 +616,7 @@ def _build_scope_prompt(
diff_text=diff_text, diff_text=diff_text,
repo_pack_placeholder=repo_pack_placeholder, repo_pack_placeholder=repo_pack_placeholder,
critical_calibration=_sr().CRITICAL_FINDING_CALIBRATION, critical_calibration=_sr().CRITICAL_FINDING_CALIBRATION,
task_evidence_section=commit_review_evidence_section(context.task_evidence or {}, delivery="packet", compact=task_evidence_compact),
) )
_SCOPE_STABLE_PREFIX_LEN.set(stable_len) _SCOPE_STABLE_PREFIX_LEN.set(stable_len)
return prompt_text return prompt_text
@ -727,6 +731,11 @@ def _build_scope_prompt(
# Even the manifest cannot fit beside the fixed part: shrink it for room. # Even the manifest cannot fit beside the fixed part: shrink it for room.
deficit = max(50_000, fixed_prompt_tokens + _atlas_min_allowance - input_limit) deficit = max(50_000, fixed_prompt_tokens + _atlas_min_allowance - input_limit)
if context.task_evidence and not task_evidence_compact:
task_evidence_compact = True
ladder_steps.append({"step": "task_evidence_excerpt_omitted", "source_ref": context.task_evidence.get("source_ref")})
continue
# Degradable never holds atlas-required-beyond-diff paths: the atlas # Degradable never holds atlas-required-beyond-diff paths: the atlas
# refuses a diff-only required artifact by design, so that rung could # refuses a diff-only required artifact by design, so that rung could
# only convert this pack into a typed refusal, never into a fit. # only convert this pack into a typed refusal, never into a fit.

View file

@ -88,6 +88,7 @@ def build_scope_session_task(
drive_root: Optional[pathlib.Path] = None, drive_root: Optional[pathlib.Path] = None,
governance_repo_dir: Optional[pathlib.Path] = None, governance_repo_dir: Optional[pathlib.Path] = None,
managed_subject: Optional[Any] = None, managed_subject: Optional[Any] = None,
task_evidence_section: str = "",
) -> Tuple[str, Dict[str, Any]]: ) -> Tuple[str, Dict[str, Any]]:
"""The scope task in SESSION delivery, plus its forensic coverage manifest. """The scope task in SESSION delivery, plus its forensic coverage manifest.
@ -184,6 +185,7 @@ def build_scope_session_task(
"your own tools)" "your own tools)"
), ),
critical_calibration=CRITICAL_FINDING_CALIBRATION, critical_calibration=CRITICAL_FINDING_CALIBRATION,
task_evidence_section=task_evidence_section,
) )
manifest: Dict[str, Any] = { manifest: Dict[str, Any] = {
# D-12's ratified spelling. It is deliberately NOT `agent_session`: that # D-12's ratified spelling. It is deliberately NOT `agent_session`: that

View file

@ -10,6 +10,8 @@ import os
import pathlib import pathlib
import re import re
import shlex import shlex
import shutil
import tempfile
import signal # noqa: F401 import signal # noqa: F401
import stat # noqa: F401 import stat # noqa: F401
import subprocess import subprocess
@ -30,7 +32,7 @@ from ouroboros.runtime_mode_policy import (
is_protected_runtime_path, # noqa: F401 is_protected_runtime_path, # noqa: F401
) )
from ouroboros.tools.commit_gate import _invalidate_advisory from ouroboros.tools.commit_gate import _invalidate_advisory
from ouroboros.shell_parse import is_absolute_path_text, recover_stringified_argv # noqa: F401 from ouroboros.shell_parse import POSIX_SHELL_HEADS, is_absolute_path_text, recover_stringified_argv # noqa: F401
from ouroboros.tools.tool_result import _publish_process_result, _wrap_run_script_process_result from ouroboros.tools.tool_result import _publish_process_result, _wrap_run_script_process_result
from ouroboros.tools.verify import check_exit_masking # noqa: F401 -- ONE exit-masking sensor shared with verify_and_record (pinned here); its disclosure lives in shell_audit from ouroboros.tools.verify import check_exit_masking # noqa: F401 -- ONE exit-masking sensor shared with verify_and_record (pinned here); its disclosure lives in shell_audit
from ouroboros.tools.registry import ( from ouroboros.tools.registry import (
@ -164,7 +166,7 @@ _SHELL_OPERATORS = frozenset(["&&", "||", "|", ";", ">", ">>", "<", "<<"])
_GLUED_REDIRECT_RE = re.compile( _GLUED_REDIRECT_RE = re.compile(
r'^(?:(?:\d+>>?|>>?&?\d*|\d*>&\d*|&>>?)(?:\S.*)?|\d+<\S*|<<\S*|<)$' r'^(?:(?:\d+>>?|>>?&?\d*|\d*>&\d*|&>>?)(?:\S.*)?|\d+<\S*|<<\S*|<)$'
) )
_SHELL_INTERPRETERS = frozenset({"sh", "bash", "zsh", "fish", "cmd", "cmd.exe", "powershell", "powershell.exe", "pwsh", "pwsh.exe"}) _SHELL_INTERPRETERS = POSIX_SHELL_HEADS | frozenset({"fish", "cmd", "cmd.exe", "powershell", "powershell.exe", "pwsh", "pwsh.exe"})
_ENV_REF_PATTERN = re.compile(r'\$(?:\{[A-Z][A-Z0-9_]*\}|[A-Z][A-Z0-9_]*)') _ENV_REF_PATTERN = re.compile(r'\$(?:\{[A-Z][A-Z0-9_]*\}|[A-Z][A-Z0-9_]*)')
@ -625,35 +627,46 @@ def _run_script(
root = pathlib.Path(ctx.drive_root) / "tmp_scripts" root = pathlib.Path(ctx.drive_root) / "tmp_scripts"
root.mkdir(parents=True, exist_ok=True) root.mkdir(parents=True, exist_ok=True)
suffix = ".py" if "python" in pathlib.PurePath(interp).name else ".sh" suffix = ".py" if "python" in pathlib.PurePath(interp).name else ".sh"
run_dir = None
script_path = root / f"script_{uuid.uuid4().hex}{suffix}" script_path = root / f"script_{uuid.uuid4().hex}{suffix}"
script_path.write_text(body, encoding="utf-8")
try:
os.chmod(script_path, 0o600)
except OSError:
pass
script_arg = str(script_path)
if executor_active:
executor = executor_ref_from_ctx(ctx)
if executor is not None and executor.kind != "local":
try:
script_arg = executor_map_host_path(executor, script_path)
except Exception as exc:
script_path.unlink(missing_ok=True)
return f"⚠️ RUN_SCRIPT_BLOCKED: executor-backed run_script could not map temp script path: {type(exc).__name__}: {exc}"
argv = [interp, script_arg, *[str(item) for item in (args or [])]]
try: try:
if active_workspace_script:
run_dir = pathlib.Path(tempfile.mkdtemp(prefix="script_", dir=root))
# Ignore only this invocation's files, not neighbouring user work.
(run_dir / ".gitignore").write_text("*\n", encoding="utf-8")
script_path = run_dir / f"script{suffix}"
script_path.write_text(body, encoding="utf-8")
try:
os.chmod(script_path, 0o600)
except OSError:
pass
script_arg = str(script_path)
if executor_active:
executor = executor_ref_from_ctx(ctx)
if executor is not None and executor.kind != "local":
try:
script_arg = executor_map_host_path(executor, script_path)
except Exception as exc:
return f"⚠️ RUN_SCRIPT_BLOCKED: executor-backed run_script could not map temp script path: {type(exc).__name__}: {exc}"
argv = [interp, script_arg, *[str(item) for item in (args or [])]]
result = _run_shell( result = _run_shell(
ctx, argv, cwd=cwd, outputs=outputs, scratch=scratch, ctx, argv, cwd=cwd, outputs=outputs, scratch=scratch,
_resolved_binding=binding, timeout_sec=timeout_sec, timeout=timeout, _resolved_binding=binding, timeout_sec=timeout_sec, timeout=timeout,
) )
finally: finally:
try: try:
script_path.unlink(missing_ok=True) if run_dir is not None:
script_path.parent.rmdir() shutil.rmtree(run_dir)
if active_workspace_script: else:
script_path.parent.parent.rmdir() script_path.unlink(missing_ok=True)
except OSError: except OSError as exc:
pass log.warning("Could not remove run_script scratch %s (%s)", run_dir or script_path, type(exc).__name__)
# These shared parents may contain another run or a user's file.
for parent in (root, root.parent) if active_workspace_script else (root,):
try:
parent.rmdir()
except OSError:
pass
if pathlib.PurePath(interp).name in {"sh", "bash"}: if pathlib.PurePath(interp).name in {"sh", "bash"}:
result = _masked_green_disclosure(ctx, result, [interp, "-c", body]) result = _masked_green_disclosure(ctx, result, [interp, "-c", body])
# POST-exec body audit: stat-confirmed user_files writes performed by the script # POST-exec body audit: stat-confirmed user_files writes performed by the script
@ -740,7 +753,7 @@ def get_tools() -> List[ToolEntry]:
"name": "run_script", "name": "run_script",
"description": ( "description": (
"Run a short task-scoped temporary script with a declared interpreter. " "Run a short task-scoped temporary script with a declared interpreter. "
"Use for multi-line diagnostics or harness helpers; generated script files live under the task drive. " "Use for multi-line diagnostics or harness helpers; generated scripts use a private run directory inside the mapped workspace or the task drive. "
"The underlying command result echoes the resolved cwd." "The underlying command result echoes the resolved cwd."
), ),
"parameters": {"type": "object", "properties": { "parameters": {"type": "object", "properties": {

View file

@ -7,6 +7,7 @@ import re
from typing import List, TYPE_CHECKING from typing import List, TYPE_CHECKING
from ouroboros.shell_parse import ( from ouroboros.shell_parse import (
POSIX_SHELL_HEADS,
collect_leading_env, collect_leading_env,
embedded_absolute_path_tokens, embedded_absolute_path_tokens,
is_absolute_path_text, is_absolute_path_text,
@ -50,7 +51,6 @@ if TYPE_CHECKING:
_UNDECLARED_OUTPUTS_MARKER = "⚠️ ARTIFACT_OUTPUT_UNDECLARED" _UNDECLARED_OUTPUTS_MARKER = "⚠️ ARTIFACT_OUTPUT_UNDECLARED"
_UNDECLARED_OUTPUT_SCAN_MAX_FILES = 5000 _UNDECLARED_OUTPUT_SCAN_MAX_FILES = 5000
_UNDECLARED_OUTPUT_METADATA_COMMANDS = frozenset({"chmod", "chown", "mkdir", "rm"}) _UNDECLARED_OUTPUT_METADATA_COMMANDS = frozenset({"chmod", "chown", "mkdir", "rm"})
_SHELL_WRAPPER_COMMANDS = frozenset({"sh", "bash", "zsh"})
def _redirect_targets_for_audit(argv: list[str]) -> set[str]: def _redirect_targets_for_audit(argv: list[str]) -> set[str]:
@ -70,7 +70,7 @@ def _writer_targets_for_output_audit(argv: list[str]) -> set[str]:
except Exception: except Exception:
command_argv = list(segment) command_argv = list(segment)
command = pathlib.PurePath(command_argv[0]).name.lower().removesuffix(".exe") if command_argv else "" command = pathlib.PurePath(command_argv[0]).name.lower().removesuffix(".exe") if command_argv else ""
if command in _SHELL_WRAPPER_COMMANDS: if command in POSIX_SHELL_HEADS:
body = shell_command_string(command_argv) body = shell_command_string(command_argv)
if body: if body:
targets.update(_writer_targets_for_output_audit(shell_argv_with_inline(body))) targets.update(_writer_targets_for_output_audit(shell_argv_with_inline(body)))

View file

@ -10,6 +10,7 @@ from typing import Any, Dict, List
from ouroboros.runtime_mode_policy import FROZEN_CONTRACT_PATH_PREFIXES, PROTECTED_RUNTIME_PATHS from ouroboros.runtime_mode_policy import FROZEN_CONTRACT_PATH_PREFIXES, PROTECTED_RUNTIME_PATHS
from ouroboros.shell_parse import ( from ouroboros.shell_parse import (
POSIX_SHELL_HEADS,
EMBEDDED_WINDOWS_ABSOLUTE_PATH_RE, EMBEDDED_WINDOWS_ABSOLUTE_PATH_RE,
collect_leading_env, collect_leading_env,
embedded_absolute_path_tokens, embedded_absolute_path_tokens,
@ -618,7 +619,7 @@ def shell_inspection_paths(
row_cwd = (row_cwd / pathlib.Path(wrapper_cwd).expanduser()).resolve(strict=False) row_cwd = (row_cwd / pathlib.Path(wrapper_cwd).expanduser()).resolve(strict=False)
head = pathlib.PurePath(argv[0]).name.lower().removesuffix(".exe") head = pathlib.PurePath(argv[0]).name.lower().removesuffix(".exe")
nested = [] nested = []
if head in _SHELL_WRAPPER_HEADS and depth < _MAX_INLINE_RECURSION: if head in POSIX_SHELL_HEADS and depth < _MAX_INLINE_RECURSION:
body = shell_command_string(argv) body = shell_command_string(argv)
nested = [body] if body else list(heredocs) if interpreter_reads_program_from_stdin(argv) else [] nested = [body] if body else list(heredocs) if interpreter_reads_program_from_stdin(argv) else []
for body in nested: for body in nested:
@ -905,7 +906,6 @@ _SED_SCRIPT_WRITE_RE = re.compile(
r"(?<![A-Za-z_])[wW]\s+\S|(?<![A-Za-z_])e(?:\s*(?:$|;)|\s+\S)|/[gpimM0-9]*[we](?=\s|$|;)" r"(?<![A-Za-z_])[wW]\s+\S|(?<![A-Za-z_])e(?:\s*(?:$|;)|\s+\S)|/[gpimM0-9]*[we](?=\s|$|;)"
) )
# A wrapper body is a command line; `cd` can move later relative writes. # A wrapper body is a command line; `cd` can move later relative writes.
_SHELL_WRAPPER_HEADS = frozenset({"sh", "bash", "zsh", "dash", "ash"})
_DIRECTORY_CHANGE_COMMANDS = frozenset({"cd", "pushd"}) _DIRECTORY_CHANGE_COMMANDS = frozenset({"cd", "pushd"})
_MAX_INLINE_RECURSION = 3 _MAX_INLINE_RECURSION = 3
def writer_target_rows(raw_cmd: Any, _depth: int = 0) -> List[tuple]: def writer_target_rows(raw_cmd: Any, _depth: int = 0) -> List[tuple]:
@ -921,7 +921,7 @@ def writer_target_rows(raw_cmd: Any, _depth: int = 0) -> List[tuple]:
continue continue
executable = pathlib.PurePath(str(argv[0])).name.lower().removesuffix(".exe") executable = pathlib.PurePath(str(argv[0])).name.lower().removesuffix(".exe")
program_argv, _stdin_redirects = split_redirections(argv) program_argv, _stdin_redirects = split_redirections(argv)
if _depth < _MAX_INLINE_RECURSION and executable in _SHELL_WRAPPER_HEADS: if _depth < _MAX_INLINE_RECURSION and executable in POSIX_SHELL_HEADS:
shell_body = shell_command_string(argv) shell_body = shell_command_string(argv)
stdin_bodies = heredoc_bodies if not shell_body and interpreter_reads_program_from_stdin(program_argv) else () stdin_bodies = heredoc_bodies if not shell_body and interpreter_reads_program_from_stdin(program_argv) else ()
nested = writer_target_rows(shell_body, _depth + 1) nested = writer_target_rows(shell_body, _depth + 1)
@ -1287,7 +1287,7 @@ def shell_writer_targets_protected(raw_cmd: Any) -> bool:
if not argv: if not argv:
return False return False
executable = pathlib.PurePath(argv[0]).name.lower().removesuffix(".exe") executable = pathlib.PurePath(argv[0]).name.lower().removesuffix(".exe")
if executable in {"bash", "sh", "zsh"}: if executable in POSIX_SHELL_HEADS:
inline = shell_command_string(argv) inline = shell_command_string(argv)
return bool(inline and shell_writer_targets_protected(inline)) return bool(inline and shell_writer_targets_protected(inline))
if not _light_writer_command(executable): if not _light_writer_command(executable):
@ -1437,7 +1437,7 @@ def light_shell_repo_mutation(
return False return False
executable = pathlib.PurePath(argv[0]).name.lower().removesuffix(".exe") executable = pathlib.PurePath(argv[0]).name.lower().removesuffix(".exe")
if executable in {"bash", "sh", "zsh"}: if executable in POSIX_SHELL_HEADS:
inline = shell_command_string(argv) inline = shell_command_string(argv)
if inline: if inline:
return light_shell_repo_mutation( return light_shell_repo_mutation(

View file

@ -92,6 +92,9 @@ class ToolContext:
# Conversation messages for safety checks. # Conversation messages for safety checks.
messages: Optional[List[Dict[str, Any]]] = None messages: Optional[List[Dict[str, Any]]] = None
# Borrowed loop trace; restored when the owning loop exits.
_execution_trace: Optional[Dict[str, Any]] = field(default=None, repr=False)
# Structured task constraints, e.g. skill repair payload confinement. # Structured task constraints, e.g. skill repair payload confinement.
task_constraint: Optional[TaskConstraint] = None task_constraint: Optional[TaskConstraint] = None
task_contract: Dict[str, Any] = field(default_factory=dict) task_contract: Dict[str, Any] = field(default_factory=dict)

View file

@ -29,7 +29,7 @@ from ouroboros.process_interpreters import (
apply_env_path_prepend, apply_env_path_prepend,
interpreter_path_overlay, interpreter_path_overlay,
) )
from ouroboros.shell_parse import normalize_check_argv from ouroboros.shell_parse import POSIX_SHELL_HEADS, normalize_check_argv, shell_tokens_typed
from ouroboros.tool_access import ( from ouroboros.tool_access import (
ResolvedResourceBinding, ResolvedResourceBinding,
active_tool_profile, active_tool_profile,
@ -123,7 +123,6 @@ _normalize_check = normalize_check_argv
# Shell stages that, as the LAST stage of a pipeline, almost always exit 0 even when an earlier # Shell stages that, as the LAST stage of a pipeline, almost always exit 0 even when an earlier
# real command failed — so the pipeline's exit (POSIX: the last stage's) MASKS the true result. # real command failed — so the pipeline's exit (POSIX: the last stage's) MASKS the true result.
_EXIT_MASK_FILTER_CMDS = frozenset({"tail", "head", "grep", "egrep", "fgrep", "sed", "awk", "cat", "tee", "tr", "sort", "uniq", "wc", "true", ":"}) _EXIT_MASK_FILTER_CMDS = frozenset({"tail", "head", "grep", "egrep", "fgrep", "sed", "awk", "cat", "tee", "tr", "sort", "uniq", "wc", "true", ":"})
_SHELL_C_HEADS = frozenset({"sh", "bash", "dash", "ash", "zsh"})
def _check_has_exit_masking(argv: List[str]) -> tuple[bool, list[str]]: def _check_has_exit_masking(argv: List[str]) -> tuple[bool, list[str]]:
@ -136,33 +135,33 @@ def _check_has_exit_masking(argv: List[str]) -> tuple[bool, list[str]]:
informs the advisory reviewer + the agent, P5-clean (it decides nothing). Returns (masked, reasons).""" informs the advisory reviewer + the agent, P5-clean (it decides nothing). Returns (masked, reasons)."""
if not argv or len(argv) < 3: if not argv or len(argv) < 3:
return False, [] return False, []
if pathlib.PurePath(str(argv[0])).name.lower() not in _SHELL_C_HEADS or str(argv[1]) not in ("-c", "-lc"): if pathlib.PurePath(str(argv[0])).name.lower() not in POSIX_SHELL_HEADS or str(argv[1]) not in ("-c", "-lc"):
return False, [] return False, []
text = str(argv[2]) typed = shell_tokens_typed(str(argv[2]))
# Operator-aware tokenization (shlex with `punctuation_chars`) so `|`/`||` are split out as if typed is None:
# standalone tokens EVEN WHEN glued to words (`pytest -q|tail`, `make test||true`) — plain
# shlex.split is whitespace-only and would miss the no-space forms. Quotes are still respected,
# so a quoted literal (e.g. a grep pattern `'| tail'`) is NOT flagged.
try:
lexer = shlex.shlex(text, posix=True, punctuation_chars="|&<>;")
lexer.whitespace_split = True
toks = list(lexer)
except ValueError:
return False, [] return False, []
# Strip only syntactic grouping, including punctuation glued to a pipe
# (``)|(``). Quoted/escaped parentheses and operators remain literal data.
toks: list[tuple[str, bool]] = []
for token, operator in typed:
if operator:
toks.extend((part, True) for part in token.replace("(", " ").replace(")", " ").split())
else:
toks.append((token, False))
reasons: list[str] = [] reasons: list[str] = []
for i, tok in enumerate(toks[:-1]): for i, tok in enumerate(toks[:-1]):
if tok == "||" and toks[i + 1] in ("true", ":"): if tok == ("||", True) and toks[i + 1] in (("true", False), (":", False)):
reasons.append("|| true") reasons.append("|| true")
break break
pipe_positions = [i for i, tok in enumerate(toks) if tok == "|"] pipe_positions = [i for i, tok in enumerate(toks) if tok == ("|", True)]
if pipe_positions: if pipe_positions:
nxt = pipe_positions[-1] + 1 nxt = pipe_positions[-1] + 1
last_stage = pathlib.PurePath(toks[nxt]).name.lower() if nxt < len(toks) else "" last_stage = pathlib.PurePath(toks[nxt][0]).name.lower() if nxt < len(toks) and not toks[nxt][1] else ""
if last_stage in _EXIT_MASK_FILTER_CMDS: if last_stage in _EXIT_MASK_FILTER_CMDS:
reasons.append(f"pipeline_{last_stage}") reasons.append(f"pipeline_{last_stage}")
if len(toks) >= 2 and toks[-1] in {"true", ":"} and toks[-2] == ";": if len(toks) >= 2 and toks[-1] in (("true", False), (":", False)) and toks[-2] == (";", True):
reasons.append(f"{toks[-2]} true") reasons.append("; true")
if len(toks) >= 3 and toks[-2:] == ["exit", "0"] and toks[-3] in {";", "||"}: if len(toks) >= 3 and toks[-2:] == [("exit", False), ("0", False)] and toks[-3] in ((";", True), ("||", True)):
reasons.append("exit 0") reasons.append("exit 0")
seen: set = set() seen: set = set()
ordered = [r for r in reasons if not (r in seen or seen.add(r))] ordered = [r for r in reasons if not (r in seen or seen.add(r))]

View file

@ -19,7 +19,7 @@ import re
import tokenize import tokenize
from typing import Any, Callable, List, Optional from typing import Any, Callable, List, Optional
from ouroboros.shell_parse import shell_argv, shell_argv_with_inline, shell_argv_with_path_tokens from ouroboros.shell_parse import POSIX_SHELL_HEADS, shell_argv, shell_argv_with_inline, shell_argv_with_path_tokens
SHELL_WRITE_INDICATORS = ( SHELL_WRITE_INDICATORS = (
"rm ", "rm\t", ">", "sed -i", "tee ", "truncate", "rm ", "rm\t", ">", "sed -i", "tee ", "truncate",
@ -322,7 +322,7 @@ def _shell_write_indicator_scan(
bodies = list(interpreter_inline_code(argv)) bodies = list(interpreter_inline_code(argv))
# An sh -c wrap hides the interpreter one level down; locate the inner # An sh -c wrap hides the interpreter one level down; locate the inner
# bodies too so a '>' comparison inside them is not read as a redirect. # bodies too so a '>' comparison inside them is not read as a redirect.
if argv and str(argv[0]).rsplit("/", 1)[-1].lower() in {"sh", "bash", "zsh"}: if argv and str(argv[0]).rsplit("/", 1)[-1].lower() in POSIX_SHELL_HEADS:
inner = shell_command_string(argv) inner = shell_command_string(argv)
if inner: if inner:
bodies.extend(interpreter_inline_code(shell_argv(inner))) bodies.extend(interpreter_inline_code(shell_argv(inner)))

View file

@ -50,7 +50,9 @@ def main(argv: list[str]) -> int:
try: try:
items = json.loads(raw) items = json.loads(raw)
except json.JSONDecodeError: except json.JSONDecodeError:
items = extract_json_array(raw) items = extract_json_array(raw, validate_fn=lambda candidate: not normalize_scope_items(candidate)[1])
if items is None:
items = extract_json_array(raw)
if items is None: if items is None:
print("invalid: no JSON array found in the receipt " print("invalid: no JSON array found in the receipt "
"(bare, fenced, or embedded)", file=sys.stderr) "(bare, fenced, or embedded)", file=sys.stderr)

View file

@ -1,7 +1,7 @@
--- ---
name: telegram name: telegram
description: Owner-only Telegram text bridge and Mini App gateway for the existing Ouroboros interface. description: Owner-only Telegram text bridge and Mini App gateway for the existing Ouroboros interface.
version: 1.2.1 version: 1.2.2
type: extension type: extension
entry: plugin.py entry: plugin.py
plugin_api: "2.0" plugin_api: "2.0"
@ -47,6 +47,8 @@ cancelled).
Version 1.2.1 marks questions that wait for an owner answer and clears the Version 1.2.1 marks questions that wait for an owner answer and clears the
waiting line after an answer submitted through Telegram. waiting line after an answer submitted through Telegram.
Version 1.2.2 re-seeds the bundled skill when these payload bytes change; it
adds no new runtime behavior.
The Mini App exposes the unchanged Ouroboros SPA through the established The Mini App exposes the unchanged Ouroboros SPA through the established
owner-authenticated sidecar and a pinned Cloudflare Quick Tunnel. It is enabled owner-authenticated sidecar and a pinned Cloudflare Quick Tunnel. It is enabled

View file

@ -1,7 +1,7 @@
--- ---
name: unix_computer_use name: unix_computer_use
description: Local and remote desktop observation/input tools with coordinate normalization (local macOS/Linux by default; optional OSWorld HTTP and SSH Mac backends). description: Local and remote desktop observation/input tools with coordinate normalization (local macOS/Linux by default; optional OSWorld HTTP and SSH Mac backends).
version: 0.4.1 version: 0.4.2
type: extension type: extension
entry: plugin.py entry: plugin.py
plugin_api: "2.0" plugin_api: "2.0"

View file

@ -106,7 +106,11 @@ def test_empty_advisory_result_is_error(monkeypatch, tmp_path):
assert items == [] assert items == []
assert raw.startswith("⚠️ ADVISORY_ERROR:") assert raw.startswith("⚠️ ADVISORY_ERROR:")
assert "empty output" in raw assert "empty output" in raw
assert any(ev.get("type") == "advisory_suspect_result" for ev in ctx.pending_events) events = [ev for ev in ctx.pending_events if ev.get("type") == "advisory_suspect_result"]
assert len(events) == 1
assert events[0]["cost_usd"] == 1.23
assert events[0]["session_id"] == 'sess-empty'
assert events[0]["prompt_chars"] == len("prompt")
def test_handle_advisory_error_persists_session_id(monkeypatch, tmp_path): def test_handle_advisory_error_persists_session_id(monkeypatch, tmp_path):
@ -181,7 +185,11 @@ def test_skill_advisory_duplicate_expected_items_warn_not_error(monkeypatch, tmp
assert len(items) == 3 assert len(items) == 3
assert not raw.startswith("⚠️ ADVISORY_ERROR:") assert not raw.startswith("⚠️ ADVISORY_ERROR:")
assert any(ev.get("type") == "advisory_contract_warning" for ev in ctx.pending_events) events = [ev for ev in ctx.pending_events if ev.get("type") == "advisory_contract_warning"]
assert len(events) == 1
assert events[0]["cost_usd"] == 0.2
assert events[0]["session_id"] == 'sess-duplicate'
assert events[0]["prompt_chars"] == len("prompt")
assert not any(ev.get("type") == "advisory_suspect_result" for ev in ctx.pending_events) assert not any(ev.get("type") == "advisory_suspect_result" for ev in ctx.pending_events)
@ -269,7 +277,11 @@ def test_skill_advisory_missing_expected_items_still_errors(monkeypatch, tmp_pat
assert "Full reviewer output:" in raw assert "Full reviewer output:" in raw
assert raw.startswith("⚠️ ADVISORY_ERROR:") assert raw.startswith("⚠️ ADVISORY_ERROR:")
assert "checklist contract mismatch" in raw assert "checklist contract mismatch" in raw
assert any(ev.get("type") == "advisory_suspect_result" for ev in ctx.pending_events) events = [ev for ev in ctx.pending_events if ev.get("type") == "advisory_suspect_result"]
assert len(events) == 1
assert events[0]["cost_usd"] == 0.2
assert events[0]["session_id"] == 'sess-partial'
assert events[0]["prompt_chars"] == len("prompt")
# --------------------------------------------------------------------------- # ---------------------------------------------------------------------------

View file

@ -0,0 +1,575 @@
"""Selected commit evidence survives compaction and each reviewer delivery."""
from __future__ import annotations
import asyncio
import base64
import json
import pathlib
import shutil
import subprocess
import time
from types import SimpleNamespace
import pytest
from ouroboros.artifacts import read_actor_source_bytes
from ouroboros.observability import persist_call, promote_child_task_refs, read_blob_ref
from ouroboros.outcomes import collect_trace_refs
from ouroboros.review_evidence import (
_ACCEPT_NOTES_CAP,
capture_commit_review_evidence,
commit_review_evidence_section,
materialize_commit_review_session_view,
pending_commit_review_evidence,
release_commit_review_session_view,
restore_commit_review_evidence,
)
from ouroboros.tools.registry import ToolContext
from tests._workspace_executor_shared import _init_repo
pytestmark = pytest.mark.serial
@pytest.fixture
def evidence_context(tmp_path):
repo, child, canonical = (tmp_path / name for name in ("repo", "child", "canonical"))
_init_repo(repo)
(repo / ".gitignore").write_text("/.review-drive/\n")
child.mkdir()
canonical.mkdir()
ctx = ToolContext(repo_dir=repo, drive_root=child, budget_drive_root=str(canonical), task_id="evidence-task")
ctx._execution_trace = {"tool_calls": [], "reasoning_notes": []}
ctx._accumulated_usage = {"llm_call_refs": []}
return ctx
def model_response(ctx, name, content, *, execution="solve", call_type="llm_response"):
trace = persist_call(ctx.drive_root, task_id=ctx.task_id, call_id=name + "_response",
call_type=call_type, payload={"message": {"content": content}},
manifest={"execution_id": execution, "llm_call_id": name})
row = {"llm_call_id": name, "execution_id": execution, "response_ref": trace["manifest_ref"]}
ctx._accumulated_usage["llm_call_refs"].append(row)
return row
def tool_response(ctx, name, parent, *, tool="view_image", result="image attached", args=None):
args = args or {"path": "/tmp/view.png"}
trace = persist_call(ctx.drive_root, task_id=ctx.task_id, call_id=name,
call_type="tool_call", payload={"tool": tool, "tool_call_id": name,
"args": args, "result": result, "parent_call_id": parent,
"execution_id": "solve", "round_id": "round-1", "semantic_ok": True,
"result_meta": {"status": "ok"}})
row = {"tool": tool, "tool_call_id": name, "args": args, "result": result, "trace_ref": trace}
ctx._execution_trace["tool_calls"].append(row)
return row
def test_large_task_selected_source_is_complete_and_native_readable(evidence_context):
ctx = evidence_context
model_response(ctx, "before", "Inspect the screenshot")
image = tool_response(ctx, "image", "before")
assessment = "Visible assessment\n" + ("The button aligns with the field.\n" * 900) + "ASSESSMENT_END"
model_response(ctx, "after", [{"type": "thinking", "text": "PRIVATE_THINKING"}, {"type": "text", "text": assessment}])
source = capture_commit_review_evidence(ctx)
exact = read_actor_source_bytes(ctx.budget_drive_root, ctx.task_id, source["source_ref"]).decode()
assert "ASSESSMENT_END" in exact
assert "PRIVATE_THINKING" not in exact
assert source["source_complete"] is True
assert source["selected_count"] == 1
for delivery in ("native", "session", "packet"):
assert len(commit_review_evidence_section(source, delivery=delivery)) <= _ACCEPT_NOTES_CAP
assert "evidence_delivery=partial" in commit_review_evidence_section(source, delivery="packet")
assert "host-retained provenance" in commit_review_evidence_section(source, delivery="packet")
ctx._execution_trace["reasoning_notes"] = ["unrelated" * 300000]
ctx._execution_trace["tool_calls"].extend([{"tool": "read_file", "result": "unrelated" * 2000}] * 100)
ctx.messages = [] # Compaction cannot remove the completed tool trace.
ctx._tool_trace_refs = {}
again = capture_commit_review_evidence(ctx)
assert again["source_ref"] == source["source_ref"]
assert again["unselected_count"] == 100
assert len(commit_review_evidence_section(again, delivery="native")) <= _ACCEPT_NOTES_CAP
image["trace_ref"]["call_id"] = "changed-after-freeze"
assert source["original_refs"][0]["call_id"] == "image"
from ouroboros.review_native_episode import inspection_registry
registry, native_ctx, _ = inspection_registry(str(ctx.repo_dir), ctx.budget_drive_root, ctx.task_id)
result = registry.execute_result("read_file", {"root": "artifact_store", "path": source["source_ref"]["path"], "max_lines": 12})
assert result.status == "ok", result.text
assert "Selected browser/vision execution sources" in result.text
assert native_ctx.last_read_view["opened_root"] == "artifact_store"
@pytest.mark.parametrize("shape", ["missing", "empty", "failed", "other_execution", "tampered"])
def test_following_response_gap_never_selects_a_later_success(evidence_context, shape):
ctx = evidence_context
model_response(ctx, "before", "Before")
tool_response(ctx, "image", "before")
if shape != "missing":
row = model_response(ctx, "following", "" if shape == "empty" else "IMMEDIATE_TEXT",
execution="other" if shape == "other_execution" else "solve",
call_type="llm_error" if shape == "failed" else "llm_response")
if shape == "tampered":
pathlib.Path(row["response_ref"]["path"]).write_text("{}")
if shape not in {"missing", "other_execution"}:
model_response(ctx, "later", "LATER_SUCCESS_MUST_NOT_BE_SELECTED")
model_response(ctx, "foreign", "FOREIGN_SUCCESS", execution="other")
source = capture_commit_review_evidence(ctx)
exact = read_actor_source_bytes(ctx.budget_drive_root, ctx.task_id, source["source_ref"]).decode()
assert "LATER_SUCCESS_MUST_NOT_BE_SELECTED" not in exact
assert "FOREIGN_SUCCESS" not in exact
if shape == "empty":
assert '"following_visible_text_status": "no_visible_text"' in exact
else:
assert "following_response_gap" in exact
assert not source["source_complete"]
def test_multiple_tools_share_one_following_response(evidence_context):
ctx = evidence_context
model_response(ctx, "before", "Before")
tool_response(ctx, "one", "before", tool="browse_page")
tool_response(ctx, "two", "before")
model_response(ctx, "after", "ONE_SHARED_ASSESSMENT")
source = capture_commit_review_evidence(ctx)
raw = read_actor_source_bytes(ctx.budget_drive_root, ctx.task_id, source["source_ref"]).decode()
assert raw.count("ONE_SHARED_ASSESSMENT") == 1
assert source["selected_count"] == 2
assert "evidence_delivery=complete_selected" in commit_review_evidence_section(source, delivery="packet")
def test_session_view_is_identical_ignored_restorable_and_disposable(evidence_context):
ctx = evidence_context
model_response(ctx, "before", "Before")
tool_response(ctx, "one", "before")
model_response(ctx, "after", "Exact assessment")
source = capture_commit_review_evidence(ctx)
view = materialize_commit_review_session_view(source, ctx.repo_dir)
assert view["session_source_status"] == "ready"
path = pathlib.Path(view["session_path"])
exact = read_actor_source_bytes(ctx.budget_drive_root, ctx.task_id, source["source_ref"])
assert path.read_bytes() == exact
status = subprocess.run(["git", "status", "--porcelain", "--untracked-files=all"], cwd=ctx.repo_dir, check=True, capture_output=True, text=True)
assert ".review-drive" not in status.stdout
path.unlink()
restored = restore_commit_review_evidence(ctx, source["source_ref"])
materialize_commit_review_session_view(restored, ctx.repo_dir)
assert path.read_bytes() == exact
release_commit_review_session_view(view)
assert not path.exists()
assert read_actor_source_bytes(ctx.budget_drive_root, ctx.task_id, source["source_ref"]) == exact
def test_unignored_session_root_keeps_a_partial_exhibit_without_widening_policy(evidence_context):
ctx = evidence_context
model_response(ctx, "before", "Before")
tool_response(ctx, "one", "before")
model_response(ctx, "after", "An assessment")
source = capture_commit_review_evidence(ctx)
(ctx.repo_dir / ".gitignore").unlink()
view = materialize_commit_review_session_view(source, ctx.repo_dir)
assert view["session_source_status"] == "unavailable"
assert not (ctx.repo_dir / ".review-drive").exists()
assert "retrieval is unavailable" in commit_review_evidence_section(view, delivery="session")
def test_original_responses_survive_child_cleanup_without_new_tool_metadata(evidence_context):
ctx = evidence_context
model_response(ctx, "before", "Before")
tool_response(ctx, "one", "before")
model_response(ctx, "after", "RETAINED_ORIGINAL")
source = capture_commit_review_evidence(ctx)
refs = collect_trace_refs(ctx._accumulated_usage, ctx._execution_trace)
result, promotion = promote_child_task_refs(pathlib.Path(ctx.budget_drive_root), ctx.drive_root, ctx.task_id, {"trace_refs": refs})
assert promotion["status"] == "complete"
shutil.rmtree(ctx.drive_root)
ref = result["trace_refs"]["llm_call_refs"][-1]["response_ref"]
manifest = json.loads(pathlib.Path(ref["path"]).read_text())
payload = read_blob_ref(pathlib.Path(ctx.budget_drive_root), manifest["redacted_projection_ref"])
assert payload["message"]["content"] == "RETAINED_ORIGINAL"
assert b"RETAINED_ORIGINAL" in read_actor_source_bytes(ctx.budget_drive_root, ctx.task_id, source["source_ref"])
assert "review_evidence_refs" not in ctx._execution_trace["tool_calls"][0]
def test_pending_reconciliation_uses_the_recorded_source(evidence_context):
ctx = evidence_context
model_response(ctx, "before", "Before")
tool_response(ctx, "one", "before")
model_response(ctx, "after", "Frozen assessment")
frozen = capture_commit_review_evidence(ctx)
prompt = persist_call(ctx.drive_root, task_id=ctx.task_id, call_id="review", call_type="review_prompt",
payload={"request": {"task_id": ctx.task_id, "evidence": {"task_execution": frozen}}})
ctx._pending_review_attempt = SimpleNamespace(triad_raw_results=[{"prompt_ref": prompt}], scope_raw_result={})
ctx._execution_trace["tool_calls"] = []
assert pending_commit_review_evidence(ctx) == frozen
@pytest.mark.parametrize("delivery", ["packet", "native", "session"])
def test_triad_request_preserves_evidence_and_native_root(evidence_context, monkeypatch, delivery):
from ouroboros.review_records import ReviewRouteKind
from ouroboros.tools.review_multi_model import _query_model
import ouroboros.review_substrate as substrate
ctx = evidence_context
model_response(ctx, "before", "Before")
tool_response(ctx, "one", "before")
model_response(ctx, "after", "Assessment")
evidence = capture_commit_review_evidence(ctx)
if delivery == "session":
evidence = materialize_commit_review_session_view(evidence, ctx.repo_dir)
captured = []
def run(request, **kwargs):
captured.append(request)
return SimpleNamespace(actors=[{"status": "ok", "raw_text": "[]"}])
monkeypatch.setattr(substrate, "run_review_request", run)
route = ReviewRouteKind.AGENT_SESSION if delivery == "session" else ReviewRouteKind.API_CHAT
messages = [{"role": "user", "content": "packet"}]
asyncio.run(_query_model(None, "model", messages, asyncio.Semaphore(1), ctx,
route=route, session_task="Review", session_root=str(ctx.repo_dir),
task_evidence=evidence, subagent_id="native" if delivery == "native" else "", use_local=False))
request = captured[0]
assert request.evidence["task_execution"]["source_ref"] == evidence["source_ref"]
assert evidence["source_ref"] in request.evidence_refs
if delivery == "native":
assert request.policy["native_data_root"] == ctx.budget_drive_root
assert "root='artifact_store'" in request.session_task
elif delivery == "session":
assert evidence["session_relative_path"] in request.session_task
assert "native_data_root" not in request.policy
else:
assert request.messages == messages
assert request.session_task == ""
@pytest.mark.parametrize("fail", [False, True])
def test_loop_borrows_trace_through_tool_calls_and_restores_it(tmp_path, monkeypatch, fail):
from ouroboros import loop
from ouroboros.tools.registry import ToolRegistry
from tests.test_loop_transport_wait import _loop_kwargs
registry = ToolRegistry(repo_dir=tmp_path, drive_root=tmp_path)
previous = {"prior": "trace"}
registry._ctx._execution_trace = previous
monkeypatch.setenv("OUROBOROS_TASK_REVIEW_MODE", "off")
observed = []
def call(model_call):
active = registry._ctx._execution_trace
assert active is not previous
observed.append(active)
if fail:
raise RuntimeError("fixture stop")
return {"role": "assistant", "content": "done"}, 0.0, model_call.active_context_mode
monkeypatch.setattr(loop, "_call_round_model", call)
if fail:
with pytest.raises(RuntimeError, match="fixture stop"):
loop.run_llm_loop(**_loop_kwargs(tmp_path, registry, []))
else:
_, _, trace = loop.run_llm_loop(**_loop_kwargs(tmp_path, registry, []))
assert observed[0] is trace
assert registry._ctx._execution_trace is previous
def test_canonical_write_failure_keeps_a_bounded_explicit_gap(evidence_context, monkeypatch):
ctx = evidence_context
model_response(ctx, "before", "Before")
tool_response(ctx, "one", "before")
model_response(ctx, "after", "Assessment")
monkeypatch.setattr("ouroboros.artifacts.store_actor_source_bytes", lambda *a, **kw: (_ for _ in ()).throw(OSError("storage failure")))
evidence = capture_commit_review_evidence(ctx)
assert evidence["source_status"] == "unavailable" and evidence["source_ref"] == {}
assert evidence["source_complete"] is False
text = commit_review_evidence_section(evidence, delivery="native")
assert "full source retrieval is unavailable" in text and "Assessment" in text
assert len(text) <= _ACCEPT_NOTES_CAP
@pytest.mark.parametrize("delivery", ["native", "session"])
def test_preflight_execution_receives_source_before_send_and_rejoins_identical_bytes(evidence_context, monkeypatch, delivery):
import ouroboros.tools.claude_advisory_review as advisory
import ouroboros.tools.preflight_review_run as run
import ouroboros.reviewer_slot_config as slots
ctx = evidence_context
model_response(ctx, "before", "Before")
tool_response(ctx, "one", "before")
model_response(ctx, "after", "Frozen assessment")
evidence = capture_commit_review_evidence(ctx)
monkeypatch.setattr(run, "advisory_review_route", lambda: "agent_session" if delivery == "session" else "api_chat")
monkeypatch.setattr(slots, "advisory_slot_config", lambda: SimpleNamespace(target_id="fixture", effort="high", subagent_id="", profile_id="", use_local=False))
monkeypatch.setattr("ouroboros.provider_models.model_has_credentials", lambda *a: True)
monkeypatch.setattr(advisory, "_predispatch_size_skip", lambda *a, **kw: None)
monkeypatch.setattr(advisory, "_api_window_skip_warning", lambda *a, **kw: "")
monkeypatch.setattr(advisory, "_mandatory_read_corpus_chars", lambda *a: 0)
monkeypatch.setattr(advisory, "_build_advisory_prompt", lambda *a, **kw: "Original work order\n" + kw["prompt_context"]["task_evidence_section"])
executions = []
execution = {}
def receive(prompt, repo, current, *args, **kwargs):
assert execution["evidence_source_ref"] == evidence["source_ref"]
source = kwargs["task_evidence"]
assert source["source_ref"] == evidence["source_ref"]
if delivery == "session":
assert pathlib.Path(source["session_path"]).read_bytes() == read_actor_source_bytes(ctx.budget_drive_root, ctx.task_id, evidence["source_ref"])
executions.append(prompt)
return SimpleNamespace(success=True, result_text="[]", source_text="[]", usage={}, session_id="", cost_usd=0.0), "fixture"
monkeypatch.setattr(advisory, "_run_advisory_delegated" if delivery == "session" else "_run_advisory_native", receive)
result = run._run_claude_advisory(ctx.repo_dir, "message", ctx, options={"include_repo_diff": False, "task_evidence": evidence, "execution": execution})
assert result[1] == "[]"
assert "Frozen assessment" in executions[0]
if delivery == "session":
# Rejoin restores from the original canonical source even after trace and
# the disposable view disappear; it does not reconstruct current facts.
execution["pending_invocation_id"] = "pending"
ctx._execution_trace["tool_calls"] = []
shutil.rmtree(ctx.repo_dir / ".review-drive")
monkeypatch.setattr("ouroboros.delegate_custody.invocation_record", lambda *a: {"request": {"prompt": executions[0]}})
again = run._run_claude_advisory(ctx.repo_dir, "message", ctx, options={"execution": execution})
assert again[1] == "[]"
assert executions[1] == executions[0]
@pytest.mark.parametrize("pending", [False, True])
def test_session_copy_lifetime_follows_existing_review_custody(evidence_context, monkeypatch, pending):
from ouroboros.tools import git_review_cycle
ctx = evidence_context
model_response(ctx, "before", "Before")
tool_response(ctx, "one", "before")
model_response(ctx, "after", "Assessment")
ctx._commit_review_evidence = materialize_commit_review_session_view(capture_commit_review_evidence(ctx), ctx.repo_dir)
path = pathlib.Path(ctx._commit_review_evidence["session_path"])
monkeypatch.setattr(git_review_cycle, "_review_custody_pending", lambda c: pending)
git_review_cycle._release_review_evidence_if_settled(ctx)
assert path.exists() is pending
assert read_actor_source_bytes(ctx.budget_drive_root, ctx.task_id, ctx._commit_review_evidence["source_ref"])
@pytest.mark.parametrize("delivery", ["packet", "native", "session"])
def test_scope_request_preserves_selected_source(evidence_context, monkeypatch, delivery):
from ouroboros.tools import scope_review
from ouroboros.review_records import ReviewRouteKind, ReviewSlot
import ouroboros.review_substrate as substrate
ctx = evidence_context
model_response(ctx, "before", "Before")
tool_response(ctx, "one", "before")
model_response(ctx, "after", "Assessment")
evidence = capture_commit_review_evidence(ctx)
captured = []
def receive(request, **kwargs):
captured.append(request)
return SimpleNamespace(actors=[{"status": "ok", "raw_text": "[]", "usage": {}}])
monkeypatch.setattr(substrate, "run_review_request", receive)
monkeypatch.setattr(scope_review, "scope_reviewer_slots", lambda *a, **kw: [ReviewSlot(slot_id="scope", model="fixture")])
monkeypatch.setattr(scope_review, "_scope_window", lambda *a, **kw: SimpleNamespace(sizing_window=lambda *a: 1000000))
scope_review._call_scope_llm("packet", ctx=ctx, scope_model="fixture", slot_id="scope",
route=ReviewRouteKind.AGENT_SESSION if delivery == "session" else ReviewRouteKind.API_CHAT,
session_root=str(ctx.repo_dir), session_task="review",
subagent_id="native" if delivery == "native" else "", task_evidence=evidence)
request = captured[0]
assert request.evidence["task_execution"] == evidence
assert evidence["source_ref"] in request.evidence_refs
assert (request.policy.get("native_data_root") == ctx.budget_drive_root) is (delivery == "native")
@pytest.mark.parametrize("image_state", ["attached", "missing", "reported_failure"])
@pytest.mark.parametrize("skill", ["unix_computer_use", "another_image_producer"])
def test_autoattached_image_process_trace_reaches_commit_evidence(evidence_context, image_state, skill):
from ouroboros.extension_surface_names import extension_surface_name
from ouroboros.loop_tool_execution import process_tool_results
from ouroboros.tools.tool_result import ToolResult
ctx = evidence_context
image = ctx.repo_dir / "captured.png"
if image_state != "missing":
image.write_bytes(base64.b64decode("iVBORw0KGgoAAAANSUhEUgAAAAEAAAABCAQAAAC1HAwCAAAAC0lEQVR42mP8/x8AAwMCAO+j8eUAAAAASUVORK5CYII="))
model_response(ctx, "before", "Inspect the application screenshot")
name = extension_surface_name(skill, "screenshot")
raw = json.dumps({"ok": image_state != "reported_failure", "path": str(image), "auto_attach_image": str(image)})
call = tool_response(ctx, "screen", "before", tool=name, result=raw)
ctx._execution_trace["tool_calls"].clear()
ctx.messages = [{"role": "assistant", "content": "", "tool_calls": [
{"id": call_id, "type": "function", "function": {"name": fn, "arguments": "{}"}}
for call_id, fn in [("screen", name), ("read", "read_file")]]}]
rows = [{"fn_name": name, "tool_call_id": "screen", "result": raw, "trace_ref": call["trace_ref"],
"is_error": False, "tool_args": call["args"], "args_for_log": call["args"], "result_meta": {}},
{"fn_name": "read_file", "tool_call_id": "read", "result": "ordinary result",
"is_error": False, "tool_args": {}, "args_for_log": {}, "result_meta": {}}]
if image_state == "reported_failure":
rows[0]["tool_result"] = ToolResult(status="error", code="TOOL_REPORTED_FAILURE", text=raw)
assert process_tool_results(rows, ctx.messages, ctx._execution_trace, lambda _: None,
tools=SimpleNamespace(_ctx=ctx)) == 0
assert [m["role"] for m in ctx.messages[:3]] == ["assistant", "tool", "tool"]
pictures = [b for m in ctx.messages if isinstance(m.get("content"), list)
for b in m["content"] if b.get("type") == "image_url"]
assert bool(pictures) is (image_state == "attached")
assert ctx._execution_trace["tool_calls"][1].get("image_attachment") is None
if pictures:
assert pathlib.Path(pictures[0]["_source_path"]).read_bytes() == image.read_bytes()
model_response(ctx, "after", [{"type": "thinking", "text": "PRIVATE_THINKING"},
{"type": "text", "text": "IMMEDIATE_VISIBLE_ASSESSMENT"}])
evidence = capture_commit_review_evidence(ctx)
if image_state == "reported_failure":
assert evidence == {}
assert "image_attachment" not in ctx._execution_trace["tool_calls"][0]
else:
assert evidence["selected_count"] == 1 and evidence["unselected_count"] == 1
exact = read_actor_source_bytes(ctx.budget_drive_root, ctx.task_id, evidence["source_ref"]).decode()
assert name in exact and raw in json.loads(exact.split("\n\n", 2)[2])["result"]
assert '"image_attachment"' in exact and "IMMEDIATE_VISIBLE_ASSESSMENT" in exact
assert "PRIVATE_THINKING" not in exact
assert evidence["source_complete"] is (image_state == "attached")
assert "proof of visual inspection" in exact
@pytest.mark.parametrize("recorded", [False, True])
@pytest.mark.parametrize("current_trace", ["changed", "empty"])
def test_pending_preflight_selection_survives_reconciliation_then_stage_dispatch(
evidence_context, monkeypatch, recorded, current_trace,
):
from ouroboros.review_state import AdvisoryRunRecord, compute_snapshot_hash, make_repo_key, update_state
from ouroboros.tools import claude_advisory_review as advisory, git, preflight_review_run as preflight, scope_review
from ouroboros.tools.review_multi_model import _query_model
from ouroboros.review_records import ReviewSlot
import ouroboros.review_substrate as substrate
ctx = evidence_context
(ctx.repo_dir / "README.md").write_text("candidate\n")
subprocess.run(["git", "add", "README.md"], cwd=ctx.repo_dir, check=True)
model_response(ctx, "before", "Before")
tool_response(ctx, "screen", "before")
model_response(ctx, "after", "ORIGINAL_ASSESSMENT")
frozen = capture_commit_review_evidence(ctx) if recorded else {}
execution = {"invocation_id": "pending-preflight", "pending_invocation_id": "pending-preflight", "operation_state": "in_flight",
"fingerprint": git._fingerprint_staged_diff(ctx.repo_dir)["fingerprint"],
"intent": {"commit_message": "candidate", "goal": "", "scope": "", "review_rebuttal": ""}}
if recorded:
execution["evidence_source_ref"] = frozen["source_ref"]
update_state(ctx.drive_root, lambda state: state.add_run(AdvisoryRunRecord(
snapshot_hash=compute_snapshot_hash(ctx.repo_dir, paths=["README.md"]),
commit_message="candidate", status="pending", ts="2026-09-09T00:00:00Z",
repo_key=make_repo_key(ctx.repo_dir), task_id=ctx.task_id, snapshot_paths=["README.md"], execution=execution)))
ctx._execution_trace["tool_calls"].clear()
if current_trace == "changed":
tool_response(ctx, "different", "before", result="NEW_TOOL_OBSERVATION")
ctx._commit_review_evidence = {"preview": "STALE_CONTEXT"}
monkeypatch.setattr(preflight, "advisory_review_route", lambda: "agent_session")
monkeypatch.setattr(advisory, "advisory_review_route", lambda: "agent_session")
monkeypatch.setattr(advisory, "advisory_slot_enabled", lambda: True)
monkeypatch.setattr(advisory, "check_worktree_readiness", lambda *a, **kw: [])
monkeypatch.setattr(advisory, "_release_metadata_preflight", lambda *a: None)
monkeypatch.setattr(advisory, "_check_worktree_version_sync_shared", lambda *a: "")
monkeypatch.setattr("ouroboros.delegate_custody.invocation_record", lambda *a: {"request": {"prompt": "ORIGINAL_PREFLIGHT_PROMPT"}})
sent = []
def rejoin(prompt, _repo, _ctx, **kwargs):
assert prompt == "ORIGINAL_PREFLIGHT_PROMPT"
sent.append(("preflight", kwargs.get("task_evidence", {}).get("source_ref", {})))
return SimpleNamespace(success=True, result_text='[{"item":"fixture","verdict":"PASS","severity":"advisory","reason":"recorded"}]',
source_text="", usage={}, session_id="existing", cost_usd=0), "fixture"
monkeypatch.setattr(advisory, "_run_advisory_delegated", rejoin)
git._reset_commit_review_state(ctx)
assert git._reconcile_advisory_before_preparation(ctx, "candidate", goal="", scope="", paths=["README.md"], review_rebuttal="") == ""
assert ctx._advisory_reconciled
assert ctx._commit_review_evidence.get("source_ref", {}) == frozen.get("source_ref", {})
monkeypatch.setattr("ouroboros.review_evidence.capture_commit_review_evidence", lambda *a: pytest.fail("rejoin must not capture the current trace"))
monkeypatch.setattr(git, "_free_cycle_gate", lambda *a, **kw: None)
monkeypatch.setattr(git, "_advisory_and_tests_gate", lambda *a, **kw: None)
monkeypatch.setattr(git, "_install_paid_dispatch_stamp", lambda *a: None)
def receive(request, **kwargs):
sent.append((request.surface, request.evidence.get("task_execution", {}).get("source_ref", {})))
return SimpleNamespace(actors=[{"status": "ok", "raw_text": "[]", "usage": {}}])
monkeypatch.setattr(substrate, "run_review_request", receive)
monkeypatch.setattr(scope_review, "scope_reviewer_slots", lambda *a, **kw: [ReviewSlot(slot_id="scope", model="fixture")])
monkeypatch.setattr(scope_review, "_scope_window", lambda *a, **kw: SimpleNamespace(sizing_window=lambda *a: 1000000))
def dispatch(_ctx, *args, **kwargs):
asyncio.run(_query_model(None, "fixture", [{"role": "user", "content": "packet"}], asyncio.Semaphore(1),
_ctx, task_evidence=_ctx._commit_review_evidence, use_local=False))
scope_review._call_scope_llm("packet", ctx=_ctx, scope_model="fixture", slot_id="scope",
task_evidence=_ctx._commit_review_evidence)
return None, None, "", []
monkeypatch.setattr(git, "_run_parallel_review", dispatch)
outcome = git._run_reviewed_stage_cycle(ctx, "candidate", time.time(), paths=["README.md"], require_release_tag=False)
assert outcome["status"] == "passed", outcome
assert sent == [(surface, frozen.get("source_ref", {})) for surface in ("preflight", "multi_model_review", "scope_review")]
if not recorded:
assert ctx._commit_review_evidence == {}
def test_fresh_stage_captures_current_evidence_after_rejoin_flag_reset(evidence_context, monkeypatch):
from ouroboros.tools import git
ctx = evidence_context
(ctx.repo_dir / "README.md").write_text("fresh candidate\n")
ctx._advisory_reconciled = True
ctx._commit_review_evidence = {"preview": "OLD_REJOIN"}
model_response(ctx, "before", "Before")
tool_response(ctx, "new", "before")
model_response(ctx, "after", "NEW_ASSESSMENT")
assert git._reconcile_advisory_before_preparation(ctx, "fresh", goal="", scope="", paths=["README.md"], review_rebuttal="") == ""
assert not ctx._advisory_reconciled
monkeypatch.setattr(git, "_free_cycle_gate", lambda *a, **kw: None)
monkeypatch.setattr(git, "_advisory_and_tests_gate", lambda *a, **kw: None)
monkeypatch.setattr(git, "_install_paid_dispatch_stamp", lambda *a: None)
observed = []
monkeypatch.setattr(git, "_run_parallel_review", lambda *a, **kw: (observed.append(ctx._commit_review_evidence) or None, None, "", []))
result = git._run_reviewed_stage_cycle(ctx, "fresh", time.time(), paths=["README.md"], require_release_tag=False)
assert result["status"] == "passed"
assert len(observed) == 1 and "NEW_ASSESSMENT" in observed[0]["preview"]
@pytest.mark.parametrize("surface", ["triad", "scope"])
def test_real_packet_assembly_omits_optional_excerpt_before_required_material(evidence_context, monkeypatch, surface):
from ouroboros.review_records import ReviewRouteKind
from ouroboros.tools import review, review_admission, scope_review, scope_review_pack
ctx = evidence_context
for path in ("BIBLE.md", "docs/DEVELOPMENT.md", "docs/DESIGN.md", "docs/ARCHITECTURE.md", "docs/CHECKLISTS.md"):
target = ctx.repo_dir / path
target.parent.mkdir(exist_ok=True)
target.write_text("GOVERNANCE_MARKER " + path)
(ctx.repo_dir / "README.md").write_text("MANDATORY_SNAPSHOT\n")
subprocess.run(["git", "add", "README.md"], cwd=ctx.repo_dir, check=True)
model_response(ctx, "before", "Before")
tool_response(ctx, "one", "before")
model_response(ctx, "after", "OPTIONAL_IMAGE_EXCERPT\n" * 400)
evidence = capture_commit_review_evidence(ctx)
ctx._commit_review_evidence = evidence
cap = [10**9]
monkeypatch.setattr(review_admission, "density_probe_before_size_refusal", lambda *a, **kw: pytest.fail("optional excerpt should fit before paid density probe"))
if surface == "triad":
monkeypatch.setattr(review, "_preflight_check", lambda *a: None)
monkeypatch.setattr(review, "_load_checklist_section", lambda: "CHECKLIST_MARKER")
monkeypatch.setattr("ouroboros.reviewer_slot_config.commit_triad_delivery", lambda: {
"models": ["fixture"], "routes": [ReviewRouteKind.API_CHAT], "slot_ids": ["triad-one"],
"session_profiles": [""], "subagent_ids": [""], "use_local": [False]})
monkeypatch.setattr(review, "reviewer_context_window", lambda *a, **kw: 1000000)
monkeypatch.setattr(review, "calibrated_input_token_limit", lambda *a, **kw: cap[0])
monkeypatch.setattr(review, "estimate_tokens", len)
def build():
prepared, early, exited = review._prepare_unified_review(ctx, "candidate", goal="INTENT_MARKER", review_rebuttal="REBUTTAL_MARKER")
assert not exited and early is None
return prepared["prompt"], prepared["stable_prefix_len"]
else:
monkeypatch.setattr(scope_review, "load_checklist_section", lambda *a: "CHECKLIST_MARKER")
monkeypatch.setattr(scope_review, "estimate_tokens", len)
monkeypatch.setattr(scope_review, "_effective_scope_input_limit", lambda **kw: cap[0])
monkeypatch.setattr(scope_review_pack, "_gather_scope_packs", lambda *a, **kw: "ATLAS_MARKER")
def build():
prompt, status = scope_review_pack._build_scope_prompt(ctx.repo_dir, "candidate", goal="INTENT_MARKER", review_rebuttal="REBUTTAL_MARKER",
context=scope_review_pack._ScopePromptContext(task_evidence=evidence))
assert status is None
return prompt, scope_review_pack._SCOPE_STABLE_PREFIX_LEN.get()
full, prefix = build()
long_exhibit = commit_review_evidence_section(evidence, delivery="packet")
short_exhibit = commit_review_evidence_section(evidence, delivery="packet", compact=True)
assert long_exhibit in full
expected = full.replace(long_exhibit, short_exhibit)
cap[0] = len(expected) + 1
assert len(full) > cap[0]
fitted, next_prefix = build()
assert fitted == expected and next_prefix == prefix
assert full[:prefix] == fitted[:prefix]
for marker in ("GOVERNANCE_MARKER", "CHECKLIST_MARKER", "INTENT_MARKER", "REBUTTAL_MARKER", "MANDATORY_SNAPSHOT", "+MANDATORY_SNAPSHOT"):
assert marker in fitted
assert "OPTIONAL_IMAGE_EXCERPT" not in fitted and "excerpt omitted to fit" in fitted
assert ctx._commit_review_evidence == evidence
if surface == "scope":
steps = scope_review_pack._current_scope_context_manifest()["ladder_steps"]
assert any(step["step"] == "task_evidence_excerpt_omitted" for step in steps)
assert all(not step.get("diff_only_files") and not step.get("zero_context_diff") for step in steps)

View file

@ -246,3 +246,25 @@ def test_scope_receipt_validator_cli_edges(tmp_path, capsys):
assert main(["validate", str(tmp_path / "absent.json")]) == 1 assert main(["validate", str(tmp_path / "absent.json")]) == 1
assert "cannot read receipt" in capsys.readouterr().err assert "cannot read receipt" in capsys.readouterr().err
def test_scope_receipt_survives_trailing_markdown_checkboxes(tmp_path, capsys):
import json
from scripts.validate_scope_receipt import main
from ouroboros.tools.scope_review_contract import SCOPE_REQUIRED_ITEMS
from ouroboros.triad_review import extract_json_array
rows = [{"item": item, "verdict": "FAIL" if i == 0 else "PASS", "severity": "advisory",
"reason": f"Checked the concrete {item} source and its consumers."}
for i, item in enumerate(sorted(SCOPE_REQUIRED_ITEMS))]
path = tmp_path / "receipt.md"
for prefix, suffix in [("", "\n- [ ] Follow up"), ("- [ ] Before\n", "\n[]\n- [ ] After"), ("```json\n", "\n```\n- [ ] Remaining")]:
path.write_text(prefix + json.dumps(rows) + suffix)
assert main(["validate", str(path)]) == 0
assert "1 FAIL row(s)" in capsys.readouterr().out
# Whole JSON objects are not repaired by selecting a nested valid array.
path.write_text(json.dumps({"nested": rows}))
assert main(["validate", str(path)]) == 1
path.write_text("- [ ] Only a checkbox")
assert main(["validate", str(path)]) == 1
assert extract_json_array("[]") == []

View file

@ -147,7 +147,9 @@ def test_desktop_bridge_version_skew_fallback_chain():
chain.""" chain."""
helper = _read("web/modules/ui_helpers.js") helper = _read("web/modules/ui_helpers.js")
assert "api?.open_external_url" in helper assert "export async function openExternalViaHostBridge(" in helper
assert "api.open_external_url ? await api.open_external_url(target) : null" in helper
assert "await openExternalViaHostBridge(url, { api, win, doc, toast });" in helper
assert "'Link copied — open it in your browser.'" in helper assert "'Link copied — open it in your browser.'" in helper
assert "api.save_bytes_to_downloads(payload.name, payload.b64)" in helper assert "api.save_bytes_to_downloads(payload.name, payload.b64)" in helper
assert "downloadBlobViaHostBridge(url, filename, { win, doc })" in helper assert "downloadBlobViaHostBridge(url, filename, { win, doc })" in helper

View file

@ -213,7 +213,7 @@ class TestAliasProducerFanOutSweep:
("ouroboros/tools/preflight_review_run.py", "cost_usd", "_llm_extract_advisory_items"): ("advisory preflight usage receipt", 1), ("ouroboros/tools/preflight_review_run.py", "cost_usd", "_llm_extract_advisory_items"): ("advisory preflight usage receipt", 1),
("ouroboros/tools/preflight_review_run.py", "cost_usd", "_advisory_failure"): ("internal advisory failure adapter; physical charges remain in usage/custody, not gateway fields", 1), ("ouroboros/tools/preflight_review_run.py", "cost_usd", "_advisory_failure"): ("internal advisory failure adapter; physical charges remain in usage/custody, not gateway fields", 1),
("ouroboros/tools/preflight_review_run.py", "cost_usd", "_run_advisory_delegated"): ("advisory preflight receipt", 1), ("ouroboros/tools/preflight_review_run.py", "cost_usd", "_run_advisory_delegated"): ("advisory preflight receipt", 1),
("ouroboros/tools/preflight_review_run.py", "cost_usd", "_run_claude_advisory"): ("advisory preflight receipt", 4), ("ouroboros/tools/preflight_review_run.py", "cost_usd", "_run_claude_advisory"): ("single advisory receipt cost reused by event projections", 1),
("ouroboros/tools/review_admission.py", "cost_usd", "triad_not_dispatched_records"): ("review admission receipt", 1), ("ouroboros/tools/review_admission.py", "cost_usd", "triad_not_dispatched_records"): ("review admission receipt", 1),
("ouroboros/tools/review_helpers.py", "cost_usd", "build_scope_actor_record"): ("review usage receipt", 1), ("ouroboros/tools/review_helpers.py", "cost_usd", "build_scope_actor_record"): ("review usage receipt", 1),
("ouroboros/tools/scope_review.py", "cost_usd", "_scope_oversize_result"): ("scope review receipt", 1), ("ouroboros/tools/scope_review.py", "cost_usd", "_scope_oversize_result"): ("scope review receipt", 1),

View file

@ -167,7 +167,8 @@ def test_headroom_is_derived_from_the_zero_diff_message(synthetic_repo, isolated
stable = mrp._governance_prefix(synthetic_repo)["stable_prefix"] stable = mrp._governance_prefix(synthetic_repo)["stable_prefix"]
dynamic = review._REVIEW_PROMPT_TEMPLATE_DYNAMIC.format( dynamic = review._REVIEW_PROMPT_TEMPLATE_DYNAMIC.format(
goal_section=build_goal_section("", "", ""), scope_section="", current_files_section="", goal_section=build_goal_section("", "", ""), scope_section="", current_files_section="",
rebuttal_section="", review_history_section="", diff_text="", changed_files="app.py") rebuttal_section="", review_history_section="", diff_text="", changed_files="app.py",
task_evidence_section="")
zero_message = head + stable + "\n" + dynamic + mrp.TRIAD_USER_TURN zero_message = head + stable + "\n" + dynamic + mrp.TRIAD_USER_TURN
assert report["zero_diff_message"]["total"]["chars"] == len(zero_message) assert report["zero_diff_message"]["total"]["chars"] == len(zero_message)
assert fit["zero_diff_message_chars_div_4"] == estimate_tokens(zero_message) assert fit["zero_diff_message_chars_div_4"] == estimate_tokens(zero_message)

View file

@ -458,6 +458,7 @@ LEAVES: dict[str, tuple[str, str, frozenset[str]]] = {
})), })),
"ouroboros/tools/preflight_review_run.py": ("ouroboros/tools/claude_advisory_review.py", "_car", frozenset({ "ouroboros/tools/preflight_review_run.py": ("ouroboros/tools/claude_advisory_review.py", "_car", frozenset({
"SEVERITY_DRIVEN_ITEMS", "_advisory_native_model", "_advisory_review_diff", "SEVERITY_DRIVEN_ITEMS", "_advisory_native_model", "_advisory_review_diff",
"_api_window_skip_warning",
"_build_advisory_prompt", "_format_advisory_error", "_get_changed_file_list", "_build_advisory_prompt", "_format_advisory_error", "_get_changed_file_list",
"_get_runtime_diagnostics", "_llm_extract_advisory_items", "_get_runtime_diagnostics", "_llm_extract_advisory_items",
"_mandatory_read_corpus_chars", "_maybe_overflow_skip", "_predispatch_size_skip", "_persist_preflight_record", "_mandatory_read_corpus_chars", "_maybe_overflow_skip", "_predispatch_size_skip", "_persist_preflight_record",

View file

@ -0,0 +1,21 @@
"""The authenticated Telegram SPA retains its presentation SDK after bootstrap."""
import pathlib
from starlette.applications import Starlette
from starlette.routing import Route
from starlette.testclient import TestClient
from ouroboros.server_web import make_index_page
def test_telegram_marker_only_changes_the_host_document():
web = pathlib.Path(__file__).resolve().parents[1] / "web"
with TestClient(Starlette(routes=[Route("/", make_index_page(web))])) as client:
plain = client.get("/")
assert plain.content == (web / "index.html").read_bytes()
marked = client.get("/", headers={"X-Ouroboros-Telegram-MiniApp": "1"})
assert marked.status_code == 200
assert '<html data-ouroboros-host="telegram"' in marked.text
assert '<script async src="https://telegram.org/js/telegram-web-app.js"></script>' in marked.text
assert 'data-ouroboros-host' not in plain.text
assert client.get("/", headers={"X-Ouroboros-Telegram-MiniApp": "0"}).content == plain.content

View file

@ -350,7 +350,7 @@ def test_read_version_returns_empty_for_missing_files(tmp_path):
assert _read_skill_manifest_version(skill) == "" assert _read_skill_manifest_version(skill) == ""
def test_telegram_owner_wait_upgrade_reseeds_version_1_2_1(tmp_path, fake_log): def test_telegram_owner_wait_upgrade_reseeds_current_version(tmp_path, fake_log):
"""The launcher-owned Telegram payload must deliver the owner-wait update. """The launcher-owned Telegram payload must deliver the owner-wait update.
The pinned string is the RESYNC KEY, not decoration: `_per_skill_version_resync` The pinned string is the RESYNC KEY, not decoration: `_per_skill_version_resync`
@ -374,6 +374,96 @@ def test_telegram_owner_wait_upgrade_reseeds_version_1_2_1(tmp_path, fake_log):
) )
assert upgraded == 1 assert upgraded == 1
assert "version: 1.2.1" in (installed / "SKILL.md").read_text(encoding="utf-8") assert "version: 1.2.2" in (installed / "SKILL.md").read_text(encoding="utf-8")
for path in ("plugin.py", "lib/telegram_quiz.py"): for path in ("plugin.py", "lib/telegram_quiz.py"):
assert (installed / path).read_bytes() == (seed_dir / "telegram" / path).read_bytes() assert (installed / path).read_bytes() == (seed_dir / "telegram" / path).read_bytes()
@pytest.mark.serial
@pytest.mark.parametrize("name,source,old_version,new_version", [
("telegram", "d5418e05b822feaf6aaa652e8cdc5b53af1232cc", "1.2.1", "1.2.2"),
("unix_computer_use", "162ad3fe6791fcaf6cf625e6b0c50d3a2a27e7f8", "0.4.1", "0.4.2"),
])
def test_resync_delivers_payload_from_real_previous_seed(tmp_path, fake_log, name, source, old_version, new_version):
"""Use the full seed before 59ce693b / 3f8db1e1, including its real payload.
These official history objects are available in CI's full checkout; a missing
object is a fixture error, not evidence that an upgrade was exercised.
"""
import io
import subprocess
import tarfile
from ouroboros.launcher_bootstrap import _per_skill_version_resync, _read_skill_manifest
from ouroboros.skill_loader import compute_content_hash
repo = pathlib.Path(__file__).resolve().parents[1]
drive = tmp_path / "data"
native = drive / "skills" / "native"
installed = native / name
installed.mkdir(parents=True)
archived = subprocess.run(["git", "archive", f"{source}:skills/{name}"], cwd=repo,
capture_output=True, check=True)
with tarfile.open(fileobj=io.BytesIO(archived.stdout)) as archive:
for member in archive:
if member.isdir():
continue
target = installed / member.name
assert member.isfile() and target.resolve().is_relative_to(installed.resolve())
target.parent.mkdir(parents=True, exist_ok=True)
target.write_bytes(archive.extractfile(member).read())
(installed / ".seed-origin").write_text(f"seeded_from={source}\n", encoding="utf-8")
assert _read_skill_manifest(installed).version == old_version
assert _per_skill_version_resync(repo / "skills", native, fake_log, drive_root=drive) == 1
manifest = _read_skill_manifest(installed)
assert manifest.version == new_version
hash_args = {"manifest_entry": manifest.entry, "manifest_scripts": manifest.scripts}
assert compute_content_hash(installed, **hash_args) == compute_content_hash(repo / "skills" / name, **hash_args)
@pytest.mark.parametrize("drift", [False, True])
def test_same_version_hash_drift_is_only_diagnostic(staging, fake_log, caplog, drift):
from ouroboros.launcher_bootstrap import _per_skill_version_resync
seed_dir, native_root, drive_root = staging
seed = _write_skill(seed_dir, "weather")
installed = _write_skill(native_root, "weather")
(installed / ".seed-origin").write_text("seeded_from=test\n")
(seed / "payload.txt").write_text("seed")
(installed / "payload.txt").write_text("local" if drift else "seed")
state = drive_root / "state" / "skills" / "weather"
state.mkdir(parents=True)
for name in ("enabled.json", "grants.json", "review.json"):
(state / name).write_text('{"unchanged":true}')
before = {str(path.relative_to(drive_root)): path.read_bytes()
for path in drive_root.rglob("*") if path.is_file()}
with caplog.at_level(logging.WARNING, logger=fake_log.name):
assert _per_skill_version_resync(seed_dir, native_root, fake_log, drive_root=drive_root) == 0
after = {str(path.relative_to(drive_root)): path.read_bytes()
for path in drive_root.rglob("*") if path.is_file()}
assert after == before
assert ("installed files retained because the manifest version is unchanged" in caplog.text) is drift
assert len(caplog.records) == int(drift)
def test_hash_comparison_failure_does_not_stop_next_skill(staging, fake_log, caplog, monkeypatch):
from ouroboros.launcher_bootstrap import _per_skill_version_resync
import ouroboros.skill_loader as loader
seed_dir, native_root, drive_root = staging
for name in ("first", "second"):
_write_skill(seed_dir, name, "1.0.0" if name == "first" else "2.0.0")
installed = _write_skill(native_root, name)
(installed / ".seed-origin").write_text("seeded_from=test\n")
original = loader.compute_content_hash
def unreadable(path, **kwargs):
if path.name == "first":
raise OSError("private error detail")
return original(path, **kwargs)
monkeypatch.setattr(loader, "compute_content_hash", unreadable)
with caplog.at_level(logging.WARNING, logger=fake_log.name):
assert _per_skill_version_resync(seed_dir, native_root, fake_log, drive_root=drive_root) == 1
assert "comparison unavailable (OSError)" in caplog.text
assert "private error detail" not in caplog.text
assert "version: 1.0.0" in (native_root / "first" / "SKILL.md").read_text()
assert "version: 2.0.0" in (native_root / "second" / "SKILL.md").read_text()

View file

@ -12,6 +12,8 @@ its own module because the engine test file sits at its 1500-line ceiling.
from __future__ import annotations from __future__ import annotations
import pytest
import dataclasses import dataclasses
import json import json
import logging import logging
@ -444,3 +446,56 @@ def test_transient_and_unknown_health_still_fail_open_for_a_pinned_slot(monkeypa
_patch_snapshot_health(monkeypatch, lambda rid, model, pin, r=reason, t=reset: (r, t)) _patch_snapshot_health(monkeypatch, lambda rid, model, pin, r=reason, t=reset: (r, t))
assert plan_panel_health_snapshot( assert plan_panel_health_snapshot(
_profile_slots(("s1", "codex=gpt-5.6-sol", "spent-acct"))) == {}, reason _profile_slots(("s1", "codex=gpt-5.6-sol", "spent-acct"))) == {}, reason
@pytest.mark.parametrize("aggregate", ["REVIEW_REQUIRED", "REVISE_PLAN", "DEGRADED"])
@pytest.mark.parametrize("enforcement", ["blocking", "advisory"])
@pytest.mark.parametrize("epoch", ["", "health"])
@pytest.mark.parametrize("cap", [None, 2, 3])
def test_whole_plan_render_respects_paid_capacity(aggregate, enforcement, epoch, cap):
from ouroboros.tools.plan_render import _render_wave, _parse_plan_review_control
import copy
wave = {"aggregate": aggregate, "closed": False, "request_fingerprint": "fp", "health_epoch": epoch,
"counts": {"configured": 3, "parseable": 0, "quorum": 2},
"findings": [{"class": "blocking", "finding_id": "s1:f1"}],
"closure_notes": ["blocking_finding_below_quorum_stays_open: let the next paid delta cycle judge the rejection",
"revise_plan_not_closable_by_disposition: next paid delta cycle",
"degraded_not_closable_by_disposition: rerun the wave",
"author note: retained unchanged"]}
before = copy.deepcopy(wave)
text = _render_wave(wave, cap=cap, cycles_paid=2, enforcement=enforcement)
assert wave == before
assert _parse_plan_review_control(text) == (aggregate, False)
assert "author note: retained unchanged" in text
assert "next paid delta cycle" not in text and "rerun the wave" not in text
if cap == 2:
assert "cycle cap is reached" in text
assert "re-dispatches a fresh panel" not in text
assert "A changed spec may start" not in text
elif aggregate == "DEGRADED":
assert "A changed spec may start another paid cycle" in text
else:
assert "another paid cycle" in text or "subsequent paid delta review" in text
def test_closed_and_pending_plan_states_do_not_advertise_new_review_at_cap():
from ouroboros.tools.plan_render import _next_step
assert _next_step({"aggregate": "GREEN", "closed": True}, enforcement="blocking", cap=2, cycles_paid=2).startswith("Closed: proceed")
pending = _next_step({"aggregate": "DEGRADED", "custody_pending": True}, enforcement="blocking", cap=2, cycles_paid=2)
assert "custody reconciliation" in pending and "fresh panel" not in pending
def test_loop_reminder_does_not_repromise_a_spent_panel(monkeypatch):
from ouroboros.owner_hurry import plan_review_reminder
monkeypatch.setenv("OUROBOROS_REVIEW_MAX_CYCLES", "2")
for outcome in ("REVISE_PLAN", "REVIEW_REQUIRED", "DEGRADED"):
text = plan_review_reminder({"outcome": outcome, "status": "open", "cycles_paid": 2,
"reviewer_slots_degraded": outcome == "DEGRADED"})
assert "cannot dispatch another paid panel" in text
assert "re-dispatches a fresh panel" not in text
assert "a changed spec starts the next paid cycle" not in text
monkeypatch.setenv("OUROBOROS_REVIEW_MAX_CYCLES", "unlimited")
assert "re-dispatches a fresh panel" in plan_review_reminder({"outcome": "DEGRADED", "reviewer_slots_degraded": True, "cycles_paid": 2})

View file

@ -360,6 +360,7 @@ def test_triad_review_prompt_includes_architecture_md(tmp_path):
review_history_section="", review_history_section="",
diff_text="DIFF", diff_text="DIFF",
changed_files="changed_file.py", changed_files="changed_file.py",
task_evidence_section="",
) )
assert "UNIQUE_MARKER_12345" in rendered, ( assert "UNIQUE_MARKER_12345" in rendered, (
"ARCHITECTURE.md content must appear in the rendered triad review prompt" "ARCHITECTURE.md content must appear in the rendered triad review prompt"

View file

@ -119,15 +119,19 @@ def test_shell_catalog_schema_bytes_and_handler_owners_are_stable():
ensure_ascii=False, ensure_ascii=False,
separators=(",", ":"), separators=(",", ":"),
).encode() ).encode()
# Re-pinned for upstream ``d0caa69b`` ("tool refusals: name the unsupported # Only run_script's description changed: its temporary file now lives in
# argument and own the timeout alias once"), which deleted the duplicate # an ignored per-invocation workspace directory or the existing task drive.
# ``timeout`` alias property from BOTH schemas because the alias now lives
# once in ``tool_resolution._TOOL_ARG_ALIASES["*"]``. That deletion is the
# only difference from the previous digest
# ``1e012faf410bf57c91a896d227aa4175cd08d38794c4b8f0404e390e79b1730a``:
# re-inserting the two removed properties into these very schema objects
# reproduces it byte for byte.
assert hashlib.sha256(schema_bytes).hexdigest() == ( assert hashlib.sha256(schema_bytes).hexdigest() == (
"e053d8164ac708c9053b5ea9cf5271b70c1a05667a33f14253101bb2b1ce9cfd"
)
original = json.loads(schema_bytes)
original[1]["description"] = (
"Run a short task-scoped temporary script with a declared interpreter. "
"Use for multi-line diagnostics or harness helpers; generated script files live under the task drive. "
"The underlying command result echoes the resolved cwd."
)
assert hashlib.sha256(json.dumps(original, sort_keys=True, ensure_ascii=False,
separators=(",", ":")).encode()).hexdigest() == (
"0c0215cafdb54ee2231e43f9edafbe8aa21d4a5aa3841eba4ab45b5ab5e3eb42" "0c0215cafdb54ee2231e43f9edafbe8aa21d4a5aa3841eba4ab45b5ab5e3eb42"
) )
assert { assert {

View file

@ -218,3 +218,17 @@ def test_review_skill_tool_result_has_no_raw_json_block(tmp_path, monkeypatch):
out = skill_exec_mod._handle_review_skill(ctx, skill="alpha") out = skill_exec_mod._handle_review_skill(ctx, skill="alpha")
assert "Raw review payload" not in out assert "Raw review payload" not in out
assert "<details>" not in out assert "<details>" not in out
def test_history_preserves_real_round_and_snapshot_numbers():
from ouroboros.skill_review import _build_skill_review_history_section
history = [{"review_round": n, "snapshot_attempt": n - 5, "snapshot_revised": n == 9,
"status": "warnings", "content_hash": "abc", "failure_signature": [f"reason-{n}"]}
for n in range(6, 10)]
text = _build_skill_review_history_section(history)
assert "Review round 7, snapshot attempt 2" in text
assert "Review round 9, snapshot attempt 4" in text
assert "snapshot_revised=True" in text and "reason-9" in text
assert "Review round 6" not in text and "Attempt 1:" not in text
legacy = _build_skill_review_history_section([{"status": "warnings"}])
assert "Review round unknown, snapshot attempt unknown" in legacy

View file

@ -81,6 +81,7 @@ def test_tool_descriptor_owner_facades_preserve_identity():
("event_queue", None), ("event_queue", None),
("task_id", None), ("task_id", None),
("messages", None), ("messages", None),
("_execution_trace", None),
("task_constraint", None), ("task_constraint", None),
("task_contract", "factory:dict"), ("task_contract", "factory:dict"),
("task_depth", 0), ("task_depth", 0),

View file

@ -374,7 +374,7 @@ def test_headless_excludes_declared_scratch_from_workspace_patch(tmp_path):
def test_check_has_exit_masking_detection(): def test_check_has_exit_masking_detection():
from ouroboros.tools.verify import _check_has_exit_masking from ouroboros.tools.verify import _check_has_exit_masking
assert _check_has_exit_masking(["sh", "-c", "node t.js -f 2>&1 | tail -5"])[0] is True assert _check_has_exit_masking(["sh", "-c", "(node t.js -f 2>&1 | tail -5)"])[0] is True
assert _check_has_exit_masking(["bash", "-c", "make test || true"])[0] is True assert _check_has_exit_masking(["bash", "-c", "make test || true"])[0] is True
assert _check_has_exit_masking(["sh", "-c", "run.sh 2>/dev/null"])[0] is False assert _check_has_exit_masking(["sh", "-c", "run.sh 2>/dev/null"])[0] is False
assert _check_has_exit_masking(["sh", "-c", "make test ; true"])[0] is True assert _check_has_exit_masking(["sh", "-c", "make test ; true"])[0] is True
@ -513,3 +513,24 @@ def test_masked_verification_nudge_one_shot_advisory_and_ordering(tmp_path):
) )
assert fired is True assert fired is True
assert any("RED" in m.get("content", "") for m in msgs2) assert any("RED" in m.get("content", "") for m in msgs2)
@pytest.mark.parametrize("head", ["sh", "bash", "zsh", "dash", "ash"])
@pytest.mark.parametrize("absolute", [False, True])
def test_subshell_masking_uses_typed_shell_grammar(head, absolute):
from ouroboros.tools.verify import check_exit_masking
executable = "/bin/" + head if absolute else head
for command, reason in [
("make test|tail", "pipeline_tail"),
("make test||true", "|| true"),
("make test; true", "; true"),
("make test; exit 0", "exit 0"),
("make test|'tail'", "pipeline_tail"),
]:
expected = check_exit_masking([executable, "-c", command])
assert expected == (True, [reason])
assert check_exit_masking([executable, "-c", "(" + command + ")"]) == expected
assert check_exit_masking([executable, "-c", "(false)|(tail)"]) == (True, ["pipeline_tail"])
for literal in ["echo '|' tail", "echo '||' true", "echo ';' true", "echo '(' false '|tail)'", r"echo \(false\|tail\)"]:
assert check_exit_masking([executable, "-c", literal]) == (False, [])

View file

@ -10,6 +10,8 @@ resolver; a resolution failure fails closed with the standard cwd block.
from __future__ import annotations from __future__ import annotations
import pathlib import pathlib
import shlex
from subprocess import CompletedProcess
import pytest import pytest
@ -191,3 +193,98 @@ def test_light_mode_versioned_interpreter_triggers_runtime_data_scan(tmp_path, m
# same way — the versioned basename must not be the weaker path. # same way — the versioned basename must not be the weaker path.
unversioned = reg.execute_result("run_command", {"cmd": ["python", "-c", read_cmd]}) unversioned = reg.execute_result("run_command", {"cmd": ["python", "-c", read_cmd]})
assert (unversioned.status, unversioned.code) == (result.status, result.code) assert (unversioned.status, unversioned.code) == (result.status, result.code)
@pytest.mark.parametrize("head", ["sh", "bash", "zsh", "dash", "ash"])
@pytest.mark.parametrize("tool_name", ["run_command", "verify_and_record"])
def test_posix_wrappers_preserve_light_read_and_deliverable_guards(
tmp_path, monkeypatch, head, tool_name,
):
"""Every accepted shell spelling reaches the same existing physical-target guards.
Inspection needs no installed shell: no process or LLM is launched. Both
process tools use their real argument normalizer and pre-execution guard.
"""
from ouroboros.tools.registry_guard_process import _run_shell_safety_check
from ouroboros.tools.shell_guards import (
process_shell_guard_args, shell_writer_targets_protected, writer_target_rows,
)
from ouroboros.tools.write_shape import interpreter_write_shape
reg = _registry(tmp_path)
repo = pathlib.Path(reg._ctx.repo_dir)
(repo / "BIBLE.md").write_text("Constitution fixture", encoding="utf-8")
user_root = tmp_path / "user-home"
deliverables = user_root / "Deliverables"
deliverables.mkdir(parents=True)
monkeypatch.setenv("OUROBOROS_USER_FILES_ROOT", str(user_root))
monkeypatch.setenv("OUROBOROS_DELIVERABLES_ROOT", str(deliverables))
def guard(body, cwd="system_repo"):
field = "check" if tool_name == "verify_and_record" else "cmd"
args = process_shell_guard_args(tool_name, {field: [head, "-c", body], "cwd": cwd})
return _run_shell_safety_check(reg, args, "light")
assert guard("cat BIBLE.md") is None
assert shell_writer_targets_protected([head, "-c", "cat BIBLE.md"]) is False
assert interpreter_write_shape([head, "-c", "python3 -c 'print(2 > 1)'"]) is False
assert interpreter_write_shape([head, "-c", "printf result > report.txt"]) is True
# Root actors retain user_files output authority even with an external
# project attached; the selected physical cwd decides the write target.
project = tmp_path / "project"
project.mkdir()
reg._ctx.workspace_root = project
reg._ctx.workspace_mode = "external"
assert guard("printf result > report.txt", "user_files") is None
assert guard("cat BIBLE.md") is None
reg._ctx.workspace_root = None
reg._ctx.workspace_mode = ""
for body in ("rm ordinary.py", "rm BIBLE.md", "rm ../drive/state/state.json"):
refusal = guard(body)
assert refusal is not None, body
assert (refusal.status, refusal.code) == ("blocked", "LIGHT_MODE_BLOCKED"), refusal
assert shell_writer_targets_protected([head, "-c", "rm BIBLE.md"]) is True
rows = writer_target_rows([head, "-c", "printf result > report.txt"])
assert [target for _argv, targets, _inline, _unknown in rows for target in targets] == ["report.txt"]
assert not (deliverables / "report.txt").exists() # inspection never executes
assert (repo / "BIBLE.md").read_text(encoding="utf-8") == "Constitution fixture"
@pytest.mark.parametrize("head", ["sh", "bash", "zsh", "dash", "ash"])
@pytest.mark.parametrize("body,returncode,masked", [
("printf verified", 0, False),
("printf verified | tail -1", 0, True),
("printf verified; exit 7", 7, False),
])
def test_posix_wrapper_verification_preserves_exit_and_masking_receipt(
tmp_path, monkeypatch, head, body, returncode, masked,
):
"""Real verification/receipt consumers; only the OS process is substituted.
Shell availability must not erase dash/ash coverage on macOS or Windows.
The supplied process exit remains verdict authority, masking stays advisory.
"""
from ouroboros.outcomes import read_verification_receipts
from ouroboros.tools.verify import _verify_and_record
reg = _registry(tmp_path)
argv = [head, "-c", body]
calls = []
def run(command, **kwargs):
calls.append((command, kwargs["cwd"]))
return CompletedProcess(command, returncode, "verified", "")
monkeypatch.setattr("ouroboros.tools.shell._tracked_subprocess_run", run)
result = _verify_and_record(
reg._ctx, contract_kind="explicit_command", check=argv, expected="verified",
)
assert calls == [(argv, str(pathlib.Path(reg._ctx.repo_dir).resolve()))]
receipt = read_verification_receipts(reg._ctx.drive_root, "t1")[-1]
assert receipt["status"] == ("pass" if returncode == 0 else "fail"), result
assert receipt["returncode"] == returncode
assert shlex.split(receipt["check"]) == argv
assert receipt.get("check_exit_masking", False) is masked
assert receipt.get("check_exit_masking_reasons", []) == (["pipeline_tail"] if masked else [])

View file

@ -233,7 +233,7 @@ def test_widgets_keep_iframe_sandbox_locked_down():
assert "const csp = moduleFrameCsp(tab.skill);" in module assert "const csp = moduleFrameCsp(tab.skill);" in module
assert "connect-src" not in source assert "connect-src" not in source
assert "'unsafe-eval'" not in source assert "'unsafe-eval'" not in source
assert "window.OuroborosWidget = { fetch: request, onEvent, download };" in source assert "window.OuroborosWidget = { fetch: request, onEvent, download, openExternal: (url) => openExternal(url) };" in source
assert "module widget fetch outside extension route prefix" in source assert "module widget fetch outside extension route prefix" in source

View file

@ -619,3 +619,97 @@ def test_start_service_local_branch_uses_case_aware_overlay(tmp_path, monkeypatc
assert env["PATH"] == "/bundle/bin:C:/stale" assert env["PATH"] == "/bundle/bin:C:/stale"
assert "Path" not in env assert "Path" not in env
assert payload.get("state") in ("running", "ready", "started", None) or payload assert payload.get("state") in ("running", "ready", "started", None) or payload
@pytest.mark.parametrize("layout", ["repo", "nested", "worktree", "plain"])
def test_live_run_scripts_keep_git_clean_and_own_only_their_scratch(tmp_path, monkeypatch, layout):
import concurrent.futures
import pathlib
import subprocess
import time
from ouroboros.tools import shell
monkeypatch.setenv("OUROBOROS_RUNTIME_MODE", "advanced")
system_repo = tmp_path / "system"
workspace = tmp_path / "workspace"
_init_repo(system_repo)
if layout == "worktree":
subprocess.run(["git", "worktree", "add", "--detach", str(workspace)], cwd=system_repo,
check=True, capture_output=True)
elif layout == "plain":
workspace.mkdir()
else:
_init_repo(workspace)
cwd = workspace / "nested" if layout == "nested" else workspace
cwd.mkdir(exist_ok=True)
data = tmp_path / "data"
data.mkdir()
releases = [data / f"release-{i}" for i in range(2)]
ready = [data / f"ready-{i}" for i in range(2)]
ctxs = [ToolContext(repo_dir=system_repo, drive_root=data, workspace_root=workspace,
workspace_mode="external", task_id=f"script-{i}") for i in range(2)]
scripts = [
"import pathlib, time, sys\n"
f"pathlib.Path({str(ready[i])!r}).write_text(sys.argv[0])\n"
f"while not pathlib.Path({str(releases[i])!r}).exists(): time.sleep(0.01)\n"
"print('finished', pathlib.Path.cwd(), sys.argv[1])\n"
for i in range(2)
]
with concurrent.futures.ThreadPoolExecutor(max_workers=2) as pool:
pending = [pool.submit(shell._run_script, ctxs[i], scripts[i], interpreter=sys.executable,
args=[f"argument-{i}"], cwd=str(cwd), timeout_sec=20) for i in range(2)]
try:
until = time.monotonic() + 15
while not all(path.exists() for path in ready) and time.monotonic() < until:
if any(f.done() for f in pending):
pytest.fail(str([f.result() for f in pending if f.done()]))
time.sleep(0.01)
assert all(path.exists() for path in ready), "scripts did not reach READY"
paths = [pathlib.Path(path.read_text()) for path in ready]
assert len({path.parent for path in paths}) == 2
assert all(path.is_file() for path in paths)
assert all((path.parent / ".gitignore").read_text() == "*\n" for path in paths)
if layout != "plain":
status = subprocess.run(["git", "status", "--porcelain", "--untracked-files=all"],
cwd=workspace, check=True, capture_output=True, text=True)
assert status.stdout == ""
neighbour = cwd / ".ouroboros" / "tmp_scripts" / "user.txt"
neighbour.write_text("user work")
if layout != "plain":
status = subprocess.run(["git", "status", "--porcelain", "--untracked-files=all"],
cwd=workspace, check=True, capture_output=True, text=True)
assert "user.txt" in status.stdout
assert "script_" not in status.stdout
releases[0].touch()
assert "argument-0" in pending[0].result(timeout=15)
assert not paths[0].parent.exists()
assert paths[1].is_file()
finally:
for release in releases:
release.touch()
results = [future.result(timeout=15) for future in pending]
assert all("exit_code=0" in result for result in results)
assert neighbour.read_text() == "user work"
assert not any(path.parent.exists() for path in paths)
if layout == "plain":
assert not (workspace / ".git").exists()
def test_run_script_mapping_refusal_cleans_its_own_directory(tmp_path, monkeypatch):
from ouroboros.tools import shell
from types import SimpleNamespace
workspace = tmp_path / "workspace"
workspace.mkdir()
data = tmp_path / "data"
data.mkdir()
ctx = ToolContext(repo_dir=workspace, drive_root=data, workspace_root=workspace,
workspace_mode="external", task_id="map-refusal")
monkeypatch.setattr(shell, "_executor_can_run_cwd", lambda *a: True)
monkeypatch.setattr(shell, "executor_ref_from_ctx", lambda *a: SimpleNamespace(kind="docker_exec"))
def refuse(*args):
raise ValueError("no mapping")
monkeypatch.setattr(shell, "executor_map_host_path", refuse)
result = shell._run_script(ctx, "print('unused')", cwd=str(workspace))
assert "could not map temp script path" in result
assert not (workspace / ".ouroboros").exists()

View file

@ -108,6 +108,11 @@ def test_docker_executor_run_script_uses_backend_script_path(tmp_path, monkeypat
captured: dict[str, object] = {} captured: dict[str, object] = {}
def fake_execute(ctx, cmd, cwd, timeout_sec, env_overlay=None): def fake_execute(ctx, cmd, cwd, timeout_sec, env_overlay=None):
backend = str(cmd[1])
host = workspace / backend.removeprefix("/workspace/")
assert host.name == "script.py"
assert host.read_text() == "print('ok')"
assert (host.parent / ".gitignore").read_text() == "*\n"
captured["cmd"] = list(cmd) captured["cmd"] = list(cmd)
captured["cwd"] = str(cwd) captured["cwd"] = str(cwd)
captured["env_overlay"] = env_overlay captured["env_overlay"] = env_overlay

View file

@ -1,6 +1,7 @@
import { apiFetch } from './api_client.js'; import { apiFetch } from './api_client.js';
import { PAGE_ICONS } from './page_icons.js'; import { PAGE_ICONS } from './page_icons.js';
import { escapeHtmlAttr as escapeHtml, normalizeTone } from './ui_primitives.js'; import { escapeHtmlAttr as escapeHtml, normalizeTone } from './ui_primitives.js';
import { safeExternalUrl } from './utils.js';
export { renderSafeField, collectSafeFieldValues, normalizeTone, setInlineStatus } from './ui_primitives.js'; export { renderSafeField, collectSafeFieldValues, normalizeTone, setInlineStatus } from './ui_primitives.js';
// Cycle note: toast.js imports normalizeTone from this module. Both edges only // Cycle note: toast.js imports normalizeTone from this module. Both edges only
// call the imported function inside function bodies (never at module eval), so // call the imported function inside function bodies (never at module eval), so
@ -377,6 +378,29 @@ async function copyShellLinkWithToast(url, win, doc, toast) {
toast('Link copied — open it in your browser.', 'info'); toast('Link copied — open it in your browser.', 'info');
} }
/** Open an external link through the active host while the user gesture is live. */
export async function openExternalViaHostBridge(url, {
win = window, doc = document, toast = showToast, api = shellBridgeApi(win),
} = {}) {
const target = safeExternalUrl(url);
if (target === '#') throw new Error('Unsupported external link');
if (api) {
const result = api.open_external_url ? await api.open_external_url(target) : null;
if (result?.ok) return { ...result, native: true };
await copyShellLinkWithToast(target, win, doc, toast);
return { ok: false, native: true, degraded: 'copy-link' };
}
const telegram = win.Telegram?.WebApp;
const telegramHost = doc.documentElement?.dataset?.ouroborosHost === 'telegram' || telegram;
if (telegramHost && /^https?:/i.test(target)) {
if (typeof telegram?.openLink !== 'function') throw new Error('Telegram link opener is not ready; try the link again');
telegram.openLink(target);
return { ok: true, native: false, host: 'telegram' };
}
win.open(target, '_blank', 'noopener');
return { ok: true, native: false, host: 'browser' };
}
async function routeShellUrl(kind, url, deps) { async function routeShellUrl(kind, url, deps) {
const { api, win, doc, toast, openFile, downloadFile, filename = '', wantsDownload = false } = deps; const { api, win, doc, toast, openFile, downloadFile, filename = '', wantsDownload = false } = deps;
try { try {
@ -393,12 +417,7 @@ async function routeShellUrl(kind, url, deps) {
if (wantsDownload) await downloadFile(url, name); if (wantsDownload) await downloadFile(url, name);
else await openFile(url, name); else await openFile(url, name);
} else if (kind === 'external') { } else if (kind === 'external') {
// Version-skew fallback (no open_external_url on an old packaged await openExternalViaHostBridge(url, { api, win, doc, toast });
// launcher) and an honest bridge failure ({ok:false}: no browser
// could be launched) degrade the same way: hand the owner the link
// instead of leaving a silently dead control.
const result = api?.open_external_url ? await api.open_external_url(url) : null;
if (!result?.ok) await copyShellLinkWithToast(url, win, doc, toast);
} else if (kind === 'bytes') { } else if (kind === 'bytes') {
const result = await downloadBlobViaHostBridge(url, filename, { win, doc }); const result = await downloadBlobViaHostBridge(url, filename, { win, doc });
if (result.unavailable) { toast(result.error, 'warn'); return; } if (result.unavailable) { toast(result.error, 'warn'); return; }

View file

@ -1,5 +1,7 @@
/* Framed widget bootstrap scripts. The parent remains the route and lifecycle owner. */ /* Framed widget bootstrap scripts. The parent remains the route and lifecycle owner. */
import { safeExternalUrl } from './utils.js';
// The parent hands every streamed body chunk to the frame as a transferred // The parent hands every streamed body chunk to the frame as a transferred
// ArrayBuffer. A reader's Uint8Array may be a window onto a larger buffer, so // ArrayBuffer. A reader's Uint8Array may be a window onto a larger buffer, so
// transfer exactly the bytes the view covers and nothing beside them. // transfer exactly the bytes the view covers and nothing beside them.
@ -12,9 +14,11 @@ export function bridgeChunkBuffer(view) {
// Child side of the one bridge grammar (nonce-bound, parent ⇄ frame): // Child side of the one bridge grammar (nonce-bound, parent ⇄ frame):
// child → parent ouro-widget-fetch {id, url, init} · ouro-widget-fetch-abort {id} // child → parent ouro-widget-fetch {id, url, init} · ouro-widget-fetch-abort {id}
// ouro-widget-fetch-pull {id} · ouro-widget-download {id, name, source} // ouro-widget-fetch-pull {id} · ouro-widget-download {id, name, source}
// ouro-widget-open-external {id, url}
// ouro-widget-events {op: subscribe | unsubscribe} · ouro-widget-disposed // ouro-widget-events {op: subscribe | unsubscribe} · ouro-widget-disposed
// ouro-widget-error {kind: error | rejection | csp, message, source, line} // ouro-widget-error {kind: error | rejection | csp, message, source, line}
// parent → child ouro-widget-fetch-chunk {id, phase: headers | data | end | error, …} // parent → child ouro-widget-fetch-chunk {id, phase: headers | data | end | error, …}
// ouro-widget-open-external-result {id, result}
// ouro-widget-event {event, data} · ouro-widget-dispose // ouro-widget-event {event, data} · ouro-widget-dispose
// Every bridged fetch streams: the child rebuilds a real Response over a // Every bridged fetch streams: the child rebuilds a real Response over a
// ReadableStream fed by `data` frames (binary by default), so text/json/blob // ReadableStream fed by `data` frames (binary by default), so text/json/blob
@ -25,6 +29,7 @@ export function moduleBridgeScript(nonce, routeBase = '') {
(() => { (() => {
const nonce = ${JSON.stringify(nonce)}; const nonce = ${JSON.stringify(nonce)};
const routeBase = ${JSON.stringify(routeBase)}; const routeBase = ${JSON.stringify(routeBase)};
const safeExternalUrl = (${safeExternalUrl.toString()});
let seq = 0; let seq = 0;
let disposing = false; let disposing = false;
let disposed = false; let disposed = false;
@ -32,6 +37,8 @@ export function moduleBridgeScript(nonce, routeBase = '') {
// frame, then feeds, ends or errors that Response's body stream. // frame, then feeds, ends or errors that Response's body stream.
const pending = new Map(); const pending = new Map();
const downloads = new Map(); const downloads = new Map();
const externalLinks = new Map();
const originalOpen = window.open;
const cleanup = new Set(); const cleanup = new Set();
const eventListeners = new Set(); const eventListeners = new Set();
const post = (message) => window.parent.postMessage({ ...message, nonce }, '*'); const post = (message) => window.parent.postMessage({ ...message, nonce }, '*');
@ -53,6 +60,8 @@ export function moduleBridgeScript(nonce, routeBase = '') {
cleanup.clear(); cleanup.clear();
await Promise.allSettled(hooks.map((fn) => Promise.resolve().then(fn))); await Promise.allSettled(hooks.map((fn) => Promise.resolve().then(fn)));
window.document?.removeEventListener('click', clickDownload); window.document?.removeEventListener('click', clickDownload);
window.document?.removeEventListener('click', clickExternal);
window.open = originalOpen;
blobUrls.clear(); blobUrls.clear();
if (createUrl) urlApi.createObjectURL = createUrl; if (createUrl) urlApi.createObjectURL = createUrl;
if (revokeUrl) urlApi.revokeObjectURL = revokeUrl; if (revokeUrl) urlApi.revokeObjectURL = revokeUrl;
@ -62,6 +71,8 @@ export function moduleBridgeScript(nonce, routeBase = '') {
pending.clear(); pending.clear();
downloads.forEach(({ reject }) => reject(new Error('widget disposed'))); downloads.forEach(({ reject }) => reject(new Error('widget disposed')));
downloads.clear(); downloads.clear();
externalLinks.forEach(({ reject }) => reject(new Error('widget disposed')));
externalLinks.clear();
eventListeners.clear(); eventListeners.clear();
window.removeEventListener('message', onMessage); window.removeEventListener('message', onMessage);
window.removeEventListener('error', onError); window.removeEventListener('error', onError);
@ -85,6 +96,14 @@ export function moduleBridgeScript(nonce, routeBase = '') {
}); });
return; return;
} }
if (msg.type === 'ouro-widget-open-external-result') {
const item = externalLinks.get(msg.id);
if (!item) return;
externalLinks.delete(msg.id);
if (msg.result?.ok || msg.result?.degraded) item.resolve(msg.result);
else item.reject(new Error(msg.result?.error || 'widget link failed'));
return;
}
if (msg.type === 'ouro-widget-download-result') { if (msg.type === 'ouro-widget-download-result') {
const item = downloads.get(msg.id); const item = downloads.get(msg.id);
if (!item) return; if (!item) return;
@ -263,8 +282,36 @@ export function moduleBridgeScript(nonce, routeBase = '') {
event.preventDefault(); event.preventDefault();
download(anchor.download, source).catch((error) => fault('error', error.message, '', 0)); download(anchor.download, source).catch((error) => fault('error', error.message, '', 0));
}; };
const openExternal = (url, trustedAnchor = false) => new Promise((resolve, reject) => {
if (disposing || disposed) { reject(new Error('widget disposed')); return; }
const activation = window.navigator?.userActivation;
if (activation ? !activation.isActive : !trustedAnchor) {
reject(new Error('Opening a link requires a user action'));
return;
}
const target = safeExternalUrl(url);
if (target === '#') { reject(new Error('Unsupported external link')); return; }
const id = ++seq;
externalLinks.set(id, { resolve, reject });
try { post({ type: 'ouro-widget-open-external', id, url: target }); }
catch (error) { externalLinks.delete(id); reject(error); }
});
const clickExternal = (event) => {
if (event.defaultPrevented || event.button > 0 || !event.isTrusted) return;
const anchor = event.target?.closest?.('a[href]');
if (!anchor || anchor.hasAttribute('download')) return;
const target = safeExternalUrl(anchor.getAttribute('href'));
if (target === '#' || (routeBase && target.startsWith(routeBase))) return;
event.preventDefault();
openExternal(target, true).catch((error) => fault('error', error.message, '', 0));
};
window.open = (url) => {
openExternal(url).catch((error) => fault('error', error.message, '', 0));
return null;
};
window.document?.addEventListener('click', clickExternal);
window.document?.addEventListener('click', clickDownload); window.document?.addEventListener('click', clickDownload);
window.OuroborosWidget = { fetch: request, onEvent, download }; window.OuroborosWidget = { fetch: request, onEvent, download, openExternal: (url) => openExternal(url) };
})(); })();
`; `;
} }

View file

@ -10,7 +10,7 @@ import { escapeHtmlAttr as escapeHtml } from './utils.js';
import { bridgeChunkBuffer, moduleBridgeScript, moduleResizeScript } from './widget_frame.js'; import { bridgeChunkBuffer, moduleBridgeScript, moduleResizeScript } from './widget_frame.js';
import { boundedNumber, WIDGET_DISPOSE_ACK_TIMEOUT_MS, WIDGET_REQUEST_TIMEOUT_MS } from './widget_job.js'; import { boundedNumber, WIDGET_DISPOSE_ACK_TIMEOUT_MS, WIDGET_REQUEST_TIMEOUT_MS } from './widget_job.js';
import { setWidgetCardFault } from './widget_card.js'; import { setWidgetCardFault } from './widget_card.js';
import { downloadViaHostBridge, downloadBlobViaHostBridge } from './ui_helpers.js'; import { downloadViaHostBridge, downloadBlobViaHostBridge, openExternalViaHostBridge } from './ui_helpers.js';
export const WIDGET_FRAME_DEFAULT_HEIGHT = 320; export const WIDGET_FRAME_DEFAULT_HEIGHT = 320;
export const WIDGET_FRAME_MAX_HEIGHT = 8192; export const WIDGET_FRAME_MAX_HEIGHT = 8192;
@ -258,6 +258,20 @@ export async function mountModuleWidget(mount, tab, render, mountSignal = null,
} }
post({ type: 'ouro-widget-download-result', id: msg.id, result }); post({ type: 'ouro-widget-download-result', id: msg.id, result });
}; };
const relayExternal = async (msg) => {
let result;
try {
if (disposing || !iframe.isConnected) throw new Error('widget disposed');
if (window.navigator?.userActivation && !window.navigator.userActivation.isActive) {
throw new Error('Opening a link requires a user action');
}
// Invoked synchronously from onMessage, before any await/fetch.
result = await openExternalViaHostBridge(msg.url);
} catch (error) {
result = { ok: false, error: error?.message || String(error) };
}
post({ type: 'ouro-widget-open-external-result', id: msg.id, result });
};
const onMessage = (event) => { const onMessage = (event) => {
if (disposed || !iframe || event.source !== iframe.contentWindow) return; if (disposed || !iframe || event.source !== iframe.contentWindow) return;
const msg = event.data || {}; const msg = event.data || {};
@ -292,6 +306,10 @@ export async function mountModuleWidget(mount, tab, render, mountSignal = null,
else if (msg.op === 'unsubscribe') messageHandlers?.delete(onWsMessage); else if (msg.op === 'unsubscribe') messageHandlers?.delete(onWsMessage);
return; return;
} }
if (msg.type === 'ouro-widget-open-external') {
relayExternal(msg);
return;
}
if (msg.type === 'ouro-widget-download') { if (msg.type === 'ouro-widget-download') {
relayDownload(msg); relayDownload(msg);
return; return;

View file

@ -9,6 +9,7 @@ import {
downloadBlobViaHostBridge, downloadBlobViaHostBridge,
installDesktopShellLinkInterceptor, installDesktopShellLinkInterceptor,
openViaHostBridge, openViaHostBridge,
openExternalViaHostBridge,
} from '../modules/ui_helpers.js'; } from '../modules/ui_helpers.js';
const BASE = 'http://127.0.0.1:8765/'; const BASE = 'http://127.0.0.1:8765/';
@ -584,3 +585,30 @@ test('direct old-shell download propagates an unavailable copy fallback', async
await assert.rejects(() => downloadViaHostBridge('/api/tasks/task/artifacts/data.bin', 'data.bin', { streaming: true }), /Desktop file download is unavailable/); await assert.rejects(() => downloadViaHostBridge('/api/tasks/task/artifacts/data.bin', 'data.bin', { streaming: true }), /Desktop file download is unavailable/);
} finally { Object.assign(globalThis, prior); } } finally { Object.assign(globalThis, prior); }
}); });
test('external host handoff preserves browser activation and accepts noopener null', async () => {
const calls = [];
const win = { open(...args) { calls.push(args); return null; } };
const result = openExternalViaHostBridge('https://example.com', { win, doc: {} });
assert.deepEqual(calls, [['https://example.com/', '_blank', 'noopener']], 'physical handoff occurs before yielding');
assert.equal((await result).ok, true);
for (const url of ['javascript:alert(1)', 'file:///tmp/x', '/relative', 'blob:https://example.com/x']) {
await assert.rejects(openExternalViaHostBridge(url, { win, doc: {} }), /Unsupported/);
}
assert.equal(calls.length, 1);
});
test('Telegram host uses the ready SDK and never queues an unavailable click', async () => {
const calls = [];
const win = { open(...args) { calls.push(['browser', ...args]); } };
const doc = { documentElement: { dataset: { ouroborosHost: 'telegram' } } };
await assert.rejects(openExternalViaHostBridge('https://example.com', { win, doc }), /not ready/);
win.Telegram = { WebApp: { openLink(url) { calls.push(['telegram', url]); } } };
assert.deepEqual(calls, [], 'SDK arrival does not replay a prior click');
const result = openExternalViaHostBridge('https://example.com', { win, doc });
assert.deepEqual(calls, [['telegram', 'https://example.com/']]);
assert.equal((await result).host, 'telegram');
await openExternalViaHostBridge('mailto:owner@example.com', { win, doc });
assert.deepEqual(calls.at(-1), ['browser', 'mailto:owner@example.com', '_blank', 'noopener']);
});

View file

@ -5,12 +5,20 @@ import { bridgeChunkBuffer, moduleBridgeScript, moduleResizeScript } from '../mo
// Runs the child bootstrap against a fake `window`; `deliver` plays a // Runs the child bootstrap against a fake `window`; `deliver` plays a
// parent → child message, `posted` records child → parent messages. // parent → child message, `posted` records child → parent messages.
function bridgeHarness() { function bridgeHarness({ active = false, activationApi = true } = {}) {
const posted = []; const posted = [];
const parent = { postMessage(message) { posted.push(message); } }; const parent = { postMessage(message) { posted.push(message); } };
const listeners = new Map(); const listeners = new Map();
const docListeners = new Map();
const originalOpen = () => null;
const window = { const window = {
parent, parent,
open: originalOpen,
navigator: activationApi ? { userActivation: { isActive: active } } : {},
document: {
addEventListener(type, fn) { const list = docListeners.get(type) || []; list.push(fn); docListeners.set(type, list); },
removeEventListener(type, fn) { docListeners.set(type, (docListeners.get(type) || []).filter((item) => item !== fn)); },
},
addEventListener(type, listener) { listeners.set(type, listener); }, addEventListener(type, listener) { listeners.set(type, listener); },
removeEventListener(type, listener) { removeEventListener(type, listener) {
if (listeners.get(type) === listener) listeners.delete(type); if (listeners.get(type) === listener) listeners.delete(type);
@ -20,7 +28,7 @@ function bridgeHarness() {
const deliver = (data, source = parent) => listeners.get('message')?.({ source, data: { nonce: 'nonce-1', ...data } }); const deliver = (data, source = parent) => listeners.get('message')?.({ source, data: { nonce: 'nonce-1', ...data } });
const chunk = (id, phase, extra = {}) => deliver({ type: 'ouro-widget-fetch-chunk', id, phase, ...extra }); const chunk = (id, phase, extra = {}) => deliver({ type: 'ouro-widget-fetch-chunk', id, phase, ...extra });
const flush = () => new Promise((resolve) => setTimeout(resolve, 0)); const flush = () => new Promise((resolve) => setTimeout(resolve, 0));
return { window, posted, listeners, deliver, chunk, flush }; return { window, posted, listeners, docListeners, originalOpen, deliver, chunk, flush };
} }
const bytes = (...values) => new Uint8Array(values).buffer; const bytes = (...values) => new Uint8Array(values).buffer;
@ -285,3 +293,41 @@ test('download uses the nonce bridge and settles from the actual host outcome',
deliver({ type: 'ouro-widget-download-result', id: 2, result: { ok: false, error: 'disk full' } }); deliver({ type: 'ouro-widget-download-result', id: 2, result: { ok: false, error: 'disk full' } });
await assert.rejects(failure, /disk full/); await assert.rejects(failure, /disk full/);
}); });
test('external-link request requires activation, checks parent and settles during disposal', async () => {
const h = bridgeHarness({ active: true });
const result = h.window.OuroborosWidget.openExternal('https://example.com');
assert.deepEqual(h.posted.at(-1), { type: 'ouro-widget-open-external', nonce: 'nonce-1', id: 1, url: 'https://example.com/' });
h.deliver({ type: 'ouro-widget-open-external-result', id: 1, result: { ok: true } }, {});
h.deliver({ type: 'ouro-widget-open-external-result', id: 1, nonce: 'wrong', result: { ok: true } });
h.deliver({ type: 'ouro-widget-open-external-result', id: 1, result: { ok: false, error: 'host unavailable' } });
await assert.rejects(result, /host unavailable/);
h.window.navigator.userActivation.isActive = false;
await assert.rejects(h.window.OuroborosWidget.openExternal('https://example.com'), /user action/);
h.window.navigator.userActivation.isActive = true;
await assert.rejects(h.window.OuroborosWidget.openExternal('/relative'), /Unsupported/);
const pending = h.window.OuroborosWidget.openExternal('https://example.com/again');
const rejected = assert.rejects(pending, /disposed/);
h.deliver({ type: 'ouro-widget-dispose' });
await rejected;
assert.equal(h.window.open, h.originalOpen);
assert.equal(h.docListeners.get('click').length, 0);
});
test('trusted anchors work on old hosts; synthetic, handled, download and internal links are untouched', async () => {
const h = bridgeHarness({ activationApi: false });
await assert.rejects(h.window.OuroborosWidget.openExternal('https://example.com'), /user action/);
const click = (href, changes = {}) => {
const anchor = { getAttribute() { return href; }, hasAttribute(name) { return name === 'download' && changes.download; } };
const event = { isTrusted: true, button: 0, defaultPrevented: false, target: { closest(selector) { return selector === 'a[href]' ? anchor : null; } },
preventDefault() { this.defaultPrevented = true; }, ...changes };
for (const fn of h.docListeners.get('click')) fn(event);
return event;
};
for (const [href, changes] of [['#x', {}], ['/relative', {}], ['https://example.com', { isTrusted: false }], ['https://example.com', { defaultPrevented: true }], ['https://example.com', { download: true }]]) click(href, changes);
assert.equal(h.posted.length, 0);
assert.equal(click('https://example.com').defaultPrevented, true);
assert.equal(h.posted.length, 1);
h.deliver({ type: 'ouro-widget-open-external-result', id: 1, result: { ok: true } });
});

View file

@ -0,0 +1,120 @@
import test from 'node:test';
import assert from 'node:assert/strict';
import { mountModuleWidget } from '../modules/widget_module.js';
// Exercise the production parent mount and host opener. This DOM double does
// not establish browser user activation; click/Enter/touch need browser proof.
async function relayHarness(t, api = null) {
const opened = [];
const replies = [];
const listeners = new Map();
const attributes = new Map();
const iframe = {
isConnected: true,
dataset: {},
style: { setProperty() {} },
setAttribute(name, value) { attributes.set(name, value); },
contentWindow: { postMessage(message) { replies.push(message); } },
remove() { this.isConnected = false; this.parentNode = null; },
};
const mount = {
replaceChildren(node) { node.parentNode = this; },
};
const win = {
location: { origin: 'https://ouroboros.test' },
navigator: { userActivation: { isActive: true } },
open(...args) { opened.push(args); return null; },
addEventListener(type, listener) { listeners.set(type, listener); },
removeEventListener(type, listener) {
if (listeners.get(type) === listener) listeners.delete(type);
},
...(api ? { pywebview: { api } } : {}),
};
t.mock.method(globalThis, 'fetch', async () => new Response('// reviewed module'));
const globals = new Map();
let dispose;
let send;
t.after(async () => {
try {
if (dispose) {
const stopped = dispose();
send({ type: 'ouro-widget-disposed' });
await stopped;
}
} finally {
for (const [name, previous] of globals) {
if (previous) Object.defineProperty(globalThis, name, previous);
else delete globalThis[name];
}
}
});
for (const [name, value] of Object.entries({
window: win,
document: { createElement() { return iframe; }, documentElement: { dataset: {} } },
})) {
globals.set(name, Object.getOwnPropertyDescriptor(globalThis, name));
Object.defineProperty(globalThis, name, { configurable: true, value });
}
dispose = await mountModuleWidget(mount, { skill: 'links' }, { entry: 'widget.js', height: 320 });
const nonce = JSON.parse(iframe.srcdoc.match(/const nonce = ("[^"]+");/)[1]);
const onMessage = listeners.get('message');
send = (data = {}, source = iframe.contentWindow) => onMessage({
source, data: { type: 'ouro-widget-open-external', id: 1, nonce, url: 'https://example.test/target', ...data },
});
const flush = () => new Promise((resolve) => setImmediate(resolve));
return { win, iframe, attributes, opened, replies, listeners, send, dispose, flush };
}
test('module parent opens synchronously, acknowledges noopener null and keeps the sandbox', async (t) => {
const h = await relayHarness(t);
h.send();
assert.deepEqual(h.opened, [['https://example.test/target', '_blank', 'noopener']]);
await h.flush();
assert.deepEqual(h.replies.at(-1).result, { ok: true, native: false, host: 'browser' });
assert.equal(h.attributes.get('sandbox'), 'allow-scripts allow-pointer-lock allow-downloads');
assert.equal(h.attributes.get('allow'), 'autoplay; fullscreen; clipboard-write');
});
test('module parent rejects foreign frames, stale nonces, inactive requests and unsafe URLs', async (t) => {
const h = await relayHarness(t);
h.send({}, {});
h.send({ nonce: 'stale-nonce' });
assert.deepEqual(h.replies, []);
h.win.navigator.userActivation.isActive = false;
h.send();
await h.flush();
assert.match(h.replies.at(-1).result.error, /requires a user action/);
h.win.navigator.userActivation.isActive = true;
h.send({ url: 'javascript:alert(1)' });
await h.flush();
assert.match(h.replies.at(-1).result.error, /Unsupported external link/);
h.iframe.isConnected = false;
h.send();
await h.flush();
assert.match(h.replies.at(-1).result.error, /disposed/);
assert.deepEqual(h.opened, []);
});
test('module parent uses the native bridge and detaches the relay after disposal', async (t) => {
const native = [];
const h = await relayHarness(t, { open_external_url(url) { native.push(url); return { ok: true }; } });
h.send();
assert.deepEqual(native, ['https://example.test/target']);
await h.flush();
assert.deepEqual(h.replies.at(-1).result, { ok: true, native: true });
assert.deepEqual(h.opened, []);
const stopped = h.dispose();
h.send();
await h.flush();
assert.match(h.replies.at(-1).result.error, /disposed/);
h.send({ type: 'ouro-widget-disposed' });
await stopped;
const replyCount = h.replies.length;
h.send();
await h.flush();
assert.equal(h.replies.length, replyCount);
assert.equal(h.listeners.has('message'), false);
assert.equal(h.iframe.isConnected, false);
assert.equal(native.length, 1);
});