mirror of
https://github.com/razzant/ouroboros.git
synced 2026-10-02 19:58:46 +00:00
Dispositions of the multi-model triad + scope wave (grok-4.6, fable-5,
gpt-5.6-sol; scope sol-xhigh), severity order:
- _evaluate_bounded lost Playwright's function-invocation semantics
(unanimous MAJOR, confirmed against the installed driver: UtilityScript
eval()s a raw-string expression and INVOKES a function-valued result
when isFunction is unset). page.evaluate(_MARKDOWN_JS), the health
snapshot and every '() =>' user evaluate serialized the uninvoked
function to undefined. The wrapper now evals and invokes exactly as the
driver does; the wrapper-shape test pins both halves.
- The settled-refresh cursor pinned its byte offset on the earliest
SETTLED row of a still-RUNNING task (grok/sol MAJOR): each tick re-read
the same 5MB window behind one long-lived parent and permanently starved
every later terminal-boundary settlement past the cap — the exact class
the cursor exists to heal. The offset now always advances; deferred
tasks ride a durable bounded map retried every tick (oldest-evicted at
500, disclosed), pinned by a starvation test.
- A superseded-by-revision acceptance run was excluded from the prior-run
lookup even on an identical paid identity (sol MAJOR): the resubmission
fell through to the wallet's binding_dispatch_already_claimed and
recorded a synthetic DEGRADED panel instead of the typed
identical_acceptance_refused terminal. Identity-matched superseded runs
now replay for free with a replayed_from_superseded disclosure —
evidence revision stays stale-detection, never a paid-cycle mint
(owner decision 5=A). Moved _prior_acceptance_run to
acceptance_dialogue.py (loop.py byte debt is shrink-only).
- The trailing-JSON scanner missed a real directive after an unmatched
brace/quote in the prose prefix (sol MAJOR with executed repro; fable
MINOR): the forced rail then leaked the raw protocol JSON — the original
O4 defect resurfacing on one input class. On primary-scan failure the
parser now retries from the last few line-start '{' anchors (bounded,
string-aware); the inline-after-unbalanced-prose case stays a disclosed
prose-degrade, pinned. The scope wave's fence-peel O(n*k) shortfall is
also fixed (no slicing per peel, 4-fence cap) and ARCHITECTURE's parser
description synced to the final implementation.
- Legacy/raw-basis density rows stayed authoritative for their 14-day TTL
after an upgrade (sol MAJOR, probe-confirmed): resolve_main_token_density
now calibrates ONLY on bounded_proxy rows (review/aggregate resolvers
keep all rows — their text-heavy witnesses measure the same on either
basis); a poisoned 0.55 image-route witness falls back to the neutral
1.0 instead of under-admitting by 45%. Tests seeding the main resolver
updated to the modern producer shape (the observer stamps the basis).
- acceptance_patch_dispositions computed the unreviewed_delegated_apply
headline AFTER truncation (fable MAJOR): a delegated apply among the
25-cap-omitted oldest rows false-negatived the one fact the attest
decision (4=A) exists to surface. Headline now derives from the complete
row set; pinned with a mixed-row scenario.
- write_patch_verdict treated custody.emit's False return as success
(grok/fable/sol): a lost attestation row was invisible — the packet's
empty section read as 'nothing recorded'. Both the exception AND the
False-return paths now disclose custody_row_write_failed on the verdict
artifact, and the row lands on the canonical custody root.
- find_child_tasks overlay test now pins CONTENT restoration (cost,
result, artifacts), not just the id (grok MINOR).
Also: reservation test keeps the read-surface pin with the offline
rationale (a cost-equality behavioral pin honestly returns None without a
live pricing route — fable m5 disposition); scope-wave (a)-items verified
against the ratified record (tools_on_path=11=A, salvage-rail in the O4
verdict, dialogue-history=R2, cycle accounting=O11 seams) — prompt gaps,
not drift.
Suites: 11775 parallel + 618 serial + 694 web green; ruff -F clean;
size-ratchet regenerated (loop.py net-shrinks further).
Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
124 lines
5.3 KiB
Python
124 lines
5.3 KiB
Python
"""D-trace (owner 4=A): patch apply/reject decisions are attested, not gated.
|
|
|
|
``integrate_delegated_patch`` applies a child's diff on five mechanical
|
|
manifest fields; no review facts exist anywhere on the path (the child is
|
|
forbidden from reviewing its own diff, and no host review of the captured
|
|
bytes runs). Rather than inventing a review, every verdict lands as a typed
|
|
``subagent_patch_verdict`` custody row and the acceptance packet carries the
|
|
host-attested disposition section — first-class, bounded, never squeezed
|
|
through the 4KB artifact-preview cliff. Nothing refuses an apply.
|
|
"""
|
|
|
|
import json
|
|
|
|
from ouroboros import delegate_custody as custody
|
|
from ouroboros.delegate_evidence import acceptance_patch_dispositions
|
|
|
|
|
|
def _emit_verdict(tmp_path, *, task_id="t-parent", child="run_r1",
|
|
pipeline="delegated", disposition="applied", applied=True,
|
|
reason="looks right", sha="abc123", write_failed=False):
|
|
assert custody.emit(tmp_path, "delegate_run_patch_verdict", {
|
|
"run_id": "", "task_id": task_id, "child_task_id": child,
|
|
"pipeline": pipeline, "disposition": disposition, "applied": applied,
|
|
"reason": reason, "patch_sha256": sha,
|
|
"verdict_artifact_write_failed": write_failed,
|
|
})
|
|
|
|
|
|
def test_no_rows_means_no_section_not_clean(tmp_path):
|
|
assert acceptance_patch_dispositions(tmp_path, "t-parent") == {}
|
|
|
|
|
|
def test_dispositions_project_with_the_unreviewed_headline(tmp_path):
|
|
_emit_verdict(tmp_path)
|
|
_emit_verdict(tmp_path, child="t-sub", pipeline="subagent",
|
|
disposition="rejected", applied=False, reason="conflicts")
|
|
section = acceptance_patch_dispositions(tmp_path, "t-parent")
|
|
assert section["total"] == 2
|
|
rows = {r["child"]: r for r in section["rows"]}
|
|
assert rows["run_r1"]["applied"] is True
|
|
assert rows["run_r1"]["pipeline"] == "delegated"
|
|
assert rows["t-sub"]["applied"] is False
|
|
# The honest headline: a delegated patch landed with no host review of
|
|
# its bytes. Visibility only — nothing on the apply path refuses.
|
|
assert section["unreviewed_delegated_apply"] is True
|
|
|
|
|
|
def test_rejected_only_delegated_rows_carry_no_apply_headline(tmp_path):
|
|
_emit_verdict(tmp_path, disposition="rejected", applied=False)
|
|
section = acceptance_patch_dispositions(tmp_path, "t-parent")
|
|
assert "unreviewed_delegated_apply" not in section
|
|
|
|
|
|
def test_other_tasks_rows_stay_out(tmp_path):
|
|
_emit_verdict(tmp_path, task_id="t-other")
|
|
assert acceptance_patch_dispositions(tmp_path, "t-parent") == {}
|
|
|
|
|
|
def test_bounded_with_exact_omitted_count(tmp_path):
|
|
for i in range(25):
|
|
_emit_verdict(tmp_path, child=f"run_r{i}")
|
|
section = acceptance_patch_dispositions(tmp_path, "t-parent")
|
|
assert section["total"] == 25
|
|
assert section["omitted"] == 5
|
|
assert len(section["rows"]) == 20
|
|
# Newest rows survive the bound (the most recent decisions matter most).
|
|
assert section["rows"][-1]["child"] == "run_r24"
|
|
|
|
|
|
def test_headline_survives_when_the_only_delegated_apply_is_truncated(tmp_path):
|
|
"""fable M1: the ONE delegated apply sits among the oldest rows past the
|
|
cap of 20; the headline must be computed over the complete set, or the
|
|
panel loses the exact fact the attest decision exists to surface."""
|
|
_emit_verdict(tmp_path, child="run_early", pipeline="delegated",
|
|
disposition="applied", applied=True)
|
|
for i in range(24):
|
|
_emit_verdict(tmp_path, child=f"t-sub{i}", pipeline="subagent",
|
|
disposition="rejected", applied=False)
|
|
section = acceptance_patch_dispositions(tmp_path, "t-parent")
|
|
assert section["omitted"] == 5
|
|
assert all(r["child"] != "run_early" for r in section["rows"])
|
|
assert section["unreviewed_delegated_apply"] is True
|
|
|
|
|
|
def test_write_verdict_emits_the_custody_row(tmp_path):
|
|
from types import SimpleNamespace
|
|
|
|
from ouroboros.tools.subagent_integration import _write_verdict
|
|
|
|
ctx = SimpleNamespace(drive_root=str(tmp_path), task_id="t-parent",
|
|
task_metadata={}, budget_drive_root=str(tmp_path))
|
|
path = _write_verdict(
|
|
ctx, "run_r9", outcome="applied", reason="ok", files=["a.py"],
|
|
manifest={"sha256": "deadbeef", "diffstat": "1 file"},
|
|
applied=True, conflicts=[], protected=[],
|
|
)
|
|
assert path, "verdict artifact write should succeed in a temp drive"
|
|
rows = [
|
|
row for row in custody._iter_rows(custody.event_log_path(tmp_path))
|
|
if row.get("type") == "delegate_run_patch_verdict"
|
|
]
|
|
assert len(rows) == 1
|
|
row = rows[0]
|
|
assert row["child_task_id"] == "run_r9"
|
|
assert row["pipeline"] == "delegated"
|
|
assert row["patch_sha256"] == "deadbeef"
|
|
assert row["applied"] is True
|
|
assert row["verdict_artifact_write_failed"] is False
|
|
|
|
|
|
def test_verdict_artifact_content_still_written(tmp_path):
|
|
from types import SimpleNamespace
|
|
|
|
from ouroboros.tools.subagent_integration import _write_verdict
|
|
|
|
ctx = SimpleNamespace(drive_root=str(tmp_path), task_id="t-parent",
|
|
task_metadata={}, budget_drive_root=str(tmp_path))
|
|
path = _write_verdict(
|
|
ctx, "child-7", outcome="rejected", reason="conflicts", files=[],
|
|
manifest={}, applied=False, conflicts=["x"], protected=[],
|
|
)
|
|
data = json.loads(open(path).read())
|
|
assert data["outcome"] == "rejected"
|
|
assert data["child_task_id"] == "child-7"
|