mirror of
https://github.com/razzant/ouroboros.git
synced 2026-10-03 04:07:04 +00:00
tz2: B3 follow-up objective_author, C2 stat-only files_rescued fact, B1 comment-only browser test
Timer follow-ups carry the same objective_author stamp a promote carries (B3). The host_task_facts row and the cancel receipt state files_rescued as a stat-only positive/zero/unknown count with hash_computed:false, walking the canonical and child-drive artifact stores (C2 fallback per owner correction; B5 receipt stays NOT_DONE pending the TZ-1 landed API). New real-server Chromium test for a zero-option question card (B1). Chapter budgets bumped with rationale. Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
This commit is contained in:
parent
7a387f7178
commit
001904a6bc
11 changed files with 352 additions and 18 deletions
|
|
@ -535,7 +535,7 @@ The advisory rows a reflection or summary reads are ATTRIBUTED. Advisory runs ar
|
|||
|
||||
Only roots synthesize; `root_phase_checkpoint` makes paid synthesis at-most-once across restart, while children contribute evidence. Durable-result persistence owes the answer as `final:<tid>:<digest>` in `supervisor/terminal_delivery.py`'s bounded outbox (§5; normal/cancel/reap). `send_message` delivers immediately; the retained buffered copy shares its ID for durable dedupe. Replays use bounded backoff. Exhaustion/eviction preserves full text on disk, emits `terminal_delivery_exhausted` and a chat notice; external delivery remains at-least-once. Buffered `task_done` stays last to retain the slot/child drive during synthesis; a hung-synthesis reap need not lose the delivered answer. Project roots keep early answers in Project. Their canonical row and deferred Main mirror use `terminal_projection.settle_terminal_projection` via task-done/checkpoint/startup/maintenance; §3 "Main rows and host-stamped card rows" owns readiness, retirement and limits.
|
||||
|
||||
Synthesis receives a sealed final package from the durable result — the submitted final text, its artifact manifest and completion_observations. Full redacted action observations live in the canonical artifact store (`task.budget_drive_root or drive_root`), in the write-once `source_handles/context_checkpoints` store with verified `task_source` refs, before compact publication and outside deliverables and inferred readiness; their native reader `get_task_result(include_completion_source=true)` returns complete length/hash first, then explicit `source_start_char`/`source_end_char` ranges (`artifacts.text_source_range_projection`, the shared work-order range contract), with bytes, kind, path containment and SHA checked before any excerpt. Packet-only reflection receives per-send-tool counts, each family's latest recorded return, and task-related skill readiness with coverage; full-source references are for later readers, not evidence the synthesizer has read. Positive observed facts correct error-trace impressions, while tool success does not prove owner receipt, empty material does not prove absence, and skill readiness does not attribute an owner's action to the task. Before context cleanup, `agent_task_pipeline.emit_task_results` also freezes `review_evidence.task_inputs` through `post_task_synthesis.capture_task_inputs`: `run_origin`, the existing task-local owner corpus, intact question/answer provenance and the canonical split-root verification-receipt union. Reflection receives the same complete redacted content through `reflection.task_inputs_prompt_section`, separate from bounded trace/review excerpts. A zero return code is positive evidence; an unrelated later pass cannot resolve another check's failure. Peer proposals stay attributed, and unavailable input is not evidence that approval or verification never existed. Recovery uses these stored observations and inputs, not a later conversation. A free `host_task_facts` row precedes paid stages (or follows result persistence when Stop skips them): no model call or narrative; its metrics, routing and cost serve history.
|
||||
Synthesis receives a sealed final package from the durable result — the submitted final text, its artifact manifest and completion_observations. Full redacted action observations live in the canonical artifact store (`task.budget_drive_root or drive_root`), in the write-once `source_handles/context_checkpoints` store with verified `task_source` refs, before compact publication and outside deliverables and inferred readiness; their native reader `get_task_result(include_completion_source=true)` returns complete length/hash first, then explicit `source_start_char`/`source_end_char` ranges (`artifacts.text_source_range_projection`, the shared work-order range contract), with bytes, kind, path containment and SHA checked before any excerpt. Packet-only reflection receives per-send-tool counts, each family's latest recorded return, and task-related skill readiness with coverage; full-source references are for later readers, not evidence the synthesizer has read. Positive observed facts correct error-trace impressions, while tool success does not prove owner receipt, empty material does not prove absence, and skill readiness does not attribute an owner's action to the task. Before context cleanup, `agent_task_pipeline.emit_task_results` also freezes `review_evidence.task_inputs` through `post_task_synthesis.capture_task_inputs`: `run_origin`, the existing task-local owner corpus, intact question/answer provenance and the canonical split-root verification-receipt union. Reflection receives the same complete redacted content through `reflection.task_inputs_prompt_section`, separate from bounded trace/review excerpts. A zero return code is positive evidence; an unrelated later pass cannot resolve another check's failure. Peer proposals stay attributed, and unavailable input is not evidence that approval or verification never existed. Recovery uses these stored observations and inputs, not a later conversation. A free `host_task_facts` row precedes paid stages (or follows result persistence when Stop skips them): no model call or narrative; its metrics, routing and cost serve history. `files_rescued` (TZ-2 C2) is a stat-only file count of the artifact stores: positive, zero or unknown if unreadable, `hash_computed: false`; the stop receipt repeats it so no salvageable text never means no files.
|
||||
|
||||
Pooled workers retain their slot until post-task work settles; early final-answer delivery is independent of that timing. Native work stays on its registered actor; `TaskModelWait` remains reachable through `POST_TASK_SYNTHESIS_INFLIGHT`, and detached work owns a separate live wait. Mailbox cleanup waits for the checkpoint. The solve-phase absolute ceiling does not cut a settled root's running post-work, but Stop, calendar deadline, monetary admission, per-call and idle rails still bind. The stage owner reads returned memory errors: budget or unresolved provider attempts skip later paid stages and degrade the checkpoint; an ordinary failure stays local — later stages still run, completed reflection actions still apply — but the stage that lost work is unfinished, so the checkpoint reads `degraded` with no skipped list, never `completed` (TZ-2 C3). A running checkpoint after restart degrades without replaying a paid request.
|
||||
|
||||
|
|
|
|||
|
|
@ -777,7 +777,10 @@ def mailbox_drain_ended(task_drive: pathlib.Path, task_id: str) -> bool:
|
|||
(TZ-2 D15). Its mailbox is then only cleaned up, never read again, so owner
|
||||
mail and quiz answers must not be labelled delivered into it — the routing
|
||||
guard and the quiz ingress both ask this one fact. The actor's drive is read,
|
||||
not the canonical row: split-root copyback can lag the settlement.
|
||||
not the canonical row: split-root copyback can lag the settlement. A receipt for
|
||||
mail that queued after the drain ended (TZ-2 B5) is not built yet: it waits for
|
||||
the artifact/forwarding API TZ-1 lands in ``origin/ouroboros`` and is not to be
|
||||
copied from provisional code.
|
||||
"""
|
||||
from ouroboros.task_results import load_task_result
|
||||
from ouroboros.task_status import SETTLED_STATUSES
|
||||
|
|
|
|||
|
|
@ -517,11 +517,15 @@ def _record_task_facts(env: Any, task: Dict[str, Any], usage: Dict[str, Any],
|
|||
task_id = str(task.get("id") or "unknown")
|
||||
try:
|
||||
from ouroboros.project_dialogue import append_canonical_task_summary, completion_status_label, outcome_phase
|
||||
from ouroboros.task_finalization import artifact_store_roots, rescued_files_fact
|
||||
|
||||
canonical_root = pathlib.Path(task.get("budget_drive_root") or drive_logs.parent)
|
||||
result_root = pathlib.Path(getattr(env, "drive_root", canonical_root))
|
||||
stored_result = _atp().load_task_result(result_root, task_id) or {}
|
||||
review_projection = _compact_review_projection(llm_trace)
|
||||
# TZ-2 C2: how many files the task rescued into its store(s) — positive, zero or
|
||||
# unknown — by stat alone; the fact discloses that no hash was computed.
|
||||
files_rescued = rescued_files_fact(task_id, artifact_store_roots(canonical_root, task_id, child_root=result_root))
|
||||
append_canonical_task_summary(canonical_root, {
|
||||
"ts": utc_now_iso(), "direction": "system", "type": "task_summary",
|
||||
"summary_kind": "host_task_facts", "summary_id": f"task-facts:{task_id}",
|
||||
|
|
@ -537,6 +541,7 @@ def _record_task_facts(env: Any, task: Dict[str, Any], usage: Dict[str, Any],
|
|||
"rounds": None if usage.get("loop_evidence_unavailable") else int(usage.get("rounds") or 0),
|
||||
"outcome_axes": normalize_outcome_axes(usage), "reason_code": str(usage.get("reason_code") or ""),
|
||||
"result_ref": {"kind": "task_result", "task_id": task_id, "reader": "get_task_result"},
|
||||
"files_rescued": files_rescued,
|
||||
**_summary_row_cost_fields(usage), **presence_provenance_fields(task),
|
||||
**({"review_projection": review_projection} if review_projection.get("panels") else {}),
|
||||
})
|
||||
|
|
|
|||
|
|
@ -27,6 +27,7 @@ from __future__ import annotations
|
|||
import hashlib
|
||||
import json
|
||||
import logging
|
||||
import os
|
||||
import pathlib
|
||||
from typing import Any, Dict, List
|
||||
|
||||
|
|
@ -599,6 +600,99 @@ def sealed_final_prompt_section(sealed_final: Dict[str, Any] | None) -> str:
|
|||
)
|
||||
|
||||
|
||||
# TZ-2 C2: a store's own bookkeeping is not a rescued file. Mirrors
|
||||
# ``artifacts._ARTIFACT_MANIFEST`` and ``workspace_patch_capture.SCRATCH_MANIFEST_NAME``
|
||||
# (importing the D05 store owner here would invert the terminal-facts direction);
|
||||
# tests/test_task_summary.py pins equality with the SSOT so the literals cannot drift.
|
||||
RESCUED_FILES_BOOKKEEPING = frozenset({".artifact_manifest.json", ".scratch_manifest.json"})
|
||||
|
||||
|
||||
def artifact_store_roots(canonical_root: Any, task_id: str, *, task: Any = None,
|
||||
child_root: Any = None) -> List[pathlib.Path]:
|
||||
"""The task's artifact store directories: the canonical one and, for a split root, the child drive's.
|
||||
|
||||
The child drive is the caller's when it knows it (the pipeline's ``env.drive_root``),
|
||||
else the task row's (``child_drive_root`` / ``drive_root``, the supervisor's own
|
||||
resolution of a running task's drive), else the durable result's; the same store
|
||||
named twice is walked once. Fail-soft: an unreadable result adds no store.
|
||||
"""
|
||||
from ouroboros.headless import ARTIFACTS_DIR
|
||||
|
||||
row = task if isinstance(task, dict) else {}
|
||||
child = str(child_root or row.get("child_drive_root") or row.get("drive_root") or "").strip()
|
||||
if not child and canonical_root:
|
||||
try:
|
||||
from ouroboros.task_results import load_task_result
|
||||
|
||||
stored = load_task_result(pathlib.Path(canonical_root), str(task_id)) or {}
|
||||
child = str(stored.get("child_drive_root") or stored.get("headless_child_drive_root")
|
||||
or stored.get("drive_root") or "").strip()
|
||||
except Exception:
|
||||
log.debug("artifact store roots: durable child drive unreadable for %s", task_id, exc_info=True)
|
||||
stores: List[pathlib.Path] = []
|
||||
for root in (str(canonical_root or ""), child):
|
||||
store = pathlib.Path(root) / ARTIFACTS_DIR / str(task_id)
|
||||
if root and store.resolve(strict=False) not in [known.resolve(strict=False) for known in stores]:
|
||||
stores.append(store)
|
||||
return stores
|
||||
|
||||
|
||||
def rescued_files_fact(task_id: str, stores: List[pathlib.Path]) -> Dict[str, Any]:
|
||||
"""How many files the task's artifact stores hold, by ``stat`` alone (TZ-2 C2).
|
||||
|
||||
``state`` is ``positive`` (files were found), ``zero`` (every store was walked and
|
||||
holds none; a store never created is one nothing was written to) or ``unknown`` (a
|
||||
store could not be read — ``count`` is then what the readable part held, a floor,
|
||||
never a total). No file is opened and no hash is computed, and the fact says so
|
||||
(``hash_computed``), so a reader never mistakes this occupancy count for the
|
||||
per-file receipt the artifact manifest (``collect_task_artifact_records``) owns;
|
||||
conversely an empty readable manifest alone never proves zero — only the walk does.
|
||||
Never raises: a failure inside the walk is an unknown store.
|
||||
"""
|
||||
rows: List[Dict[str, Any]] = []
|
||||
for store in stores:
|
||||
try:
|
||||
count, readable = _stat_only_file_count(pathlib.Path(store))
|
||||
except Exception:
|
||||
count, readable = 0, False
|
||||
rows.append({"store": str(store), "count": count, "readable": readable})
|
||||
total = sum(int(row["count"]) for row in rows)
|
||||
unreadable = not rows or any(not row["readable"] for row in rows)
|
||||
state = "unknown" if unreadable else ("positive" if total else "zero")
|
||||
return {"count": total, "state": state, "hash_computed": False, "stores": rows}
|
||||
|
||||
|
||||
def _stat_only_file_count(store: pathlib.Path) -> tuple[int, bool]:
|
||||
"""(regular files below ``store`` minus bookkeeping, whether every directory was readable)."""
|
||||
count, readable, pending = 0, True, [store]
|
||||
while pending:
|
||||
directory = pending.pop()
|
||||
try:
|
||||
with os.scandir(directory) as entries:
|
||||
for entry in entries:
|
||||
if entry.is_dir(follow_symlinks=False):
|
||||
pending.append(pathlib.Path(entry.path))
|
||||
elif entry.is_file(follow_symlinks=False) and entry.name not in RESCUED_FILES_BOOKKEEPING:
|
||||
count += 1
|
||||
except FileNotFoundError:
|
||||
if directory != store:
|
||||
readable = False # a subdirectory vanished mid-walk: the count is a floor
|
||||
except OSError:
|
||||
readable = False # permission denied, a file where the store should be, an I/O error
|
||||
return count, readable
|
||||
|
||||
|
||||
def rescued_files_sentence(fact: Dict[str, Any]) -> str:
|
||||
"""ONE owner sentence for the stop receipt: the count, its state, and that no hash was computed."""
|
||||
state, count = str(fact.get("state") or "unknown"), int(fact.get("count") or 0)
|
||||
if state == "positive":
|
||||
return f"Files rescued: {count} in the task's artifact store (counted by stat; hashes not computed)."
|
||||
if state == "zero":
|
||||
return "Files rescued: none — the task's artifact store was walked and holds no files (hashes not computed)."
|
||||
seen = f"; {count} seen before the failure" if count else ""
|
||||
return f"Files rescued: unknown — the task's artifact store could not be read{seen} (hashes not computed)."
|
||||
|
||||
|
||||
def model_execution_projection(usage: Dict[str, Any]) -> Dict[str, Any] | None:
|
||||
"""Project the last usable ordinary solve response, not post-task authorship."""
|
||||
initial = usage.get("initial_model_request")
|
||||
|
|
|
|||
|
|
@ -431,6 +431,9 @@ def _register_followup(ctx: ToolContext, task_id: str, drive_root: Any,
|
|||
# `or ""` before str(): an absent/None root_task_id must fall back
|
||||
# to task_id, never become the literal string "None".
|
||||
"origin_root_task_id": str(root_task_id or "") or task_id,
|
||||
# TZ-2 B3: the same author stamp a promote carries — the successor's
|
||||
# first turn is this task's note, framed as such, never an owner directive.
|
||||
"objective_author": {"kind": "task", "task_id": task_id},
|
||||
},
|
||||
**({"chat_id": source_chat_id} if source_chat_id not in (None, "") else {}),
|
||||
},
|
||||
|
|
|
|||
|
|
@ -34,9 +34,8 @@ from typing import Any, Dict, List, Optional
|
|||
|
||||
from ouroboros.utils import update_json_locked, utc_now_iso
|
||||
from ouroboros.task_finalization import (
|
||||
HOST_AUTHORED_TERMINAL_ORIGINS,
|
||||
TERMINAL_ORIGIN_HOST_SALVAGE,
|
||||
TERMINAL_ORIGIN_MODEL_FINAL,
|
||||
HOST_AUTHORED_TERMINAL_ORIGINS, TERMINAL_ORIGIN_HOST_SALVAGE, TERMINAL_ORIGIN_MODEL_FINAL,
|
||||
artifact_store_roots, rescued_files_fact, rescued_files_sentence,
|
||||
)
|
||||
|
||||
log = logging.getLogger(__name__)
|
||||
|
|
@ -1281,7 +1280,7 @@ def _persist_cancel_receipt(
|
|||
preserved_path: str, preview_omitted: int,
|
||||
children: Optional[List[Dict[str, Any]]] = None,
|
||||
unreconciled_runs: Optional[List[str]] = None,
|
||||
reason_code: str = "",
|
||||
reason_code: str = "", files_rescued: Optional[Dict[str, Any]] = None,
|
||||
) -> None:
|
||||
"""Q5=A: the technical stop facts live in the task DETAILS panel.
|
||||
|
||||
|
|
@ -1293,6 +1292,7 @@ def _persist_cancel_receipt(
|
|||
digest for a cascade root. Never creates the result file (a later full
|
||||
write would clobber a block-only row) and never clobbers previously
|
||||
persisted non-empty facts with an emptier rebuild. Fail-soft.
|
||||
``files_rescued`` (TZ-2 C2) is the typed stat-only store count the receipt text speaks.
|
||||
"""
|
||||
tid = str(task_id or "")
|
||||
try:
|
||||
|
|
@ -1310,6 +1310,7 @@ def _persist_cancel_receipt(
|
|||
else {"path": "", "preserved": False}
|
||||
),
|
||||
"ts": utc_now_iso(),
|
||||
**({"files_rescued": dict(files_rescued)} if files_rescued else {}),
|
||||
}
|
||||
if str(reason_code or ""):
|
||||
block["reason_code"] = str(reason_code) # TZ-2 C1: the typed rail beside its sentence
|
||||
|
|
@ -1443,6 +1444,7 @@ def build_unreviewed_salvage_event(
|
|||
details panel — not in chat.
|
||||
``reason_code`` (TZ-2 C1): the TYPED rail; with an empty ``outcome`` the owner
|
||||
sentence comes from TASK_CAUSE_PHRASES here, and the code rides the receipt, not the prose.
|
||||
TZ-2 C2: text and receipt state the files rescued (positive/zero/unknown, stat-only walk, no hashes).
|
||||
"""
|
||||
from ouroboros.task_results import STATUS_COMPLETED
|
||||
|
||||
|
|
@ -1467,21 +1469,19 @@ def build_unreviewed_salvage_event(
|
|||
if str(settled_status or "").strip().lower() == STATUS_COMPLETED:
|
||||
# Completion-wins (owner 4=A): the kept result is the real answer, not a
|
||||
# salvage — but it still bypassed the normal delivery path, so say so.
|
||||
lines = [
|
||||
f"✅ Task {tid} {outcome_text}. Its completed result is preserved below."
|
||||
+ descendants,
|
||||
]
|
||||
lines = [f"✅ Task {tid} {outcome_text}. Its completed result is preserved below." + descendants]
|
||||
else:
|
||||
lines = [
|
||||
f"⚠️ Task {tid} was {outcome_text}. Below is the last persisted "
|
||||
"intermediate model message, preserved WITHOUT review (salvaged "
|
||||
"best-effort; NOT a final answer)." + descendants,
|
||||
]
|
||||
if preview:
|
||||
lines += ["", preview, "", _preview_note_line(preserved_path, omitted)]
|
||||
else:
|
||||
lines += ["", "(no salvageable agent output was found for this task)"]
|
||||
disclosure_lines: List[str] = []
|
||||
lines += ["", preview, "", _preview_note_line(preserved_path, omitted)] if preview else [
|
||||
"", "(no salvageable agent output was found for this task)"]
|
||||
# TZ-2 C2: "no salvageable text" never implies "no files": the stat-only store count
|
||||
# rides the mutable disclosure (never the content-derived identity), hashes not computed.
|
||||
rescued = rescued_files_fact(tid, artifact_store_roots(drive_root, tid, task=task_row))
|
||||
disclosure_lines: List[str] = ["", rescued_files_sentence(rescued)]
|
||||
if unreconciled_runs:
|
||||
# GR3-7: an audit-failure marker means run state is UNKNOWN — a
|
||||
# different honest sentence than "these named runs stayed open".
|
||||
|
|
@ -1523,7 +1523,7 @@ def build_unreviewed_salvage_event(
|
|||
settled_status=str(settled_status or ""), outcome=outcome_text,
|
||||
delivery_id=did, preserved_path=str(preserved_path or ""),
|
||||
preview_omitted=omitted, children=children,
|
||||
unreconciled_runs=unreconciled_runs, reason_code=code,
|
||||
unreconciled_runs=unreconciled_runs, reason_code=code, files_rescued=rescued,
|
||||
)
|
||||
event = {
|
||||
"type": "send_message",
|
||||
|
|
|
|||
|
|
@ -641,3 +641,52 @@ def test_receipt_names_the_stop_cause_before_and_after_the_settle(tmp_path):
|
|||
)
|
||||
plain = load_task_result(tmp_path, "task-no-cause")["cancel_receipt"]
|
||||
assert "stop_reason" not in plain and "stop_requested_at" not in plain
|
||||
|
||||
|
||||
def test_salvage_receipt_states_files_rescued_even_without_salvageable_text(tmp_path):
|
||||
"""TZ-2 C2: "(no salvageable agent output ...)" must not read as "no files". The
|
||||
receipt states the stat-only artifact-store count — positive, zero or unknown — and
|
||||
that no hashes were computed; the typed fact rides ``cancel_receipt``. The count is a
|
||||
mutable disclosure, never part of the content-derived delivery identity."""
|
||||
from ouroboros.headless import task_artifacts_dir
|
||||
from supervisor import terminal_delivery as td
|
||||
|
||||
def build(tid, task=None):
|
||||
return td.build_unreviewed_salvage_event(
|
||||
tmp_path, task or {"chat_id": 4}, tid, outcome="cancelled", settled_status="cancelled")
|
||||
|
||||
write_task_result(tmp_path, "files-1", STATUS_RUNNING, result="working")
|
||||
store = task_artifacts_dir(tmp_path, "files-1")
|
||||
(store / "draft.docx").write_bytes(b"x")
|
||||
event = build("files-1")
|
||||
assert "(no salvageable agent output was found for this task)" in event["text"]
|
||||
assert "Files rescued: 1 " in event["text"] and "hashes not computed" in event["text"], event["text"]
|
||||
receipt = load_task_result(tmp_path, "files-1")["cancel_receipt"]
|
||||
assert receipt["files_rescued"] == {"count": 1, "state": "positive", "hash_computed": False,
|
||||
"stores": [{"store": str(store), "count": 1, "readable": True}]}
|
||||
(store / "more.txt").write_bytes(b"y")
|
||||
rebuilt = build("files-1")
|
||||
assert rebuilt["delivery_id"] == event["delivery_id"] and "Files rescued: 2 " in rebuilt["text"]
|
||||
assert load_task_result(tmp_path, "files-1")["cancel_receipt"]["files_rescued"]["count"] == 2
|
||||
|
||||
write_task_result(tmp_path, "files-0", STATUS_RUNNING, result="working")
|
||||
task_artifacts_dir(tmp_path, "files-0")
|
||||
zero = build("files-0")
|
||||
assert "Files rescued: none" in zero["text"] and "hashes not computed" in zero["text"], zero["text"]
|
||||
assert load_task_result(tmp_path, "files-0")["cancel_receipt"]["files_rescued"]["state"] == "zero"
|
||||
|
||||
write_task_result(tmp_path, "files-x", STATUS_RUNNING, result="working")
|
||||
task_artifacts_dir(tmp_path, "files-x", create=False).write_text("not a directory", encoding="utf-8")
|
||||
unknown = build("files-x")
|
||||
assert "Files rescued: unknown" in unknown["text"], unknown["text"]
|
||||
assert load_task_result(tmp_path, "files-x")["cancel_receipt"]["files_rescued"]["state"] == "unknown"
|
||||
|
||||
# A split root: the child drive named on the task row is walked beside the canonical store.
|
||||
child = tmp_path / "child-drive"
|
||||
write_task_result(tmp_path, "files-s", STATUS_RUNNING, result="working")
|
||||
(task_artifacts_dir(child, "files-s") / "out.txt").write_bytes(b"o")
|
||||
split = build("files-s", {"chat_id": 4, "child_drive_root": str(child)})
|
||||
assert "Files rescued: 1 " in split["text"], split["text"]
|
||||
stores = load_task_result(tmp_path, "files-s")["cancel_receipt"]["files_rescued"]["stores"]
|
||||
assert [row["store"] for row in stores] == [str(task_artifacts_dir(tmp_path, "files-s", create=False)),
|
||||
str(task_artifacts_dir(child, "files-s", create=False))]
|
||||
|
|
|
|||
|
|
@ -96,7 +96,9 @@ CHAPTER_BYTE_BUDGETS: dict[str, int] = {
|
|||
# projection write per turn (the unbounded drain and per-event write they replace had
|
||||
# no sentence of their own), and the projection paragraph states the writer's slim read,
|
||||
# its retry interval and the crossing rule of the OpenRouter check.
|
||||
"docs/architecture/05-supervisor-loop.md": 32400,
|
||||
# 32400 -> 32600 (tz2 aafaa3713): the D15 late-answer / drain-ended sentence in the
|
||||
# owner-wait paragraph took the chapter to 32502 before this diff; nothing displaced.
|
||||
"docs/architecture/05-supervisor-loop.md": 32600,
|
||||
# 286850 -> 287600: "an answer that has not arrived is a gap" is a new invariant of
|
||||
# plan review and task acceptance (the slot census vocabulary, the `awaiting`
|
||||
# projection, the only-awaited task outcome); the in-flight sentence it grew from is
|
||||
|
|
@ -247,7 +249,9 @@ CHAPTER_BYTE_BUDGETS: dict[str, int] = {
|
|||
# descendant is a Presence caller (inherited binding authority, never speaker metadata).
|
||||
# 95150 -> 95200 (TZ2 repair, measured 95191): the promotion/follow-up clause names the
|
||||
# one carrier it copies instead of "the Presence metadata".
|
||||
"docs/development/06-rules-by-change-class.md": 95200,
|
||||
# 95200 -> 95400 (tz2 7a387f717): the C4 explicit-stop rule took the chapter to 95223
|
||||
# before this diff; nothing displaced.
|
||||
"docs/development/06-rules-by-change-class.md": 95400,
|
||||
"docs/development/07-managed-update-rule.md": 4166,
|
||||
"docs/development/08-mutation-attribution-rule.md": 2899,
|
||||
"docs/development/09-process-custody-rule.md": 10028,
|
||||
|
|
|
|||
|
|
@ -212,6 +212,38 @@ def test_schedule_followup_registers_a_one_shot_entry(tmp_path):
|
|||
assert record["task"]["context"] == "plan review for root-1 was quorum-unreachable"
|
||||
|
||||
|
||||
def test_a_timer_follow_up_is_framed_as_task_authored_and_its_note_never_enters_the_owner_corpus(tmp_path):
|
||||
"""TZ-2 B3: the successor of a timer follow-up reads the same ``objective_author``
|
||||
stamp a promote / route_to_project carries, so its first turn is framed as drafted
|
||||
by the scheduling task and the note text never becomes an owner directive. The
|
||||
owner door's stamp (``origin_message_ref``) is still not carried (#1271)."""
|
||||
from types import SimpleNamespace
|
||||
|
||||
from ouroboros.context import build_user_content
|
||||
from ouroboros.dialogue_provenance import run_origin
|
||||
from ouroboros.loop_messages import _initialize_owner_directives
|
||||
from supervisor import queue
|
||||
|
||||
ctx = _ctx(tmp_path)
|
||||
ctx.task_metadata["origin_message_ref"] = {"chat_id": 1, "message_id": "owner-7"}
|
||||
note = "Re-check the reviewer window and resume the plan."
|
||||
assert _followup(ctx, objective=note).startswith("FOLLOWUP_SCHEDULED")
|
||||
[record] = queue.list_scheduled_tasks(tmp_path / "data")["tasks"]
|
||||
author = {"kind": "task", "task_id": "root-1"}
|
||||
assert record["task"]["metadata"]["objective_author"] == author
|
||||
assert "origin_message_ref" not in record["task"]["metadata"]
|
||||
successor = queue._task_from_schedule(record)
|
||||
assert successor["metadata"]["objective_author"] == author
|
||||
origin = run_origin(successor)
|
||||
assert (origin["owner_ingress"], origin["objective_author"]) == (False, author)
|
||||
content = build_user_content(successor)
|
||||
assert content.startswith("[OBJECTIVE_AUTHOR] The objective below was drafted by task root-1, "), content
|
||||
assert content.endswith("[/OBJECTIVE_AUTHOR]\n\n" + note), content
|
||||
live = SimpleNamespace(task_metadata=successor["metadata"])
|
||||
_initialize_owner_directives(live, [{"role": "user", "content": content}])
|
||||
assert getattr(live, "_owner_directives", []) == [] # the note is the task's, not the owner's
|
||||
|
||||
|
||||
def test_presence_followup_preserves_ceiling_and_return_context(tmp_path):
|
||||
ctx = _ctx(tmp_path)
|
||||
ctx.task_metadata["presence"] = {"binding_id": "b" * 32}
|
||||
|
|
|
|||
|
|
@ -246,3 +246,59 @@ def test_build_trace_summary_shows_structured_failure_facts():
|
|||
"reasoning_notes": ["note" * 2000],
|
||||
}
|
||||
assert "OMISSION NOTE" in pipeline.build_trace_summary(long_trace)
|
||||
|
||||
|
||||
def test_facts_row_states_files_rescued_from_a_stat_only_walk(tmp_path, no_model_calls):
|
||||
"""TZ-2 C2: at terminal the free facts row says how many files reached the task's
|
||||
artifact store — a positive count, a confirmed zero, or unknown — from a stat-only
|
||||
walk that discloses it computed no hashes. Store bookkeeping is not a rescued file,
|
||||
an empty readable manifest alone never proves zero (the walk does), an unreadable
|
||||
store is unknown (never zero), and a split root walks the child-drive store too."""
|
||||
from ouroboros.headless import task_artifacts_dir
|
||||
|
||||
drive_logs = tmp_path / "logs"
|
||||
drive_logs.mkdir()
|
||||
|
||||
def fact(task_id, env=None, **task_extra):
|
||||
pipeline._record_task_facts(env=env, task={"id": task_id, "chat_id": 1, **task_extra},
|
||||
usage={"rounds": 1}, llm_trace={"tool_calls": []}, drive_logs=drive_logs)
|
||||
[row] = [r for r in _rows(drive_logs) if r["summary_id"] == f"task-facts:{task_id}"]
|
||||
return row["files_rescued"]
|
||||
|
||||
store = task_artifacts_dir(tmp_path, "pos-1")
|
||||
(store / "report.md").write_text("r", encoding="utf-8")
|
||||
(store / "nested").mkdir()
|
||||
(store / "nested" / "data.csv").write_text("1,2", encoding="utf-8")
|
||||
(store / ".artifact_manifest.json").write_text("{}", encoding="utf-8")
|
||||
(store / ".scratch_manifest.json").write_text("{}", encoding="utf-8")
|
||||
assert fact("pos-1") == {"count": 2, "state": "positive", "hash_computed": False,
|
||||
"stores": [{"store": str(store), "count": 2, "readable": True}]}
|
||||
|
||||
store = task_artifacts_dir(tmp_path, "zero-1")
|
||||
(store / ".artifact_manifest.json").write_text('{"schema_version": 1, "artifacts": {}}', encoding="utf-8")
|
||||
assert fact("zero-1") == {"count": 0, "state": "zero", "hash_computed": False,
|
||||
"stores": [{"store": str(store), "count": 0, "readable": True}]}
|
||||
never_created = task_artifacts_dir(tmp_path, "none-1", create=False)
|
||||
assert fact("none-1")["state"] == "zero" and not never_created.exists()
|
||||
|
||||
blocked = task_artifacts_dir(tmp_path, "unk-1", create=False)
|
||||
blocked.write_text("a file where the store directory should be", encoding="utf-8")
|
||||
assert fact("unk-1") == {"count": 0, "state": "unknown", "hash_computed": False,
|
||||
"stores": [{"store": str(blocked), "count": 0, "readable": False}]}
|
||||
|
||||
child = tmp_path / "child-drive"
|
||||
(task_artifacts_dir(child, "split-1") / "out.txt").write_text("o", encoding="utf-8")
|
||||
canonical = task_artifacts_dir(tmp_path, "split-1")
|
||||
assert fact("split-1", env=SimpleNamespace(drive_root=child), budget_drive_root=str(tmp_path)) == {
|
||||
"count": 1, "state": "positive", "hash_computed": False,
|
||||
"stores": [{"store": str(canonical), "count": 0, "readable": True},
|
||||
{"store": str(task_artifacts_dir(child, "split-1", create=False)), "count": 1, "readable": True}]}
|
||||
|
||||
|
||||
def test_rescued_files_walk_excludes_exactly_the_store_bookkeeping_names():
|
||||
"""The bookkeeping names the walk skips are the SSOT literals, pinned so they cannot drift."""
|
||||
from ouroboros.artifacts import _ARTIFACT_MANIFEST
|
||||
from ouroboros.task_finalization import RESCUED_FILES_BOOKKEEPING
|
||||
from ouroboros.workspace_patch_capture import SCRATCH_MANIFEST_NAME
|
||||
|
||||
assert RESCUED_FILES_BOOKKEEPING == frozenset({_ARTIFACT_MANIFEST, SCRATCH_MANIFEST_NAME})
|
||||
|
|
|
|||
88
tests/test_ui_quiz_comment_only_browser.py
Normal file
88
tests/test_ui_quiz_comment_only_browser.py
Normal file
|
|
@ -0,0 +1,88 @@
|
|||
"""TZ-2 B1: a comment-only question (an ``escalate`` with zero options) reaches the real
|
||||
web UI as a quiz card with no option buttons and the free-answer box, and the owner's
|
||||
typed answer travels verbatim through the real gateway: the durable quiz block records
|
||||
no chosen option and the exact text, and the card shows it as the owner's answer.
|
||||
|
||||
Only model judgment is a fixture (the scripted stub asks, then finishes); the server,
|
||||
the worker, the WebSocket ingress and the decision endpoint are the production ones.
|
||||
"""
|
||||
import json
|
||||
import os
|
||||
import uuid
|
||||
from pathlib import Path
|
||||
|
||||
import pytest
|
||||
|
||||
from tests.test_owner_wait_integration import wait_clone as clone_fixture
|
||||
from tests.test_native_owner_wait_browser import chat_connection
|
||||
from tests.system_e2e.harness import (
|
||||
ArtifactOracle, KeylessIsolatedServer, ScriptedStubModel, keyless_settings,
|
||||
wait_durable_result, wait_until, write_settings_file,
|
||||
)
|
||||
|
||||
wait_clone = clone_fixture
|
||||
pytestmark = [pytest.mark.serial, pytest.mark.browser]
|
||||
|
||||
QUESTION = "What deadline should the report state on its cover page?"
|
||||
ANSWER = "Friday, 3 October — and say it is provisional."
|
||||
|
||||
|
||||
def test_comment_only_question_renders_and_takes_a_free_text_answer(wait_clone, tmp_path):
|
||||
from playwright.sync_api import sync_playwright
|
||||
|
||||
root = tmp_path / "instance" / "data"
|
||||
root.mkdir(parents=True)
|
||||
screenshots = Path(os.environ.get("OUROBOROS_BROWSER_EVIDENCE_OUT") or tmp_path / "screenshots")
|
||||
screenshots.mkdir(parents=True, exist_ok=True)
|
||||
steps = [
|
||||
{"tool": "escalate", "arguments": {"question": QUESTION, "options": [],
|
||||
"stake": "The deadline is printed on the cover page.", "wait_for_answer": True}},
|
||||
{"final": "The report states the deadline you gave."},
|
||||
]
|
||||
with ScriptedStubModel(steps) as stub:
|
||||
settings_path = root / "settings.json"
|
||||
write_settings_file(settings_path, keyless_settings(stub, OUROBOROS_MAX_WORKERS=1))
|
||||
server = KeylessIsolatedServer(wait_clone, root, settings_path)
|
||||
server.start(ready_timeout=120)
|
||||
oracle = ArtifactOracle(root)
|
||||
try:
|
||||
with sync_playwright() as pw:
|
||||
browser = pw.chromium.launch()
|
||||
page = browser.new_page(viewport={"width": 1280, "height": 900}, reduced_motion="reduce")
|
||||
try:
|
||||
page.goto(server.base_url, wait_until="domcontentloaded")
|
||||
with chat_connection(server) as ws:
|
||||
message_id = uuid.uuid4().hex
|
||||
ws.send(json.dumps({"type": "chat", "content": "Ask me for the deadline, then wait for my answer.",
|
||||
"client_message_id": message_id, "chat_id": 1}))
|
||||
task = wait_until(lambda: next((row["task"] for row in oracle.events("task_received")
|
||||
if (row.get("task", {}).get("metadata", {}).get("origin_message_ref") or {}).get("client_message_id") == message_id), None), 90)
|
||||
assert task
|
||||
wait = wait_until(lambda: (block if (block := oracle.task_result(task["id"]).get("owner_wait", {})).get("state") == "waiting" else None), 90)
|
||||
assert wait and wait.get("quiz_id"), wait
|
||||
card = page.locator(f'#chat-messages .chat-quiz-card[data-task-id="{task["id"]}"][data-quiz-id="{wait["quiz_id"]}"]')
|
||||
card.get_by_text(QUESTION, exact=True).wait_for(timeout=30000)
|
||||
card.locator('.chat-quiz-comment').wait_for(timeout=30000)
|
||||
# Zero options: no option button at all, only the free-answer box and its send.
|
||||
assert card.locator('.chat-quiz-option').count() == 0
|
||||
assert card.locator('.chat-quiz-comment').count() == 1
|
||||
assert card.locator('.chat-quiz-send').is_disabled()
|
||||
card.scroll_into_view_if_needed()
|
||||
card.screenshot(animations="disabled", path=str(screenshots / "chromium-comment-only-question.png"))
|
||||
card.locator('.chat-quiz-comment').fill(ANSWER)
|
||||
assert card.locator('.chat-quiz-send').is_enabled()
|
||||
card.locator('.chat-quiz-send').click()
|
||||
result = wait_durable_result(oracle, task["id"], timeout=120)
|
||||
block = result["owner_quiz"][wait["quiz_id"]]
|
||||
assert block.get("answered_index") is None and block["comment"] == ANSWER, block
|
||||
assert result["status"] == "completed", result.get("status")
|
||||
card.locator('.chat-quiz-answer').filter(has_text=f"Owner's answer: {ANSWER}").wait_for(timeout=30000)
|
||||
wait_until(lambda: card.get_attribute("data-state") == "answered" or None, 30)
|
||||
assert card.locator('.chat-quiz-comment').count() == 0 # a settled card takes no second answer
|
||||
card.screenshot(animations="disabled", path=str(screenshots / "chromium-comment-only-answered.png"))
|
||||
(screenshots / "comment-only.json").write_text(json.dumps({
|
||||
"task_id": task["id"], "quiz_id": wait["quiz_id"], "owner_quiz": block}, indent=2))
|
||||
finally:
|
||||
browser.close()
|
||||
finally:
|
||||
server.stop()
|
||||
Loading…
Add table
Add a link
Reference in a new issue