diff --git a/docs/DESIGN.md b/docs/DESIGN.md index 1e27baf3f..92a14b7e4 100644 --- a/docs/DESIGN.md +++ b/docs/DESIGN.md @@ -280,7 +280,14 @@ in details and Logs, but do not relabel the whole still-working task. A failed c authoritative status. Internal reason codes belong in details and diagnostics, not compact headlines. Where a card does show a cause, it says it in the owner's words while the record keeps the machine code; a cause with no sentence yet stays -raw rather than borrowing a wrong one. The routing receipt under an owner +raw rather than borrowing a wrong one. The rails that end a task are such +causes: the loop's forced finalization (round limit, deadline, grace window, +context, unabsorbed children) and the supervisor's timeout reaper (maximum +running time, deadline, idle silence) keep their typed codes on the record and +on the incident key, and the card, the reaper's grace toast, kill notice and +salvage line, and the loop's own fallback text all say the one sentence from +the shared table (`project_dialogue.TASK_CAUSE_PHRASES`, whose browser twin +lives in `log_events.js`). The routing receipt under an owner message is such a surface: a refused addressing act carries the host-composed `cause` sentence (`project_dialogue.routing_refusal_cause` — one host table for the receipt line, the System row and the picker toast), a landed act carries @@ -630,7 +637,7 @@ answer keep both forms readable. Anatomy, top to bottom: one word that answers "is there an unanswered question for me?": `Waiting for your answer` needs positive wait evidence (the task's live wait record, or the original required flag before any record exists); a resumed - wait — the owner typed instead, or the bound closed — reads `Unanswered · the + wait — owner input or other mail woke the task, or the bound closed — reads `Unanswered · the task continued; an answer is still accepted`; an open question without any wait evidence reads `Unanswered · an answer is still accepted`; `Unanswered · the task finished; a late answer is accepted as your message` @@ -647,11 +654,11 @@ answer keep both forms readable. Anatomy, top to bottom: does not infer a title from the first line or rewrite authored Markdown to make it fit. 3. **Stake** — optional one-liner (`At stake: …`), `--type-meta`, `--text-meta`. -4. **Options** — real owner actions: buttons with `--text-primary` labels, +4. **Options** — zero to six real owner actions: buttons with `--text-primary` labels, legible at rest; an optional per-option detail steps down to meta ink. After settlement buttons drop to `--text-disabled`; the chosen option keeps the ok pair. Options are capped by the shared Python↔JS constant - (`MAX_QUIZ_OPTIONS`). + (`MAX_QUIZ_OPTIONS`); with none, the free answer is the whole answer. 5. **Free answer** — while the card is open, a compact always-visible field (`Your answer or comment…`) with a `Send my answer` button, enabled only once something is typed. No option ever has to be the least wrong one: the diff --git a/docs/architecture/05-supervisor-loop.md b/docs/architecture/05-supervisor-loop.md index 72876ba79..b4d21fa75 100644 --- a/docs/architecture/05-supervisor-loop.md +++ b/docs/architecture/05-supervisor-loop.md @@ -26,7 +26,7 @@ At actual execution start, `agent._persist_running_record` mirrors a split root' Pooled completion separates a finished file-save attempt from publishable terminal truth: the worker prepares its terminal task files (`headless.prepare_terminal_task_files`) before its buffered `task_done`, and `events_task_done` publishes only when `headless.terminal_task_files_ready` confirms CURRENT — on a split drive the child-bound copyback, with no workspace artifact finalization pending. Legacy or faulted completions go through `enqueue_terminal_file_recovery` (`worker_health` prepares and recovers, `task_reaper` runs the queues): unreadable or missing CURRENT keeps RUNNING ownership and retries on the health cadence — never a false Done, never replayed model work — and only CONFIRMED absence of a terminal source reaches the lifecycle-fault owner, which marks execution `infra_failed` while an early sticky completed status, the authored answer, review and cost survive. -A required owner wait keeps its task RUNNING and retains the same worker, command queue, live browser and services; the park is durable (completed-tool checkpoint, `owner_wait` projection) before the snapshot confirms it. The worker lends only its `active_capacity`, which the assignment tick replenishes; addressed input requests a wake, which waits for active capacity, and neither a failed grant nor maintenance reservations can buy an extra replacement beyond the configured cap. Resumption never replays: same attempt, same start timestamp, continuation authority consumed before dispatch. Waiting spares only the idle rail — stop, deadline and absolute ceiling keep their authority. Ordinary Main/Project roots wait in-process instead (`owner_wait.direct_owner_wait`): no pooled slot is held, lent or synthesized, and the saved source is evidence, never a cold-restart grant — an addressed text or quiz answer resumes the same live stack and browser. +A required owner wait keeps the task RUNNING, its worker and browser; a completed-tool checkpoint precedes parking. The worker lends active capacity; resumption needs a grant within the configured cap. Resumption never replays: same attempt, same start timestamp, continuation authority consumed before dispatch. Any mail wakes; only an owner answer answers. Stop, deadline and ceiling still bind. Ordinary Main/Project roots wait in-process instead (`owner_wait.direct_owner_wait`): no pooled slot is held, lent or synthesized, and the saved source is evidence, never a cold-restart grant — an addressed text or quiz answer resumes the same live stack and browser. A crash or automatic timeout retry keeps the current-attempt checkpoint as evidence and never blindly replays the task; cold continuation exists only for a confirmed planned restart, which cannot preserve an OS browser session. An owner-requested MANAGED UPDATE is such a restart: its writer fence prepares the handoff before stopping the pool and passes the parked ids as `preserve_running_task_ids`, so a wait is requeued instead of interrupted, and a preparation failure blocks the update rather than terminalizing the wait; the re-exec seam arms the transaction only in the phases `_safe_restart_serialized` allows (launcher mode observes exit code 42), no-resume flags suppress it, and an aborted update's leftover record never authorizes a later manual Restart or rollback. A manual ROLLBACK returns to an older runtime, so it deliberately parks nothing. Cold continuation resumes the saved economics — original CostCeiling, hard clocks, the saved model's ContextFit, TaskModelWait choices, the pending budget decision of the saved round — and calendar deadlines and owner-wait time are unchanged. @@ -51,7 +51,7 @@ Beside stopping sits the owner "hurry" control: a typed task-local `kind=hurry` The event bus is process-lifetime rather than worker-generation-lifetime: respawns reuse one manager-backed queue shared by workers and direct chat, because a force-killed producer can corrupt a raw multiprocessing feeder frame and a queue rebuilt on pool rotation strands surviving producers on the old endpoint. Live-frame publication of persisted rows is exactly-once and process-symmetric: `ouroboros/utils.py::append_jsonl` streams only runtime `logs/*.jsonl` rows into the process log sink (never `chat.jsonl`, never state/memory/receipt stores), and each process suppresses the types whose live delivery has a dedicated owner (`WORKER_LOG_SINK_SUPPRESSED_TYPES`, the server superset `SERVER_LOG_SINK_SUPPRESSED_TYPES`). One persisted event produces exactly one live frame (`tests/test_log_forwarding.py`); an LLM call failure is one durable `llm_api_error` row and nothing else. -Heartbeat and progress are different evidence: a heartbeat proves a process or loop is alive; owner-visible progress and model-usage events prove the task advanced. Fresh descendant progress or queued descendants can keep an orchestrator alive, while an explicit deadline, absolute ceiling, cancellation and budget stop remain hard. After the typed finalization episode (§6), timeout handling freezes its decision under the queue lock, marks the worker `reaping`, and hands kill, join, salvage, retry and respawn to the single off-loop reaper; an orchestrator with live descendants is not blindly retried, because a retry would replay its plan and spawn a competing tree. No retry or new assignment may occupy a timed-out slot until the original process is provably dead: if kill and join cannot establish death, the reaper keeps a low-rank RUNNING result and the `reaping` slot, emits a visible `task_reaper_wedged` receipt and restart hint, and writes no terminal, `task_done`, retry or respawn — one slot is sacrificed rather than letting a still-running process race a replacement and overwrite its result; the next supervisor generation reconciles the record after old-generation process custody. +Heartbeat and progress are different evidence: a heartbeat proves a process or loop is alive; owner-visible progress and model-usage events prove the task advanced. Fresh descendant progress or queued descendants can keep an orchestrator alive, while an explicit deadline, absolute ceiling, cancellation and budget stop remain hard. After the typed finalization episode (§6), timeout handling freezes its decision under the queue lock, marks the worker `reaping`, and hands kill, join, salvage, retry and respawn to the single off-loop reaper; an orchestrator with live descendants is not blindly retried, because a retry would replay its plan and spawn a competing tree. No retry or new assignment may occupy a timed-out slot until the original process is provably dead: if kill and join cannot establish death, the reaper keeps a low-rank RUNNING result and the `reaping` slot, emits a visible `task_reaper_wedged` receipt and restart hint, and writes no terminal, `task_done`, retry or respawn — one slot is sacrificed rather than letting a still-running process race a replacement and overwrite its result; the next supervisor generation reconciles the record after old-generation process custody. Typed timeout codes (`queue_timeouts.TIMEOUT_TERMINAL_REASONS`) retain their `task_incident` identity; owner-facing grace, kill and salvage text use `project_dialogue.TASK_CAUSE_PHRASES`, with unknown codes still raw. A spawned or respawned slot is not assignable until its child's PID-bound `worker_ready` row arrives (`supervisor/worker_pool_lifecycle.py`). A live child's own `worker_starting` row, emitted before extension loading and agent construction, permits one extension of `WORKER_READY_WINDOW_SEC` to `WORKER_READY_CEILING_SEC` (300 seconds from birth, both in `runtime_limits.py`); foreign or pre-spawn rows cannot extend another slot. `worker_ready_window_extended` records that decision. A silent child keeps the original window, and logging failure cannot block startup. After `WORKER_READY_MAX_ATTEMPTS` failed attempts, `Worker.readiness_exhausted` is final for that exact slot — late events cannot reopen it. Total exhaustion, distinguished from busy/booting/reaping capacity and from a live owner-wait stack, closes pooled ingress (owner `/review` included) without blocking direct chat/control or boot/update recovery; once RUNNING completion custody has settled, `disable_exhausted_worker_pool` fails unstarted PENDING work honestly with a Restart hint, and a new task cannot clear the latch. Readiness stays separate from liveness and task idle time; a watcher error releases only still-booting, non-exhausted slots to the crash detector (`worker_ready_released`). Linux workers use forkserver; macOS and Windows use spawn. diff --git a/docs/architecture/06-agent-core.md b/docs/architecture/06-agent-core.md index 1a202a74e..0959f4567 100644 --- a/docs/architecture/06-agent-core.md +++ b/docs/architecture/06-agent-core.md @@ -48,7 +48,7 @@ The packet is delivery-conditional; the FULL packet is not. Every triad row reac Paid identity binds that semantic subject together with substantive nonempty obligation dispositions (`acceptance_paid_identity`); forensic source hashes and ingress counters alone do not buy a panel. A resubmit with the same paid identity reuses its recorded verdict for free: a clean replay can authorize acceptance, a non-clean replay keeps its verdict and the `identical_acceptance_refused` outcome — no repeated payment, and no cosmetic edit needed when real criteria or evidence change. -The configured slots are independent actors with adaptive quorum (`config.adaptive_quorum`: 2-of-N for N≥3, both for N=2, a single reviewer as loud `single_reviewer_no_diversity`; a fewer-responded shortfall stays a loud infra quorum failure); each receives one substantive interaction on its bound route — at most two physical sends for a packet row, one bounded episode for a native row, one delegated session for a session row — and a retrieving verdict is equally authoritative. Transport status, parse status, semantic verdict, criterion support, route, quorum contribution and binding hashes stay distinct, so an unavailable or malformed response cannot masquerade as a negative judgment; a panel that refuses before any transport projects `not_dispatched` on every row and on the panel, and a slot released at the dispatch barrier projects `awaiting` for transport and parse until it settles — distinct from `success`, `timeout` and `provider_transport_error`, never a failure or a verdict. `PASS`/`FAIL`/`DEGRADED` are reviewer verdicts; the host-owned completion decision is separately `accepted`, `revision_requested` or `finalized_unaccepted`, written only by `loop_acceptance._set_acceptance_decision`. A clean quorum supplies critic approval; an informed Advisory author finish supplies separate current-author authority. Neither changes the original verdict. Material FAIL and typed unavailable outcomes reach Main before it chooses how to respond, including after the last paid panel; a no-quorum outcome with real minority findings retains that partial feedback. A terminal technical failure keeps `reason=review_degraded` when no permitted author completion follows. +The configured slots are independent actors with adaptive quorum (`config.adaptive_quorum`: 2-of-N for N≥3, both for N=2, a single reviewer as loud `single_reviewer_no_diversity`; a fewer-responded shortfall stays a loud infra quorum failure); each receives one substantive interaction on its bound route — at most two physical sends for a packet row, one bounded episode for a native row, one delegated session for a session row — and a retrieving verdict is equally authoritative. Transport status, parse status, semantic verdict, criterion support, route, quorum contribution and binding hashes stay distinct, so an unavailable or malformed response cannot masquerade as a negative judgment; a panel that refuses before any transport projects `not_dispatched` on every row and on the panel, and a slot released at the dispatch barrier projects `awaiting` for transport and parse until it settles — distinct from `success`, `timeout` and `provider_transport_error`, never a failure or a verdict. `PASS`/`FAIL`/`DEGRADED` are reviewer verdicts; the host-owned completion decision is separately `accepted`, `revision_requested` or `finalized_unaccepted`, written only by `loop_acceptance._set_acceptance_decision`. A clean quorum supplies critic approval; an Advisory author finish is author authority, never a host tier or stance. Neither changes the original verdict. Material FAIL and typed unavailable outcomes reach Main before it chooses how to respond, including after the last paid panel; a no-quorum outcome with real minority findings retains that partial feedback. A terminal technical failure keeps `reason=review_degraded` when no permitted author completion follows. A clean criterion is evidence-resolved, not only well argued: reviewer `evidence_refs` must be exact members of the packet's enumerable reference vocabulary, and a claim id resolves only through `acceptance_support_refs` linked to a passing host receipt for that claim. Agent-supplied, declared-intent, unattested and non-resolving sections never certify success; an OPEN plan wave binds nothing — its claims are disclosed as `acceptance_claims_source='none_open_plan_wave'` beside a non-binding `plan_claims_exhibit` inside `DECLARED_INTENT_SECTIONS`, so citing it never resolves and the task is distinguishable from one that never had claims. An unresolved reference keeps the actor's record for audit but removes its clean contribution (`criteria_refs_unresolved`). This total, fail-closed resolver is why the task cannot certify itself by echoing its expected outcome. @@ -525,7 +525,7 @@ Availability follows `deep_review_route`, never a window floor. An API row needs #### Post-task reflection -Typed root post-task triggers decide whether a run warrants Experience Review. `reflection.generate_reflection` sends the Light route one open prompt with the EXACT initial text (never a prefix) and its host-recorded `task_inputs.run_origin` beside it (provenance, never by itself the accepted requirement) plus tool-use, error, review and child projections and the same frozen non-final cost snapshot the task summary uses; it runs outside the tool loop, records its own usage, and its failure never erases the delivered result or changes a review verdict. Its execution trace is the ALL-CALLS listing (`build_trace_summary(all_calls=True)`): every call in order with every argument, identical consecutive calls folded into one `×N` row with their rounds, the first line of a failed or repeated call's result, and one header count of rounds whose every call was non-ok — no positional window and no literal cut, because the consolidation seam fits the call to the Light route whenever that route's window is known (an unknown window sends the prompt unchecked — the accepted residual of enlarging it); the STORED `trace_summary` (task card, parents, children) stays the bounded two-argument preview. The trace row carries the `round_id` of the model round that issued the call (absent when unknown), when the listing really cut an argument value or a failed/repeated call's answer, the redacted per-call record is retained through `retain_memory_source` and named in the prompt as OPTIONAL reading (never a required source), and its claim states the STORED bounds on both axes, not completeness: an argument already passed `sanitize_tool_args_for_log` (an oversized value carries a marker with its length and sha) and a result is the stored actor-visible cap — more than the listing, which shows only the first line of a failed or repeated answer — a partial one naming its own `FULL_RESULT_SOURCE_JSON` or `FULL_RESULT_SOURCE_UNAVAILABLE`, with a call's recorded manifest named only when it has one — claiming results "in full" or an unconditional manifest overstated a cognitive artifact. Cut detection reads the ONE shared marker list (`artifacts.SANITIZER_OMISSION_MARKERS`), because a width test over already-sanitized args measured the widest argument in the task as a small one and retained nothing at all, while a hand-rolled subset missed the `_repr` and `_error` shapes whose arguments survive only in the call blob. Unavailable source retention is disclosed, error details group by full redacted content before display clipping, and post-task synthesis — the reflection, its Pattern Register update and the episodic summary — thinks at the owner's Task / Chat effort (`settings_scales.resolve_effort("task")`), never a literal. Admission to the Pattern Register is typed, not a word scan: it opens on a call the loop recorded as errored (its stamped `tool_result_code`, or the recorded status for a legacy row), on a producer fact naming a failure the ok status cannot carry (a preserved commit whose post-commit tests failed publishes `post_commit_tests`), on typed codes already stored with an entry, or on a genuinely FAILED child — a cancelled, soft-landed best-effort or degraded child is not a failure. Those failed-child classes reach the root through the child evidence the synthesis walk already collects and make the run error-bearing for both the trigger and the prompt's error details: children do not reflect, so a short clean root that delegated the work is the only place its child's failure can be learned from at all. Deliberately not "a reason code exists", which would open the register on every terminal. +Typed root post-task triggers decide whether a run warrants Experience Review. `reflection.generate_reflection` sends the Light route one open prompt with the EXACT initial text (never a prefix) and its host-recorded `task_inputs.run_origin` beside it (provenance, never by itself the accepted requirement) plus tool-use, error, review and child projections and the same frozen non-final cost snapshot the facts row carries; it runs outside the tool loop, records its own usage, and its failure never erases the delivered result or changes a review verdict. Its execution trace is the ALL-CALLS listing (`build_trace_summary(all_calls=True)`): every call in order with every argument, identical consecutive calls folded into one `×N` row with their rounds, the first line of a failed or repeated call's result, and one header count of rounds whose every call was non-ok — no positional window and no literal cut, because the consolidation seam fits the call to the Light route whenever that route's window is known (an unknown window sends the prompt unchecked — the accepted residual of enlarging it); the STORED `trace_summary` (task card, parents, children) stays the bounded two-argument preview. The trace row carries the `round_id` of the model round that issued the call (absent when unknown), when the listing really cut an argument value or a failed/repeated call's answer, the redacted per-call record is retained through `retain_memory_source` and named in the prompt as OPTIONAL reading (never a required source), and its claim states the STORED bounds on both axes, not completeness: an argument already passed `sanitize_tool_args_for_log` (an oversized value carries a marker with its length and sha) and a result is the stored actor-visible cap — more than the listing, which shows only the first line of a failed or repeated answer — a partial one naming its own `FULL_RESULT_SOURCE_JSON` or `FULL_RESULT_SOURCE_UNAVAILABLE`, with a call's recorded manifest named only when it has one — claiming results "in full" or an unconditional manifest overstated a cognitive artifact. Cut detection reads the ONE shared marker list (`artifacts.SANITIZER_OMISSION_MARKERS`), because a width test over already-sanitized args measured the widest argument in the task as a small one and retained nothing at all, while a hand-rolled subset missed the `_repr` and `_error` shapes whose arguments survive only in the call blob. Unavailable source retention is disclosed, error details group by full redacted content before display clipping, and post-task synthesis — the reflection and its Pattern Register update — thinks at the owner's Task / Chat effort (`settings_scales.resolve_effort("task")`), never a literal. Admission to the Pattern Register is typed, not a word scan: it opens on a call the loop recorded as errored (its stamped `tool_result_code`, or the recorded status for a legacy row), on a producer fact naming a failure the ok status cannot carry (a preserved commit whose post-commit tests failed publishes `post_commit_tests`), on typed codes already stored with an entry, or on a genuinely FAILED child — a cancelled, soft-landed best-effort or degraded child is not a failure. Those failed-child classes reach the root through the child evidence the synthesis walk already collects and make the run error-bearing for both the trigger and the prompt's error details: children do not reflect, so a short clean root that delegated the work is the only place its child's failure can be learned from at all. Deliberately not "a reason code exists", which would open the register on every terminal. A reflection lands where it durably belongs: a non-project root appends the full entry to the canonical `logs/task_reflections.jsonl`; a project-scoped root appends the full entry to its project drive and the canonical log receives only a bounded pointer row — full project text never enters the canonical log, which feeds future global context. A project-bound task's context includes a bounded labeled tail of its own project's reflections; the headless mirror drive of a split root is never the reflection home, and the Pattern Register update stays canonical in both cases. Every entry carries task identity, evidence, lessons, backlog candidates, and validated memory actions. `MEMORY_ACTIONS_JSON` permits only `scratchpad_append`, `knowledge_write`, and `identity_update_candidate`, at bounded count and size. `apply_memory_actions` routes accepted actions through provenance-preserving memory and knowledge APIs. An `identity_update_candidate` is recorded in the scratchpad for review and is never auto-written to `identity.md`. For a project-scoped task, reflection applies knowledge actions only (the project store by default, explicit global allowed); its scratchpad and identity-candidate actions are skipped because this automatic Light pass lacks the conversation's full view — the conversation writes identity and scratchpad from any room through its own tools. Reflection may propose a future campaign or backlog item, but it cannot enqueue, review, commit, or enable one. @@ -535,7 +535,7 @@ The advisory rows a reflection or summary reads are ATTRIBUTED. Advisory runs ar Only roots synthesize; `root_phase_checkpoint` makes paid synthesis at-most-once across restart, while children contribute evidence. Durable-result persistence owes the answer as `final::` in `supervisor/terminal_delivery.py`'s bounded outbox (§5; normal/cancel/reap). `send_message` delivers immediately; the retained buffered copy shares its ID for durable dedupe. Replays use bounded backoff. Exhaustion/eviction preserves full text on disk, emits `terminal_delivery_exhausted` and a chat notice; external delivery remains at-least-once. Buffered `task_done` stays last to retain the slot/child drive during synthesis; a hung-synthesis reap need not lose the delivered answer. Project roots keep early answers in Project. Their canonical row and deferred Main mirror use `terminal_projection.settle_terminal_projection` via task-done/checkpoint/startup/maintenance; §3 "Main rows and host-stamped card rows" owns readiness, retirement and limits. -Synthesis receives a sealed final package from the durable result — the submitted final text, its artifact manifest and completion_observations. Full redacted action observations live in the canonical artifact store (`task.budget_drive_root or drive_root`), in the write-once `source_handles/context_checkpoints` store with verified `task_source` refs, before compact publication and outside deliverables and inferred readiness; their native reader `get_task_result(include_completion_source=true)` returns complete length/hash first, then explicit `source_start_char`/`source_end_char` ranges (`artifacts.text_source_range_projection`, the shared work-order range contract), with bytes, kind, path containment and SHA checked before any excerpt. Packet-only summary/reflection receive per-send-tool counts, each family's latest recorded return, and task-related skill readiness with coverage; full-source references are for later readers, not evidence the synthesizer has read. Positive observed facts correct error-trace impressions, while tool success does not prove owner receipt, empty material does not prove absence, and skill readiness does not attribute an owner's action to the task. Before context cleanup, `agent_task_pipeline.emit_task_results` also freezes `review_evidence.task_inputs` through `post_task_synthesis.capture_task_inputs`: `run_origin`, the existing task-local owner corpus, intact question/answer provenance and the canonical split-root verification-receipt union. Summary and reflection receive the same complete redacted content through `reflection.task_inputs_prompt_section`, separate from bounded trace/review excerpts. A zero return code is positive evidence; an unrelated later pass cannot resolve another check's failure. Peer proposals stay attributed, and unavailable input is not evidence that approval or verification never existed. Recovery uses these stored observations and inputs, not a later conversation. Summary trace pointers name the task and existing archive-aware reader (`ouroboros tasks watch --jsonl`), not guessed flat log paths. +Synthesis receives a sealed final package from the durable result — the submitted final text, its artifact manifest and completion_observations. Full redacted action observations live in the canonical artifact store (`task.budget_drive_root or drive_root`), in the write-once `source_handles/context_checkpoints` store with verified `task_source` refs, before compact publication and outside deliverables and inferred readiness; their native reader `get_task_result(include_completion_source=true)` returns complete length/hash first, then explicit `source_start_char`/`source_end_char` ranges (`artifacts.text_source_range_projection`, the shared work-order range contract), with bytes, kind, path containment and SHA checked before any excerpt. Packet-only reflection receives per-send-tool counts, each family's latest recorded return, and task-related skill readiness with coverage; full-source references are for later readers, not evidence the synthesizer has read. Positive observed facts correct error-trace impressions, while tool success does not prove owner receipt, empty material does not prove absence, and skill readiness does not attribute an owner's action to the task. Before context cleanup, `agent_task_pipeline.emit_task_results` also freezes `review_evidence.task_inputs` through `post_task_synthesis.capture_task_inputs`: `run_origin`, the existing task-local owner corpus, intact question/answer provenance and the canonical split-root verification-receipt union. Reflection receives the same complete redacted content through `reflection.task_inputs_prompt_section`, separate from bounded trace/review excerpts. A zero return code is positive evidence; an unrelated later pass cannot resolve another check's failure. Peer proposals stay attributed, and unavailable input is not evidence that approval or verification never existed. Recovery uses these stored observations and inputs, not a later conversation. A free `host_task_facts` row precedes paid stages: no model call or narrative; its metrics, routing and cost serve history. Pooled workers retain their slot until root post-task synthesis settles, for API-only and subscription tasks alike; early final-answer delivery keeps the response independent from that queue timing. Ordinary native post-work, including an inline Presence turn after its durable result is returned to the adapter, stays on its registered actor thread without a pooled worker slot: its `TaskModelWait` owner remains available through `POST_TASK_SYNTHESIS_INFLIGHT` after ordinary dialogue admission closes, detached server post-work binds its own live owner in the same registry, and the task mailbox stays available until the terminal post-task checkpoint. An open phase remains finalizing rather than appearing completed. Typed quota/auth waits resume only the unsettled call, stop or unknown outcomes degrade the phase without repeating finished stages, and restart recovery degrades an indeterminate `running` phase rather than replaying a possibly paid request. diff --git a/docs/architecture/11-frozen-contracts-v1.md b/docs/architecture/11-frozen-contracts-v1.md index 51e65a950..9a27d9baf 100644 --- a/docs/architecture/11-frozen-contracts-v1.md +++ b/docs/architecture/11-frozen-contracts-v1.md @@ -28,7 +28,7 @@ This chapter owns the ABI promise: which typed shapes and their parsing, normali | `ChatOutbound.initiator` — additive origin label of a self-initiated turn (`"consciousness"` on every frame and chat/progress/summary row of a consciousness wake-up and of the roots it starts; absent on an owner's turn); stamped by the turn's own event queue and the agent's frame meta, persisted by `log_chat`/the authored summary row/the task result, replayed by history on each row | `ouroboros/gateway/contracts.py`, `supervisor/log_addressing.py`, `ouroboros/subagent_messages.py`, `supervisor/message_bus.py`, `ouroboros/gateway/history.py`, `web/modules/api_types.js` | `tests/test_consciousness_initiator_label.py`, `tests/test_consciousness_wake_lane.py`, `web/tests/consciousness_label.test.js` | | `project_thread` stamp on all seven outbound frame types, stamped at the message-bus broadcast choke; a stamped frame is never adopted by Main (`chat_activity.mainThreadAccepts`) | `supervisor/message_bus.py`, `ouroboros/projects_registry.py` | `tests/test_message_bus.py`, `web/tests/chat_thread_routing.test.js` | | Media/link envelopes — media `task_id`/`size_bytes`/`download_url`; `LinkAction {label,url}` with at most twelve absolute HTTP(S) actions; `links` in `WS_MESSAGE_TYPES`; `chat.links` host topic | `ouroboros/gateway/contracts.py`, `ouroboros/tools/core.py`, `ouroboros/event_bus.py` | `tests/test_contracts.py` | -| Owner quiz ABI — `QuizOption {label, detail?, recommended?}` (the asker marks its recommendation on that option; the web card badges it, Telegram stars its button, the durable block keeps `recommended_index`), `QuizOutbound` (quiz_id, question, options, stake, `assumption` (required for optional clarification), additive `wait_for_answer` for a live pooled or ordinary-conversation root that must wait, lifecycle state open/answered/expired_terminal/superseded), separate `QuizStateOutbound` discriminator, `chat.quiz` host topic; the producer is the one escalation verb `escalate(question, options, stake, assumption, wait_for_answer=False, max_wait_minutes=None)` (the bound applies to a required wait only, never past the task's own deadline: the wait resumes with a system notice and the card stays open; named on an optional question it takes the omitted path and the asker's receipt says so) — a ROOT asks the owner, a SUBAGENT delivers a typed frame to its nearest LIVE ancestor, which answers via `forward_to_worker` or escalates verbatim, so the owner sees only what no ancestor answered; answers arrive through the ONE ingress `POST /api/decisions` (family ids `quiz:{task_id}:{quiz_id}`, `routing:{client_message_id}:{routing_token}`; `interaction:` reserved), request-id idempotent, first answer wins, validated against the STORED options; `option_index` is optional for the quiz family alone — a comment-only answer writes NO `answered_index`, because a stored 0 would replay as "chose the first option"; injected as the typed `KIND_QUIZ_ANSWER` mailbox control and broadcast as `quiz_state` (carrying the recorded `comment` when the owner answered in their own words, so the live card shows `Owner's answer:` exactly as replay does); expiry is structural only (the task-done seam flips open quizzes to `expired_terminal`, and the SAME reconcile closes the paired `owner_wait` so a terminal task never projects `quiz=expired_terminal` beside `owner_wait=waiting`; `owner_wait.set_owner_wait`'s refusal to continue waiting on a terminal result is preserved, not caught); a LATE answer to such an expired card is nevertheless ACCEPTED at the same ingress (the projection records `answered_after_terminal`) and, because no mailbox will ever be drained, is delivered as the owner's OWN message into the card's chat through the named ingress `supervisor.message_bus.accept_local_message`, idempotent on `client_message_id = quiz_late_answer::`, its provenance in the message's own `late_answer` metadata rather than a substituted `client_surface`; the 2xx says `forwarded` so no surface claims a delivery that did not happen, and 409 is left for what is genuinely settled (an already answered card, a non-root addressee). History replay merges the projection state | `ouroboros/gateway/contracts.py`, `ouroboros/gateway/task_decision.py`, `ouroboros/owner_quiz.py`, `ouroboros/tools/core.py` | `tests/test_gateway_parity.py`, `tests/test_quiz_display.py`, `tests/test_quiz_answer.py`, `web/tests/chat_decision.test.js` | +| Owner quiz ABI — `QuizOption {label, detail?, recommended?}` (the asker marks its recommendation on that option; the web card badges it, Telegram stars its button, the durable block keeps `recommended_index`), `QuizOutbound` (quiz_id, question, options (0–6; empty is an open question), stake, `assumption` (required for optional clarification), additive `wait_for_answer` for a live pooled or ordinary-conversation root that must wait, lifecycle state open/answered/expired_terminal/superseded), separate `QuizStateOutbound` discriminator, `chat.quiz` host topic; the producer is the one escalation verb `escalate(question, options, stake, assumption, wait_for_answer=False, max_wait_minutes=None)` (the bound applies to a required wait only, never past the task's own deadline: the wait resumes with a system notice and the card stays open; named on an optional question it takes the omitted path and the asker's receipt says so) — a ROOT asks the owner, a SUBAGENT delivers a typed frame to its nearest LIVE ancestor, which answers via `forward_to_worker` or escalates verbatim, so the owner sees only what no ancestor answered; answers arrive through the ONE ingress `POST /api/decisions` (family ids `quiz:{task_id}:{quiz_id}`, `routing:{client_message_id}:{routing_token}`; `interaction:` reserved), request-id idempotent, first answer wins, validated against the STORED options; `option_index` is optional for the quiz family alone — a comment-only answer writes NO `answered_index`, because a stored 0 would replay as "chose the first option"; injected as the typed `KIND_QUIZ_ANSWER` mailbox control and broadcast as `quiz_state` (carrying the recorded `comment` when the owner answered in their own words, so the live card shows `Owner's answer:` exactly as replay does); expiry is structural only (the task-done seam flips open quizzes to `expired_terminal`, and the SAME reconcile closes the paired `owner_wait` so a terminal task never projects `quiz=expired_terminal` beside `owner_wait=waiting`; `owner_wait.set_owner_wait`'s refusal to continue waiting on a terminal result is preserved, not caught); a LATE answer to such an expired card is nevertheless ACCEPTED at the same ingress (the projection records `answered_after_terminal`) and, because no mailbox will ever be drained, is delivered as the owner's OWN message into the card's chat through the named ingress `supervisor.message_bus.accept_local_message`, idempotent on `client_message_id = quiz_late_answer::`, its provenance in the message's own `late_answer` metadata rather than a substituted `client_surface`; the 2xx says `forwarded` so no surface claims a delivery that did not happen, and 409 is left for what is genuinely settled (an already answered card, a non-root addressee). History replay merges the projection state | `ouroboros/gateway/contracts.py`, `ouroboros/gateway/task_decision.py`, `ouroboros/owner_quiz.py`, `ouroboros/tools/core.py` | `tests/test_gateway_parity.py`, `tests/test_quiz_display.py`, `tests/test_quiz_answer.py`, `web/tests/chat_decision.test.js` | | Managed update ABI — preflight, `UpdateMergePlan`, pinned apply, process-local `update_progress`, `update_progress_changed` invalidation and boot-only `update_status_ready` | `ouroboros/gateway/contracts.py` | `tests/test_update_apply_routing.py` | | `ChatOutbound.review_projection` — bounded actor findings via `utils.truncate_review_artifact`, at most `MAX_PROJECTED_ACTOR_FINDINGS` rows (`review_execution_projection.py`) | `ouroboros/gateway/contracts.py` | `tests/test_review_substrate_v2.py`, `web/tests/review_truth.test.js` | | Skill preflight statuses — `preflight_failed` is fresh-only; a stale failure surfaces as `preflight_failed_stale`; absence means the caller could not know | `ouroboros/skill_review_status.py` | `tests/test_skill_preflight_repair.py`, `web/tests/skill_preflight_repair.test.js` | diff --git a/docs/inventories/FACADE_INVENTORY.md b/docs/inventories/FACADE_INVENTORY.md index 8a4f4a047..85eef8d58 100644 --- a/docs/inventories/FACADE_INVENTORY.md +++ b/docs/inventories/FACADE_INVENTORY.md @@ -2,13 +2,13 @@ AST-derived inventory of compatibility facades, regenerated by `python scripts/regenerate_inventories.py`. Do not edit. A facade row is any runtime module whose top-level `from import ...` statements carry the `noqa: F401` re-export marker — the codebase's declared "this binding exists for its binding, not for this module's own use" convention (reference FACADE_CONSUMERS method). Leaf domains come from `ouroboros/domains.toml`; a leaf outside the facade's domain is marked ✗ (that edge also appears in the manifest's pinned direction matrix). `tests/test_generated_inventories.py` pins byte-identity, so any re-export surface change must regenerate this file. -- facade modules: **63**; marked re-export bindings: **2343**; cross-domain facade→leaf pairs: **133** +- facade modules: **63**; marked re-export bindings: **2342**; cross-domain facade→leaf pairs: **133** | facade | domain | bindings | leaves | |---|---|---:|---| | `launcher.py` | D18 | 2 | `ouroboros/launcher_windows_runtime.py` (2) | | `ouroboros/agent.py` | D01 | 32 | `ouroboros/agent_dispatch.py` (15)
`ouroboros/agent_startup_checks.py` (4)
`ouroboros/config.py` (2 ✗D12)
`ouroboros/subagent_dispatch_notes.py` (4 ✗D07)
`ouroboros/subagents.py` (7 ✗D07) | -| `ouroboros/agent_task_pipeline.py` | D01 | 28 | `ouroboros/dialogue_provenance.py` (2 ✗D15)
`ouroboros/post_task_synthesis.py` (11)
`ouroboros/synthesis_cost_text.py` (5)
`ouroboros/task_finalization.py` (10) | +| `ouroboros/agent_task_pipeline.py` | D01 | 27 | `ouroboros/dialogue_provenance.py` (2 ✗D15)
`ouroboros/post_task_synthesis.py` (10)
`ouroboros/synthesis_cost_text.py` (5)
`ouroboros/task_finalization.py` (10) | | `ouroboros/config.py` | D12 | 142 | `ouroboros/model_slots.py` (17)
`ouroboros/provider_models.py` (6 ✗D02)
`ouroboros/review_model_routes.py` (10)
`ouroboros/runtime_limits.py` (67)
`ouroboros/settings_defaults.py` (19)
`ouroboros/settings_integrity.py` (4)
`ouroboros/settings_scales.py` (17)
`ouroboros/update_channels.py` (2) | | `ouroboros/context.py` | D03 | 4 | `ouroboros/context_runtime_facts.py` (4) | | `ouroboros/delegate_custody.py` | D07 | 10 | `ouroboros/delegate_custody_reconcile.py` (9)
`ouroboros/delegate_evidence.py` (1) | diff --git a/docs/inventories/FROZEN_CONTRACTS_INVENTORY.md b/docs/inventories/FROZEN_CONTRACTS_INVENTORY.md index c39b7c927..4c42907b1 100644 --- a/docs/inventories/FROZEN_CONTRACTS_INVENTORY.md +++ b/docs/inventories/FROZEN_CONTRACTS_INVENTORY.md @@ -2,7 +2,7 @@ Machine extraction of `docs/ARCHITECTURE.md` §11.1 (the frozen-ABI SSOT), regenerated by `python scripts/regenerate_inventories.py`. Do not edit — edit the owning chapter named in the Source line and regenerate; `tests/test_generated_inventories.py` pins byte-identity and the resolution invariants (a §11.1 row whose owner or anchor file disappeared from the tree = red). -Source: `docs/architecture/11-frozen-contracts-v1.md`, physical LF lines 7-40; UTF-8 SHA-256 `902c0a2935eb46a9cea18f4d0405f27e5facf7e400c7c8efcd914dafca74ffd4`. +Source: `docs/architecture/11-frozen-contracts-v1.md`, physical LF lines 7-40; UTF-8 SHA-256 `5aa09c5a7dde9832e4f796a36e0de9580bad4eb8fb68a953bc96df9ee2fb5dda`. - table rows: **29** - browser-envelope prose owners: diff --git a/ouroboros/agent.py b/ouroboros/agent.py index 7d275bc94..405c9079b 100644 --- a/ouroboros/agent.py +++ b/ouroboros/agent.py @@ -563,6 +563,7 @@ class OuroborosAgent: # passes the start-message identity to the next binding. "origin_message_ref", "origin_message_text", + "objective_author", "owner_corpus", # The complete work-order source reader needs the original typed # constraint mapping, not the normalized dataclass repr, to rebuild # the exact canonical serializer bytes during a source-range answer. diff --git a/ouroboros/agent_task_pipeline.py b/ouroboros/agent_task_pipeline.py index 2e97628b0..0534d36f8 100644 --- a/ouroboros/agent_task_pipeline.py +++ b/ouroboros/agent_task_pipeline.py @@ -171,6 +171,9 @@ def _run_post_task_processing_async( checkpoint_status = "degraded" skipped: list[str] = [] try: + # The free facts row precedes every paid stage, so neither Stop nor a + # failed paid stage costs the card its facts; it is not a stage. + _record_task_facts(env, task_snapshot, usage_snapshot, trace_snapshot, drive_logs) from ouroboros.llm import LLMClient from ouroboros.memory import Memory @@ -198,18 +201,13 @@ def _run_post_task_processing_async( on_reflection(reflection_entry, llm_client) # All late model work belongs to this one scoped worker. This keeps - # the root checkpoint non-final until consolidation, summary, - # reflection, and promotion have all stopped billing. Summary - # before reflection: chat.jsonl is more durable than best-effort - # reflection/backlog. + # the root checkpoint non-final until consolidation, reflection, + # and promotion have all stopped billing. stages: List[tuple[str, Callable[[], Any]]] = [ ("chat_consolidation", lambda: _run_chat_consolidation( env, task_memory, llm_client, task_snapshot, drive_logs)), ("scratchpad_consolidation", lambda: _run_scratchpad_consolidation( env, task_memory, llm_client)), - ("summary", lambda: _run_task_summary( - env, llm_client, task_snapshot, usage_snapshot, trace_snapshot, drive_logs, - review_evidence=review_evidence_snapshot, sealed_final=sealed_snapshot)), ("reflection", lambda: result.__setitem__("reflection_entry", _run_reflection( env, llm_client, task_snapshot, usage_snapshot, trace_snapshot, review_evidence_snapshot, sealed_final=sealed_snapshot))), @@ -1303,9 +1301,8 @@ from ouroboros.post_task_synthesis import ( # noqa: E402, F401 -- intentional p _child_task_evidence, _pre_synthesis_usage_snapshot, _compact_review_projection, - _run_task_summary, + _record_task_facts, _run_chat_consolidation, _run_scratchpad_consolidation, _run_reflection, - _TASK_SUMMARY_PROMPT, ) diff --git a/ouroboros/context.py b/ouroboros/context.py index 59be6cef1..94ff14027 100644 --- a/ouroboros/context.py +++ b/ouroboros/context.py @@ -72,6 +72,11 @@ def build_user_content(task: Dict[str, Any]) -> Any: text = task.get("text", "") metadata = task.get("metadata") if isinstance(task.get("metadata"), dict) else {} + author = metadata.get("objective_author") + if isinstance(author, dict) and author.get("kind") == "task": + text = (f"[OBJECTIVE_AUTHOR] The objective below was drafted by task {author.get('task_id')}, " + "not spoken by the owner. The owner's words retain their own source. " + "[/OBJECTIVE_AUTHOR]\n\n" + str(text or "")) if metadata.get("force_plan"): source = str(metadata.get("force_plan_source") or "operator").strip() or "operator" from ouroboros.config import get_review_enforcement diff --git a/ouroboros/gateway/task_decision.py b/ouroboros/gateway/task_decision.py index 69fa96333..f7b8d20c1 100644 --- a/ouroboros/gateway/task_decision.py +++ b/ouroboros/gateway/task_decision.py @@ -119,10 +119,7 @@ def _quiz_answer_frame( f"Question was: {block.get('question')}", ] if option_index is None: - lines.append( - "The owner rejected all offered options and answered verbatim: " - f"{comment}" - ) + lines.append(f"No option was selected; the owner wrote verbatim: {comment}") else: label = str(options[option_index]) if 0 <= option_index < len(options) else "" lines.append(f"The owner chose option {option_index + 1}: {label}") @@ -238,10 +235,11 @@ def _forward_late_quiz_answer( if chat_id is None: return False, reason index = block.get("answered_index") - text = _quiz_answer_frame( + text = (f"[Late answer to a question asked by task {task_id}, which had finished]\n" + + _quiz_answer_frame( block, index if isinstance(index, int) else None, str(block.get("comment") or ""), - ) + )) client_message_id = f"quiz_late_answer:{task_id}:{quiz_id}" bridge = message_bus.get_bridge() row, rejoined = message_bus.accept_local_message( diff --git a/ouroboros/loop_acceptance_review.py b/ouroboros/loop_acceptance_review.py index 2c567d7c8..f71ebb53d 100644 --- a/ouroboros/loop_acceptance_review.py +++ b/ouroboros/loop_acceptance_review.py @@ -602,8 +602,11 @@ def _finish_cyber_acceptance(ctx: _TaskAcceptanceContext, result: Any) -> bool: ctx.tools._ctx._task_acceptance_reviewed = False # final ingress, not review, owns delivery sealing clean = not pending and task_acceptance_is_clean(result) signal = "" if pending else str(getattr(result, "aggregate_signal", "") or "") + # The submitted final is Main's act (finish); Main stated no stance toward + # the criticism, so the record carries the act and no invented disposition. author = build_author_disposition( - disposition="accepted", rationale="Main submitted this complete response for delivery; independent review remains advisory.", + disposition="", action="finish", + rationale="Main submitted this complete response for delivery; independent review remains advisory.", # Evidence assembly can fail before a review binding exists. Bind the # author's decision to its real subject without inventing a reviewed pack. subject_hash=ctx.review_binding.get("binding_hash") or delivery_subject_hash(ctx.tools._ctx, ctx.llm_trace, ctx.content), diff --git a/ouroboros/loop_messages.py b/ouroboros/loop_messages.py index 8efbf95e9..921cf460a 100644 --- a/ouroboros/loop_messages.py +++ b/ouroboros/loop_messages.py @@ -254,12 +254,24 @@ def _initialize_owner_directives(ctx: Any, messages: List[Dict[str, Any]]) -> No existing = getattr(ctx, "_owner_directives", None) if isinstance(existing, list) and existing: return + metadata = getattr(ctx, "task_metadata", None) + metadata = metadata if isinstance(metadata, dict) else {} + author = metadata.get("objective_author") + if isinstance(author, dict) and author.get("kind") == "task": + for row in metadata.get("owner_corpus") or []: + if isinstance(row, dict) and row.get("source") in { + "owner_mailbox", "owner_quiz_answer", "origin_message", "owner_corpus", "direct_incoming"}: + _loop()._record_owner_directive( + ctx, source=str(row["source"]), content=row.get("content"), + msg_id=str(row.get("msg_id") or ""), + origin={key: row[key] for key in ("source_task_id", "relayed_from_task_id") if row.get(key)}, + ) + return # The task-drafted objective is never an owner directive. for message in messages: if isinstance(message, dict) and str(message.get("role") or "") == "user": from ouroboros.dialogue_provenance import run_origin - metadata = getattr(ctx, "task_metadata", None) - stamped = run_origin({"metadata": metadata if isinstance(metadata, dict) else {}})["owner_ingress"] + stamped = run_origin({"metadata": metadata})["owner_ingress"] _loop()._record_owner_directive( ctx, source="initial_user" if stamped else "initial_text", diff --git a/ouroboros/loop_round_limits.py b/ouroboros/loop_round_limits.py index cccd84469..0d105f730 100644 --- a/ouroboros/loop_round_limits.py +++ b/ouroboros/loop_round_limits.py @@ -479,7 +479,12 @@ def _handle_forced_finalization(ctx: _RoundLimitContext, reason: str) -> Tuple[s return _handle_owner_stop_finalization(ctx, str(reason)) if reason_lines and reason_lines[0].strip() == REASON_OWNER_STOPPED_DIRECT_TURN: return _handle_direct_turn_hard_stop(ctx) - fallback = f"⚠️ Task reached {reason or 'deadline'}; finalization grace produced no answer." + from ouroboros.project_dialogue import TASK_CAUSE_PHRASES + + # The host fallback speaks the rail's owner sentence; an unknown rail stays raw. + rail = (reason_lines[0].strip() if reason_lines else "") or "deadline" + cause = TASK_CAUSE_PHRASES.get(rail, f"Task reached {rail}") + fallback = f"⚠️ {cause}; finalization grace produced no answer." prompt = ( f"[FINALIZE_NOW] The supervisor opened a finalization grace window (reason: {reason or 'deadline'}). " "The task will be stopped shortly. Produce your best final answer NOW from the verified " diff --git a/ouroboros/outcomes.py b/ouroboros/outcomes.py index 1d3dcb4d8..f2c6095d9 100644 --- a/ouroboros/outcomes.py +++ b/ouroboros/outcomes.py @@ -639,8 +639,8 @@ def _objective_axis(review: Dict[str, Any]) -> Dict[str, Any]: local_failure = (decision.get("acceptance_incident") or {}).get("status") == "failed" if (not local_failure and _decision_reason == "author_finish" and author and author.get("enforcement") == "advisory" and author.get("action", "finish") == "finish"): - return {"status": OBJECTIVE_PASS, "source": "author_acceptance", "review_status": status, - "outcome_tier": OUTCOME_TIER_SOLVED, "reason": "author_finish"} + # The author's own finish, not a grade: the host derives no tier from it. + return {"status": OBJECTIVE_PASS, "source": "author_acceptance", "review_status": status, "reason": "author_finish"} if ( str(decision.get("status") or "") == ACCEPTANCE_FINALIZED_UNACCEPTED and (_decision_reason in _ACCEPTANCE_BLOCKED_TERMINAL_REASONS diff --git a/ouroboros/owner_quiz.py b/ouroboros/owner_quiz.py index 8ecb5dfaf..3b48ef0ba 100644 --- a/ouroboros/owner_quiz.py +++ b/ouroboros/owner_quiz.py @@ -254,7 +254,8 @@ def mark_wait_ended(drive_root: Any, task_id: str, quiz_id: str) -> bool: def _mutator(quizzes: Dict[str, Dict[str, Any]]) -> Any: block = quizzes.get(str(quiz_id)) - if not isinstance(block, dict) or not block.get("wait_for_answer"): + if (not isinstance(block, dict) or block.get("state") != STATE_OPEN + or not block.get("wait_for_answer")): return _KEEP block.pop("wait_for_answer", None) block["wait_ended_at"] = utc_now_iso() diff --git a/ouroboros/owner_wait.py b/ouroboros/owner_wait.py index f036b5965..46d04ce19 100644 --- a/ouroboros/owner_wait.py +++ b/ouroboros/owner_wait.py @@ -9,11 +9,9 @@ bytes outlive their one-use resume authority in the ordinary task result. An optional bound (``escalate(max_wait_minutes=N)``) rides the checkpoint as an ABSOLUTE stamp (``wait_deadline_at``), so a planned restart resumes the same -bound instead of starting it over. Both callbacks return ``"owner_input"`` or -``"timeout"``; the existing control axes (Stop, cancel, task deadline, absolute -ceiling) are still checked FIRST, so the soft bound can never overtake them, and -a timeout introduces no new wait state — the row resumes with the additive -``resume_reason: "timeout"``. +bound instead of starting it over. Both callbacks report a typed wake cause (answer, owner input, other mail, +timeout or control). Stop, cancel, deadline and absolute ceiling are checked +first; a timeout is not an answer and creates no new wait state. """ from __future__ import annotations @@ -34,6 +32,37 @@ from ouroboros.task_results import _TRULY_TERMINAL_STATUSES, load_task_result log = logging.getLogger(__name__) +def classify_wake(entries: list[dict], quiz_id: str) -> str: + """Classify unread mailbox facts without treating a wake as an answer.""" + from ouroboros.owner_mailbox import KIND_FINALIZE_NOW, KIND_OWNER_TEXT, KIND_QUIZ_ANSWER + + if any(row.get("kind") == KIND_FINALIZE_NOW for row in entries): + return "control:finalize_now" + if any(row.get("kind") == KIND_QUIZ_ANSWER + and row.get("msg_id") == f"quiz_answer:{quiz_id}" for row in entries): + return "answer" + if any(row.get("kind", KIND_OWNER_TEXT) in {KIND_OWNER_TEXT, KIND_QUIZ_ANSWER} for row in entries): + return "owner_input" + return "mail" if entries else "unknown" + + +def _wait_entries(ctx: Any) -> list[dict]: + """Observe on a private seen-set; only the loop may ACK or deliver.""" + from ouroboros.owner_mailbox import drain_owner_entries + + return drain_owner_entries( + pathlib.Path(ctx.drive_root), ctx.task_id, + set(getattr(ctx, "_loop_mailbox_seen_ids", None) or ()), ctx.task_attempt or 1, + ) + + +def _fresh_wake(ctx: Any, quiz_id: str, outcome: str) -> str: + """An owner answer can land while a pooled wait is awaiting capacity.""" + observed = classify_wake(_wait_entries(ctx), quiz_id) + return (observed if observed in {"answer", "owner_input", "control:finalize_now"} + or (observed == "mail" and outcome == "unknown") else outcome) + + def set_owner_wait(root: Any, task_id: str, wait: dict, expected_wait_id: str | None = None) -> dict: """Update only the existing continuation projection, preserving siblings.""" @@ -254,7 +283,7 @@ def worker_owner_wait(wid: int, in_q: Any, out_q: Any, ctx: Any, # The bound's absolute instant; None = an unbounded wait. deadline = parse_deadline_ts((checkpoint or {}).get("wait_deadline_at")) parked = resume_requested = False - outcome = "owner_input" + outcome = "unknown" while True: try: command = in_q.get(timeout=1.0) @@ -266,6 +295,11 @@ def worker_owner_wait(wid: int, in_q: Any, out_q: Any, ctx: Any, if phase == "parked": parked = True elif phase == "resume_granted": + outcome = _fresh_wake(ctx, str(checkpoint.get("quiz_id") or ""), outcome) + root = pathlib.Path(ctx.budget_drive_root or ctx.drive_root) + wait = (load_task_result(root, ctx.task_id, strict=True) or {}).get("owner_wait") or {} + if wait.get("wait_id") == checkpoint["wait_id"] and wait.get("state") == "resumed": + set_owner_wait(root, ctx.task_id, {**wait, "resume_reason": outcome}, wait["wait_id"]) return outcome elif phase == "refused": raise RuntimeError(str(command.get("reason") or "owner wait refused")) @@ -273,7 +307,8 @@ def worker_owner_wait(wid: int, in_q: Any, out_q: Any, ctx: Any, if peek.pending( pathlib.Path(ctx.drive_root), ctx.task_id, set(getattr(ctx, "_loop_mailbox_seen_ids", set())), ctx.task_attempt or 1): - out_q.put({**identity, "phase": "resume"}) + outcome = classify_wake(_wait_entries(ctx), str(checkpoint.get("quiz_id") or "")) + out_q.put({**identity, "phase": "resume", "resume_reason": outcome}) resume_requested = True elif deadline is not None and utc_now() >= deadline: outcome = "timeout" @@ -300,7 +335,7 @@ def direct_owner_wait(ctx: Any, checkpoint: dict) -> str: wait = set_owner_wait(root, ctx.task_id, {**checkpoint, "state": "waiting"}) peek = OwnerMailboxPeek() deadline = parse_deadline_ts((checkpoint or {}).get("wait_deadline_at")) # None = unbounded - outcome = "owner_input" + outcome = "unknown" while not control.control_reason() and not peek.pending( pathlib.Path(ctx.drive_root), ctx.task_id, set(getattr(ctx, "_loop_mailbox_seen_ids", set())), ctx.task_attempt or 1): @@ -308,11 +343,12 @@ def direct_owner_wait(ctx: Any, checkpoint: dict) -> str: outcome = "timeout" break time.sleep(1.0) + control_reason = control.control_reason() + outcome = (f"control:{control_reason}" if control_reason else + _fresh_wake(ctx, str(checkpoint.get("quiz_id") or ""), outcome)) set_owner_wait(root, ctx.task_id, - {**wait, "state": "resumed", - **({"resume_reason": outcome} if outcome == "timeout" else {})}, - wait["wait_id"]) - if outcome == "timeout": + {**wait, "state": "resumed", "resume_reason": outcome}, wait["wait_id"]) + if not outcome.startswith("control:"): announce_wait_ended(root, ctx.task_id, str(checkpoint.get("quiz_id") or ""), int(getattr(ctx, "current_chat_id", 0) or 0)) return outcome @@ -327,9 +363,12 @@ def announce_wait_ended(root: Any, task_id: str, quiz_id: str, chat_id: int) -> try: from ouroboros.owner_quiz import mark_wait_ended - mark_wait_ended(root, task_id, quiz_id) + changed = mark_wait_ended(root, task_id, quiz_id) except Exception: log.debug("owner-wait end not recorded on quiz %s", quiz_id, exc_info=True) + return + if not changed: # an answered card must never be broadcast as open + return try: from supervisor.message_bus import get_bridge @@ -377,10 +416,24 @@ def wait_after_tools(ctx: Any, messages: list, trace: dict, usage: dict, checkpoint = checkpoint_owner_wait(ctx, messages, trace, usage, round_idx, tool_schemas, seen, review_binding=review_binding) outcome = callback(ctx, checkpoint) - if outcome == "timeout" and not review_binding: - # The acceptance-review park shares this seam and must never receive a - # quiz-timeout notice. - messages.append(owner_wait_timeout_notice(ctx, checkpoint)) + if outcome in {"timeout", "mail", "unknown"} and not review_binding: + # The acceptance park is not an owner question. An answer may have + # landed after the wake but before this round: never claim silence then. + from ouroboros.owner_quiz import quiz_states + + root = pathlib.Path(ctx.budget_drive_root or ctx.drive_root) + block = quiz_states(root, ctx.task_id).get(str(checkpoint.get("quiz_id") or "")) or {} + if block.get("state") != "answered": + if outcome == "timeout": + messages.append(owner_wait_timeout_notice(ctx, checkpoint)) + else: + entries = _wait_entries(ctx) + sources = ", ".join(sorted({str(row.get("source_task_id") or row.get("kind") or "mail") + for row in entries})) or "unconfirmed input" + messages.append({"role": "user", "content": ( + f"[SYSTEM NOTICE]\nThe wait on question {checkpoint.get('quiz_id')} ended " + f"after {sources}, not a confirmed owner answer. The question remains open " + "and answerable. Continue by your judgment; ask again only if you must wait.")}) ctx._owner_wait_requested = "" ctx._owner_wait_deadline_at = "" ctx._owner_wait_max_minutes = 0 diff --git a/ouroboros/post_task_synthesis.py b/ouroboros/post_task_synthesis.py index 7f7d68e80..93c3b01b0 100644 --- a/ouroboros/post_task_synthesis.py +++ b/ouroboros/post_task_synthesis.py @@ -2,7 +2,7 @@ The LLM-heavy best-effort memory work the post-task orchestrator (``agent_task_pipeline._run_post_task_processing_async``) dispatches after a -task ends: the tool-trace summary, the episodic task summary, chat/scratchpad +task ends: the tool-trace summary, the free host facts row, chat/scratchpad consolidation, the execution reflection with its child-task evidence, the durable improvement backlog and reflection memory actions, plus the shared pre-synthesis usage snapshot and the compact review projection those prompts @@ -21,7 +21,7 @@ from ouroboros.dialogue_provenance import presence_provenance_fields from ouroboros.llm_claudexor import propagate_model_error from ouroboros.outcomes import normalize_outcome_axes from ouroboros.subagent_messages import initiator_meta -from ouroboros.synthesis_cost_text import _summary_row_cost_fields, _synthesis_cost_text, _synthesis_cost_usd, _synthesis_usage_snapshot_text +from ouroboros.synthesis_cost_text import _summary_row_cost_fields, _synthesis_cost_usd, _synthesis_usage_snapshot_text from ouroboros.task_finalization import sealed_final_prompt_section from ouroboros.tool_capabilities import routing_action_for_tool from ouroboros.utils import append_jsonl, truncate_review_artifact as _truncate_with_notice, utc_now_iso @@ -502,109 +502,46 @@ def _compact_review_projection(llm_trace: Dict[str, Any]) -> Dict[str, Any]: return {"panels": []} -def _run_task_summary(env, llm, task, usage, llm_trace, drive_logs, review_evidence=None, - sealed_final=None): - """Generate a detailed task summary and inject it into chat.jsonl.""" +def _record_task_facts(env: Any, task: Dict[str, Any], usage: Dict[str, Any], + llm_trace: Dict[str, Any], drive_logs: pathlib.Path) -> None: + """Append the task's free host facts row to chat.jsonl: no model call, no prose. + + The owner removed the paid task narrative (TZ-2 decision 2=A). This row keeps + the facts its readers take from a ``task_summary`` row when the result file is + gone: the direct-turn fact, origin label, typed routing action, tool metrics, + rounds, cost and review projection. Its kind is ``host_task_facts``, never + ``authored_root_summary``, and it persists no continuation narrative, so a + Main continuation sees the typed ``continuation_narrative_unavailable`` gap. + Text stays empty: existing Project and result references own navigation. + """ + task_id = str(task.get("id") or "unknown") try: - from ouroboros.project_dialogue import append_authored_task_summary, completion_status_label, outcome_phase - from ouroboros.projects_registry import project_thread_note_for_task - from ouroboros.consolidator import _consolidation_route - from ouroboros.settings_scales import resolve_effort - task_id = str(task.get("id") or "unknown") + from ouroboros.project_dialogue import append_canonical_task_summary, completion_status_label, outcome_phase + canonical_root = pathlib.Path(task.get("budget_drive_root") or drive_logs.parent) - summary_id = f"task-narrative:{task_id}" - tool_metrics = task_tool_metrics(llm_trace) - n_tool_calls = tool_metrics["tool_calls"] - rounds = None if usage.get("loop_evidence_unavailable") else int(usage.get("rounds") or 0) - round_text = "round count unknown" if rounds is None else f"{rounds}r" - cost_text = _synthesis_cost_text(usage) - outcome_axes = normalize_outcome_axes(usage) - reason_code = str(usage.get("reason_code") or "") - review_projection = _compact_review_projection(llm_trace) - presence_fields = presence_provenance_fields(task) result_root = pathlib.Path(getattr(env, "drive_root", canonical_root)) stored_result = _atp().load_task_result(result_root, task_id) or {} - result_ref = {"kind": "task_result", "task_id": task_id, "reader": "get_task_result"} - - def _append_summary(value: str) -> None: - row = { - "ts": utc_now_iso(), "direction": "system", "type": "task_summary", - "summary_kind": "authored_root_summary", "summary_id": summary_id, - "task_id": task_id, "parent_task_id": str(task.get("parent_task_id") or ""), "root_task_id": str(task.get("root_task_id") or task_id), - "project_id": str(task.get("project_id") or ""), "chat_id": int(task.get("chat_id") or 0), "delegation_role": str(task.get("delegation_role") or ""), "role": str(task.get("role") or ""), - "status": str(stored_result.get("status") or "completed"), "outcome": completion_status_label(stored_result, usage), "outcome_phase": outcome_phase(stored_result, usage), - "outcome_final": False, "outcome_authority": "pre_finalization_narrative_context", - # The chat block reads its chrome, the addressing fact and the - # origin label from this row when the task result has been pruned. - "_is_direct_chat": bool(task.get("_is_direct_chat")), **initiator_meta(task), - **({"typed_routing_action": str(usage["typed_routing_action"])} if usage.get("typed_routing_action") else {}), - "text": value, **tool_metrics, "rounds": rounds, "outcome_axes": outcome_axes, "reason_code": reason_code, - "result_ref": result_ref, "source_coverage": {"task_result": result_ref}, **_summary_row_cost_fields(usage), **presence_fields, - **({"review_projection": review_projection} if review_projection.get("panels") else {}), - } - append_authored_task_summary( - canonical_root, result_root, row, status=str(stored_result.get("status") or ""), - ) - # Skip LLM summary for trivial tasks. - if n_tool_calls in (None, 0) and (rounds is None or rounds <= 1): - goal = _truncate_with_notice(task.get("text", ""), 200) - summary_text = ( - f"Task {task_id} ({task.get('type', 'user')}): " - f"{goal}. {round_text}, {cost_text}." + project_thread_note_for_task(task) - ) - _append_summary(summary_text) - return - - summary_model, summary_use_local = _consolidation_route() - goal = _truncate_with_notice(task.get("text", ""), 500) - trace = build_trace_summary(llm_trace) - try: - from ouroboros.review_evidence import format_review_evidence_for_prompt - review_section = format_review_evidence_for_prompt(review_evidence or {}, max_chars=8000, acceptance_panels=review_projection.get("panels")) - except Exception: - review_section = "(review evidence unavailable)" - from ouroboros.reflection import task_inputs_prompt_section - - prompt = _TASK_SUMMARY_PROMPT.format( - task_id=task_id, goal=goal or "(no goal text)", - task_type=task.get("type", "user"), rounds="unknown" if rounds is None else rounds, - cost_text=cost_text, - usage_snapshot=_synthesis_usage_snapshot_text(usage), - sealed_final=sealed_final_prompt_section(sealed_final), - task_inputs=task_inputs_prompt_section(review_evidence), - trace_summary=trace, - review_evidence=review_section, - ) - try: - from ouroboros.llm_observability import chat_observed - - msg, _usage = chat_observed(llm, drive_root=canonical_root, task_id=task_id, - call_type="task_summary", messages=[{"role": "user", "content": prompt}], - model=summary_model, - model_role="light", - reasoning_effort=resolve_effort("task"), # the owner's Task / Chat level: one SSOT, no literal - max_tokens=16384, - use_local=summary_use_local) - summary_text = (msg.get("content") or "").strip() - if _usage.get("cost"): - try: - from supervisor.state import update_budget_from_usage - update_budget_from_usage(_usage) - except Exception: - pass - except Exception as exc: - propagate_model_error(exc) - log.warning("Task summary LLM call failed, using fallback", exc_info=True) - summary_text = ( - f"Task {task_id} ({task.get('type', 'user')}): " - f"{_truncate_with_notice(goal, 200)}. {round_text}, {cost_text}." - ) - if summary_text: - summary_text += project_thread_note_for_task(task) - _append_summary(summary_text) - except Exception as exc: - propagate_model_error(exc) - log.debug("Task summary generation failed (non-critical)", exc_info=True) + review_projection = _compact_review_projection(llm_trace) + append_canonical_task_summary(canonical_root, { + "ts": utc_now_iso(), "direction": "system", "type": "task_summary", + "summary_kind": "host_task_facts", "summary_id": f"task-facts:{task_id}", + "task_id": task_id, "parent_task_id": str(task.get("parent_task_id") or ""), "root_task_id": str(task.get("root_task_id") or task_id), + "project_id": str(task.get("project_id") or ""), "chat_id": int(task.get("chat_id") or 0), "delegation_role": str(task.get("delegation_role") or ""), "role": str(task.get("role") or ""), + "status": str(stored_result.get("status") or "completed"), "outcome": completion_status_label(stored_result, usage), "outcome_phase": outcome_phase(stored_result, usage), + "outcome_final": False, "outcome_authority": "pre_finalization_host_facts", + # The chat block reads its chrome, the addressing fact and the + # origin label from this row when the task result has been pruned. + "_is_direct_chat": bool(task.get("_is_direct_chat")), **initiator_meta(task), + **({"typed_routing_action": str(usage["typed_routing_action"])} if usage.get("typed_routing_action") else {}), + "text": "", **task_tool_metrics(llm_trace), + "rounds": None if usage.get("loop_evidence_unavailable") else int(usage.get("rounds") or 0), + "outcome_axes": normalize_outcome_axes(usage), "reason_code": str(usage.get("reason_code") or ""), + "result_ref": {"kind": "task_result", "task_id": task_id, "reader": "get_task_result"}, + **_summary_row_cost_fields(usage), **presence_provenance_fields(task), + **({"review_projection": review_projection} if review_projection.get("panels") else {}), + }) + except Exception: + log.warning("Task facts row was not recorded for %s (non-critical)", task_id, exc_info=True) def _run_chat_consolidation(env, memory, llm, task, drive_logs): @@ -756,33 +693,3 @@ def _run_reflection(env: Any, llm: Any, task: Dict[str, Any], propagate_model_error(error) log.debug("Execution reflection setup failed", exc_info=True) return None - - -_TASK_SUMMARY_PROMPT = """\ -Summarize this completed task for Ouroboros's episodic memory. -Be specific about: what was tried, what worked, what failed, key decisions made. -Include file names, tool names, error messages when relevant. -Treat tool statuses and exit/signal facts as authoritative. Agent notes are supplementary only. -Never claim a tool succeeded when the trace shows non-zero exit, timeout, install_error, or any error status. -If structured review evidence contains critical/advisory findings or open obligations, -mention them individually with severity, item/tag identity, and whether they blocked -the commit, remained open, or were resolved. -If the task was trivial (0 tool calls and ≤1 round), keep it to 1-2 sentences and DO NOT add meta-reflection. -If the task was non-trivial, end with a short meta-reflection section: -- What friction, errors, or weak assumptions slowed the work? -- What should Ouroboros change in its own process or prompts to avoid repeating that class of mistake? -Keep the meta-reflection concrete and operational, not narrative. -End with a task-scoped trace pointer: task_id={task_id}, task-events reader -(CLI: ouroboros tasks watch {task_id} --jsonl). Do not guess flat log-file paths; -the existing task reader merges this task's retained local, project and archived events. -## Task -Initial text: {goal} -Type: {task_type} -Rounds: {rounds}, Cost: {cost_text} - -{usage_snapshot}{sealed_final}{task_inputs}## Execution trace -{trace_summary} - -## Structured review evidence -{review_evidence} -""" diff --git a/ouroboros/project_dialogue.py b/ouroboros/project_dialogue.py index f703d57cb..4395143b8 100644 --- a/ouroboros/project_dialogue.py +++ b/ouroboros/project_dialogue.py @@ -147,7 +147,7 @@ def project_question_pointer(row: Dict[str, Any], block: Any, project: Any, "text": f"{lead} in {name}", "is_progress": False, "markdown": False, # Display fields only when known: a narrower producer must never blank a complete row. **({"question": question} if question else {}), - **({"options": labels} if labels else {}), + **({"options": labels} if isinstance(quiz.get("options"), list) or isinstance(block.get("options"), list) else {}), **({"option_details": details} if details else {}), **({"stake": stake} if stake else {}), **({"assumption": assumption} if assumption else {}), @@ -779,6 +779,26 @@ TASK_CAUSE_PHRASES = { "host_child_status_suffix": "A child task had not settled when the answer was delivered", "invalid_delivery_control_after_repair": "Ouroboros's final delivery instruction could not be read even after repair, so the answer stands as delivered.", "budget_exhausted": "The task ran out of budget before it could finish cleanly", + # The other forced-finalization rails (outcomes.BEST_EFFORT_REASON_CODES and + # the keys of ACCEPTANCE_BYPASS_REASON_BY_RAIL): the loop's typed reason_code + # when a limit ended the task. Each sentence names only the limit its code + # states; whether an answer was still delivered is the status word's to say. + "round_limit": "The task hit its round limit before it could finish cleanly", + "finalization_grace": "The task hit a time limit and had to wrap up before it could finish cleanly", + "deadline_local": "The task reached its deadline before it could finish cleanly", + "context_overflow": "The task outgrew its context before it could finish cleanly", + "children_unabsorbed": "Some sub-task results were never folded in, so the task had to wrap up", + # The supervisor's timeout rails (queue_timeouts.TIMEOUT_TERMINAL_REASONS): the + # reaper's task_done reason_code, spoken on its grace toast, its kill notice + # and its salvage line through this same table, never as the code. + "absolute_ceiling": "The task reached its maximum running time", + "deadline": "The task reached its deadline", + "idle_timeout": "The task made no progress for too long", + # The reason codes outcomes.derive_loop_outcome stamps from typed terminal facts. + "provider_failure": "The model provider failed to answer, so the task could not finish", + "empty_final_text": "The task ended without a final answer", + "deep_self_review_unavailable": "The deep self-review could not run", + "deep_self_review_error": "The deep self-review stopped on an error", # #869: the provider-death rail's terminal words; the amount of retained text is # said by the notice, this clause only names why the task ended. "provider_unavailable": "The model provider stopped answering, so the task could not finish", diff --git a/ouroboros/review_records.py b/ouroboros/review_records.py index eb8a6e583..931d015df 100644 --- a/ouroboros/review_records.py +++ b/ouroboros/review_records.py @@ -21,6 +21,7 @@ from ouroboros.review_execution import ReviewRouteKind, delivery_retrieves # their existing storage and reviewer evidence; this vocabulary only makes an # author's final stance explicit and hash-bound when a review is advisory. AUTHOR_DISPOSITION_VALUES = frozenset({"accepted", "rejected", "partial", "deferred"}) +AUTHOR_ACTION_VALUES = frozenset({"finish", "stop"}) def build_author_disposition( @@ -32,6 +33,7 @@ def build_author_disposition( enforcement: str = "", source: str = "author", recorded_at: str = "", + action: str = "", ) -> Dict[str, Any]: """Build one bounded, current-subject author-finality record. @@ -39,11 +41,15 @@ def build_author_disposition( returned object in their existing plan/skill/acceptance/commit owners and continue to retain raw reviewer rows beside it. A missing hash or reason is rejected so an author finish can never look like an unbound PASS. + The author's act (``action``) is a final stance by itself: a record may + carry it with an empty disposition, so no caller has to invent a stance. """ value = str(disposition or "").strip().lower() reason = " ".join(str(rationale or "").split()).strip() subject = str(subject_hash or "").strip() - if value not in AUTHOR_DISPOSITION_VALUES: + if action and action not in AUTHOR_ACTION_VALUES: + raise ValueError("AUTHOR_DISPOSITION_INVALID: unknown action") + if value not in AUTHOR_DISPOSITION_VALUES and (value or not action): raise ValueError("AUTHOR_DISPOSITION_INVALID: unknown disposition") if not subject: raise ValueError("AUTHOR_DISPOSITION_INVALID: subject_hash is required") @@ -63,6 +69,7 @@ def build_author_disposition( "enforcement": str(enforcement or "").strip().lower(), "recorded_at": str(recorded_at), "source": str(source or "author"), + **({"action": action} if action else {}), } @@ -76,6 +83,8 @@ def validate_author_disposition( if not isinstance(record, dict): return None try: + if "action" in record and record["action"] not in AUTHOR_ACTION_VALUES: + return None normalized = build_author_disposition( disposition=record.get("disposition", ""), rationale=record.get("rationale", ""), @@ -84,13 +93,10 @@ def validate_author_disposition( enforcement=record.get("enforcement", ""), source=record.get("source", "author"), recorded_at=record.get("recorded_at", ""), + action=record.get("action", ""), ) except (TypeError, ValueError): return None - if "action" in record: - if record["action"] not in {"finish", "stop"}: - return None - normalized["action"] = record["action"] if "review_reference" in record: import json try: diff --git a/ouroboros/tools/control_routing.py b/ouroboros/tools/control_routing.py index 03aebb65e..449dff9e3 100644 --- a/ouroboros/tools/control_routing.py +++ b/ouroboros/tools/control_routing.py @@ -461,6 +461,14 @@ def _promote_chat_to_task( evt.update(consciousness_origin_metadata(metadata)) _attach_origin_from_metadata(ctx, evt) + # A model-drafted objective has a different author from its owner source. + evt["objective_author"] = {"kind": "task", "task_id": str(getattr(ctx, "task_id", "") or "")} + owner_rows = [dict(row) for row in (getattr(ctx, "_owner_directives", None) or []) + if isinstance(row, dict) and row.get("source") in { + "owner_mailbox", "owner_quiz_answer", "origin_message", "owner_corpus", "direct_incoming"}] + if evt.get("source_text") and not any(row.get("source") == "origin_message" for row in owner_rows): + owner_rows.insert(0, {"source": "origin_message", "content": evt["source_text"]}) + evt["owner_corpus"] = owner_rows predecessor_error = _attach_predecessor_authority_from_metadata( ctx, evt, predecessor_task_id, ) diff --git a/ouroboros/tools/core.py b/ouroboros/tools/core.py index fd506d3ed..42649cbd8 100644 --- a/ouroboros/tools/core.py +++ b/ouroboros/tools/core.py @@ -1449,7 +1449,7 @@ def get_tools() -> List[ToolEntry]: "name": "escalate", "description": ( "Escalate a decision up the responsibility chain instead of guessing. " - "List 2-6 real options, mark your recommendation with recommended=true on that option, " + "Offer 0-6 real options (none for an open question); mark one recommendation if useful, " "and let each option's detail name what it gains and what it costs. " "A root task asks the OWNER (a typed quiz card with option buttons); " "a subagent asks its PARENT task (a typed mailbox frame the parent " @@ -1460,7 +1460,7 @@ def get_tools() -> List[ToolEntry]: "is irreversible or costly to redo, or the choice is the owner's to make " "(spending, publishing, deleting); your judgment decides. The task then waits after the " "current tool batch without model calls; waiting questions in one batch share one wait, " - "which ends on the first incoming message." + "which ends on the first incoming message, not necessarily an owner answer. A plain-text clarification ends this turn; a waited question keeps it alive." ), "parameters": {"type": "object", "properties": { "question": {"type": "string", "description": "The decision being escalated (markdown renders in chat)"}, @@ -1468,12 +1468,12 @@ def get_tools() -> List[ToolEntry]: "label": {"type": "string", "description": "Short option label (button text, max 120)"}, "detail": {"type": "string", "description": "Optional one-line consequence of this option (max 500)"}, "recommended": {"type": "boolean", "description": "True on the ONE option you recommend"}, - }, "required": ["label"]}, "description": "2-6 mutually exclusive options"}, + }, "required": ["label"]}, "description": "Optional 0-6 choices; omit for an open question answered in the owner's words"}, "stake": {"type": "string", "description": "What depends on this decision (optional, max 500)"}, "assumption": {"type": "string", "description": "For optional clarification, the assumption you continue under (max 500); may be empty for required waiting."}, "wait_for_answer": {"type": "boolean", "default": False, "description": "Live roots: wait for addressed owner input before another model round."}, "max_wait_minutes": {"type": "integer", "description": "Optional bound for wait_for_answer: resume with a system notice after N minutes if no answer arrives (the card stays open)."}, - }, "required": ["question", "options"]}, + }, "required": ["question"]}, }, _escalate), ToolEntry("forward_to_worker", { "name": "forward_to_worker", diff --git a/ouroboros/tools/core_artifacts.py b/ouroboros/tools/core_artifacts.py index 713aaf61f..f4cd020d6 100644 --- a/ouroboros/tools/core_artifacts.py +++ b/ouroboros/tools/core_artifacts.py @@ -325,10 +325,12 @@ def validate_quiz_payload( "QUIZ_QUESTION_INVALID", f"question must be 1..{_MAX_QUIZ_QUESTION_CHARS} characters.", ) - if not isinstance(options, list) or not 2 <= len(options) <= _MAX_QUIZ_OPTIONS: + if options is None: + options = [] + if not isinstance(options, list) or len(options) > _MAX_QUIZ_OPTIONS: raise QuizValidationError( "QUIZ_OPTIONS_INVALID", - f"provide 2..{_MAX_QUIZ_OPTIONS} options.", + f"provide at most {_MAX_QUIZ_OPTIONS} options.", ) cleaned: List[Dict[str, Any]] = [] for item in options: @@ -543,7 +545,8 @@ def _escalate( "under your stated assumption and record the open question " "in your result.") parent_task_id = target_id - lines = [f"ESCALATION (decision requested): {payload['question']}", "Options:"] + lines = [f"ESCALATION (decision requested): {payload['question']}"] + lines.append("Options:" if payload["options"] else "Open question — answer in your own words.") lines += [ f"{i + 1}. {row['label']}" + (f" — {row['detail']}" if row.get("detail") else "") + (" [recommended]" if row.get("recommended") else "") @@ -612,7 +615,7 @@ def _escalate( "task_id": task_id, **({"wait_for_answer": True} if wait_for_answer else {}), }) - delivered = "delivered to the owner" if mode == "live" else "queued for the owner" + delivered = "accepted for delivery" if mode == "live" else "queued for delivery" if wait_for_answer: bound = payload.get("max_wait_minutes") ctx._owner_wait_requested = quiz_id @@ -632,7 +635,8 @@ def _escalate( "continues with a notice" if bound else "the task waits after this tool batch") return (f"OK: quiz {quiz_id} {delivered}; {limit}, " "preserving its live browser and releasing active execution capacity. " - "Addressed owner text resumes your judgment; existing Stop and task deadlines still apply.") + "Any incoming mail ends the wait; only an owner answer answers the question. " + "Other mail leaves the card open. Stop and task deadlines still apply.") return (f"OK: quiz {quiz_id} {delivered}; continuing under assumption: " f"{payload['assumption']}. The answer (if any) arrives as an owner " "quiz answer in a later round; the card stays answerable after this " diff --git a/skills/telegram/lib/telegram_quiz.py b/skills/telegram/lib/telegram_quiz.py index 0cc1602b5..3b0aec21b 100644 --- a/skills/telegram/lib/telegram_quiz.py +++ b/skills/telegram/lib/telegram_quiz.py @@ -32,6 +32,7 @@ HostPost = Callable[[Any, str, Dict[str, Any]], Awaitable[Tuple[int, Dict[str, A _TEXTS = { "en": { "hint": "Tap an option, or reply to this message with your own answer.", + "hint_open": "Reply to this message with your answer.", "recorded": "✅ Answer delivered to the task.", "late_delivered": "✅ The task had already finished — your answer was delivered to the chat.", "late_recorded": "✅ Answer recorded. The task had already finished and this card has no chat to deliver it to.", @@ -43,6 +44,7 @@ _TEXTS = { }, "ru": { "hint": "Нажмите вариант или ответьте на это сообщение своим текстом.", + "hint_open": "Ответьте на это сообщение своим текстом.", "recorded": "✅ Ответ передан задаче.", "late_delivered": "✅ Задача уже завершилась — ответ доставлен в чат.", "late_recorded": "✅ Ответ записан. Задача уже завершилась, а доставлять его в чат некуда.", @@ -63,6 +65,10 @@ def hint(lang: str) -> str: return _texts(lang)["hint"] +def hint_open(lang: str) -> str: + return _texts(lang)["hint_open"] + + def mint_token(task_id: str, quiz_id: str) -> str: """Short stable token for ``callback_data`` (Telegram's 64-byte cap).""" return hashlib.sha256(f"{task_id}:{quiz_id}".encode("utf-8")).hexdigest()[:12] diff --git a/skills/telegram/plugin.py b/skills/telegram/plugin.py index 53d62d888..1f4ae93b4 100644 --- a/skills/telegram/plugin.py +++ b/skills/telegram/plugin.py @@ -1315,7 +1315,7 @@ def _make_quiz(api): labels.append(f"★ {label}" if option.get("recommended") is True else label) # Shared quiz contract cap: ouroboros.tools.core._MAX_QUIZ_OPTIONS. labels = labels[:6] - if not question or len(labels) < 2: + if not question: return task_id = str(event.get("task_id") or "").strip() quiz_id = str(event.get("quiz_id") or "").strip() @@ -1331,10 +1331,15 @@ def _make_quiz(api): token = telegram_quiz.mint_token(task_id, quiz_id) # One button per option; a reply to the card is a free-form answer. # Both reach the host's decision ingress (#472). - message_id = await client.send_message_with_inline_keyboard( - chat_id, f"{body}\n{telegram_quiz.hint(lang)}", - telegram_quiz.quiz_keyboard(token, labels), parse_mode="", - ) + if labels: + message_id = await client.send_message_with_inline_keyboard( + chat_id, f"{body}\n{telegram_quiz.hint(lang)}", + telegram_quiz.quiz_keyboard(token, labels), parse_mode="", + ) + else: + message_id = await client.send_message( + chat_id, f"{body}\n{telegram_quiz.hint_open(lang)}", parse_mode="", + ) telegram_quiz.remember_quiz(api, token, { "task_id": task_id, "quiz_id": quiz_id, "chat_id": chat_id, "message_id": int(message_id or 0), "options": labels, diff --git a/supervisor/events_chat_delivery.py b/supervisor/events_chat_delivery.py index dd5ef71c6..14ceb4a9d 100644 --- a/supervisor/events_chat_delivery.py +++ b/supervisor/events_chat_delivery.py @@ -441,7 +441,7 @@ def _handle_send_quiz(evt: Dict[str, Any], ctx: Any) -> None: try: chat_id = _delivery_chat_id(evt, ctx) options = evt.get("options") - if chat_id is None or not isinstance(options, list) or not options: + if chat_id is None or not isinstance(options, list): return if chat_id == 0: # Deliberate exception to the "0 is a real hidden session" policy: diff --git a/supervisor/queue_timeouts.py b/supervisor/queue_timeouts.py index ec62a1cdf..1d660e88a 100644 --- a/supervisor/queue_timeouts.py +++ b/supervisor/queue_timeouts.py @@ -39,6 +39,14 @@ def _queue(): log = logging.getLogger(__name__) +# The supervisor's timeout rails in priority order: the typed ``terminal_reason`` +# the reaper stamps as the task_done ``reason_code`` and the ``task_incident`` key. +# ``project_dialogue.TASK_CAUSE_PHRASES`` carries one owner sentence per member; +# the code itself never reaches a chat. +REASON_ABSOLUTE_CEILING, REASON_DEADLINE, REASON_IDLE_TIMEOUT = TIMEOUT_TERMINAL_REASONS = ( + "absolute_ceiling", "deadline", "idle_timeout", +) + def _task_deadline_ts(task: Dict[str, Any]) -> float: raw = str(task.get("deadline_at") or "").strip() @@ -279,11 +287,11 @@ def _enforce_task_timeouts_locked( continue if ceiling_reached: - terminal_reason = "absolute_ceiling" + terminal_reason = REASON_ABSOLUTE_CEILING elif deadline_reached: - terminal_reason = "deadline" + terminal_reason = REASON_DEADLINE else: - terminal_reason = "idle_timeout" + terminal_reason = REASON_IDLE_TIMEOUT finalization_requested_at = float(meta.get("finalization_requested_at") or 0.0) if finalization_requested_at <= 0 and _queue().FINALIZATION_GRACE_SEC > 0: meta["finalization_requested_at"] = now diff --git a/supervisor/task_reaper.py b/supervisor/task_reaper.py index 3254ceb64..874ba70ad 100644 --- a/supervisor/task_reaper.py +++ b/supervisor/task_reaper.py @@ -211,14 +211,14 @@ def request_finalization_grace( # would only decorate its summary. return control_msg_id try: + from ouroboros.project_dialogue import TASK_CAUSE_PHRASES from supervisor import workers as _workers_mod + cause = TASK_CAUSE_PHRASES.get(terminal_reason, terminal_reason) # an unknown rail stays raw _workers_mod.get_event_q().put({ "type": "send_message", "chat_id": chat_id, - "text": str(toast_text or "") or ( - f"⏳ Task {task_id} reached {terminal_reason}. " - "Finalize artifacts/results now; supervisor will stop the task after the grace window." - ), + "text": str(toast_text or "") or (f"⏳ Task {task_id}: {cause}. Finalize artifacts/results now; " + "supervisor will stop the task after the grace window."), "role": "system", "system_type": "finalization_notice", "format": "markdown", "is_progress": True, @@ -1132,14 +1132,14 @@ def _deliver_reap_salvage( return try: from ouroboros.observability import latest_llm_response_text, preserved_salvage_path + from ouroboros.project_dialogue import TASK_CAUSE_PHRASES from supervisor.terminal_delivery import deliver_unreviewed_salvage - salvage_text = latest_llm_response_text( - pathlib.Path(_q._task_drive_for_task(task, task_id)), task_id, - ) + salvage_text = latest_llm_response_text(pathlib.Path(_q._task_drive_for_task(task, task_id)), task_id) deliver_unreviewed_salvage( pathlib.Path(_q.DRIVE_ROOT), task, task_id, - outcome=f"stopped by {terminal_reason}", + # Reads "was stopped by the supervisor. ."; an unknown rail stays raw. + outcome=f"stopped by the supervisor. {TASK_CAUSE_PHRASES.get(terminal_reason, terminal_reason)}", salvaged_text=salvage_text, preserved_path=preserved_salvage_path(pathlib.Path(_q.DRIVE_ROOT), task_id), unreconciled_runs=list(unreconciled_runs or []), @@ -1553,19 +1553,20 @@ def reap_timed_out_task(job: Dict[str, Any]) -> None: incident_chat_id = _incident_chat_id(task, owner_chat_id, _q) if incident_chat_id is not None: try: + from ouroboros.project_dialogue import TASK_CAUSE_PHRASES + cause = TASK_CAUSE_PHRASES.get(terminal_reason, terminal_reason) # an unknown rail stays raw + headline = f"🛑 {cause}: task {task_id} killed after {int(runtime_sec)}s.\n" if requeued: send_with_budget( incident_chat_id, - f"🛑 {terminal_reason}: task {task_id} killed after {int(runtime_sec)}s.\n" - f"Worker {worker_id} restarted. Task queued for retry attempt={new_attempt}.", + headline + f"Worker {worker_id} restarted. Task queued for retry attempt={new_attempt}.", is_progress=True, task_id=task_id, progress_meta={"task_incident": "task_reaper_retry", "toast_once": incident_toast_once}, role="system", system_type="task_reaper_notice") elif retry_suppression.get("kind") == "cancel_intent": send_with_budget( incident_chat_id, - f"🛑 {terminal_reason}: task {task_id} killed after {int(runtime_sec)}s.\n" - "Its retry was suppressed because cancellation won the " + headline + "Its retry was suppressed because cancellation won the " "admission race; cancellation custody is settling the task.", is_progress=True, task_id=task_id, @@ -1578,8 +1579,7 @@ def reap_timed_out_task(job: Dict[str, Any]) -> None: stop_detail = _stop_detail(ceiling_reached, deadline_reached, orchestrator) send_with_budget( incident_chat_id, - f"🛑 {terminal_reason}: task {task_id} killed after {int(runtime_sec)}s.\n" - f"Worker {worker_id} restarted. {stop_detail}", + headline + f"Worker {worker_id} restarted. {stop_detail}", is_progress=True, task_id=task_id, progress_meta={"task_incident": "task_reaper_stopped", "toast_once": incident_toast_once}, role="system", system_type="task_reaper_notice") diff --git a/supervisor/worker_owner_wait.py b/supervisor/worker_owner_wait.py index b0029913d..daff4538e 100644 --- a/supervisor/worker_owner_wait.py +++ b/supervisor/worker_owner_wait.py @@ -170,7 +170,8 @@ def _grant_resume( log.warning("Owner-wait rollback remains unpersisted for %s", task_id, exc_info=True) raise meta.pop("owner_wait_resume_requested", None) - if str(resumed.get("resume_reason") or "") == "timeout" and str(resumed.get("quiz_id") or ""): + if (str(resumed.get("quiz_id") or "") + and not str(resumed.get("resume_reason") or "").startswith("control:")): # The bound closed and the pooled task resumed: one seam with the direct lane. from ouroboros.owner_wait import announce_wait_ended diff --git a/supervisor/worker_promotion.py b/supervisor/worker_promotion.py index 9e3d73704..66a2a9236 100644 --- a/supervisor/worker_promotion.py +++ b/supervisor/worker_promotion.py @@ -466,6 +466,8 @@ def promote_chat_to_task(evt: dict, ctx: Any) -> dict: "title": title, "suggested_name": suggested_name, "source": "promote_chat_to_task", + "objective_author": dict(evt.get("objective_author") or {}), + "owner_corpus": list(evt.get("owner_corpus") or []), "_require_unique_task_id": True, "_require_worker_pool": True, "_admission_token": admission_token, @@ -520,6 +522,8 @@ def promote_chat_to_task(evt: dict, ctx: Any) -> dict: # The door's other stamp (an owner message it never logged) rides the root # in METADATA, where run_origin reads it, the way a ref rides by value. task.setdefault("metadata", {})["origin_suppressed"] = True + if task.get("objective_author"): + task.setdefault("metadata", {})["objective_author"] = dict(task["objective_author"]) if isinstance(evt.get("predecessor_authority_source"), dict): task["predecessor_authority_source"] = dict(evt["predecessor_authority_source"]) # Owner Surface Fact: the promoting turn's sending-surface fact lands in diff --git a/tests/system_e2e/test_system_scenarios_w6.py b/tests/system_e2e/test_system_scenarios_w6.py index 569309dc0..f2ef89a93 100644 --- a/tests/system_e2e/test_system_scenarios_w6.py +++ b/tests/system_e2e/test_system_scenarios_w6.py @@ -52,8 +52,8 @@ WHAT S26 ASSERTS, in order: and the keepalive script was never run down; the DURABLE artifacts show no post-stop paid work either — the tool-bearing gate cannot see a tool-less summary/reflection call, so the pin reads the task result (no open - ``root_phase_checkpoint``), chat.jsonl (no ``authored_root_summary`` row) and - ``task_reflections.jsonl`` (no row) for the turn; + ``root_phase_checkpoint``), chat.jsonl (no paid ``authored_root_summary`` row; + the free host facts row is not paid work) and ``task_reflections.jsonl`` (no row); 5. the chat CONCLUDED (no "Working…" forever) — over the SAME /ws the SPA opens the turn announced itself (typing frame: activity_id = task id, kind ``direct_chat``, the client_message_id link), the toast frame carried @@ -282,9 +282,9 @@ def _post_task_synthesis_is_open(stored: dict) -> bool: def _post_stop_synthesis_rows(oracle: ArtifactOracle, task_id: str) -> list: - """Durable traces of the post-task worker for the turn: the authored summary - row (written even for a trivial turn, so its absence means the phase never - ran) and any reflection row.""" + """Durable traces of PAID post-task work for the turn: any reflection row and + any paid narrative row (the paid summary no longer exists, so this stays a + regression guard). The free host facts row is not paid work and may exist.""" summaries = [r for r in oracle._jsonl("logs/chat.jsonl", type_filter="task_summary") if str(r.get("task_id") or "") == task_id and r.get("summary_kind") == "authored_root_summary"] diff --git a/tests/test_acceptance_optional_control.py b/tests/test_acceptance_optional_control.py index 3f64fb1d9..694297da5 100644 --- a/tests/test_acceptance_optional_control.py +++ b/tests/test_acceptance_optional_control.py @@ -211,11 +211,20 @@ def test_cyber_explicit_nomination_keeps_its_existing_no_wait_power(full_loop, m return {"content": revised}, 0.0 monkeypatch.setattr(loop, "call_llm_with_retry", main) - result, _usage, trace = f.run() + result, usage, trace = f.run() assert result == revised and len(f.review_sends) == 1 and not f.waits assert trace["acceptance_decision"]["reason"] == "author_finish" - assert trace["acceptance_decision"]["author_disposition"]["source"] == "author_final_response" + author = trace["acceptance_decision"]["author_disposition"] + assert author["source"] == "author_final_response" + # The submitted final is Main's act: the host records "finish" and invents + # neither an "accepted" stance nor a "solved" tier (TZ-2 C4). + assert author["action"] == "finish" and author["disposition"] == "" assert trace["review_decision"]["review_pending"] is True + from ouroboros.outcomes import derive_loop_outcome + + objective = derive_loop_outcome(result, usage, trace)["outcome_axes"]["objective"] + assert (objective["status"], objective["source"], objective["reason"]) == ("pass", "author_acceptance", "author_finish") + assert "outcome_tier" not in objective @pytest.mark.parametrize("early", ["settled", "queued_wake"]) diff --git a/tests/test_agent_task_pipeline.py b/tests/test_agent_task_pipeline.py index bd76737e0..47fbea356 100644 --- a/tests/test_agent_task_pipeline.py +++ b/tests/test_agent_task_pipeline.py @@ -417,10 +417,11 @@ def test_stopped_direct_turn_pays_no_post_task_synthesis(tmp_path, monkeypatch): end through the real post-task lane. The loop's hard stop records the existing ``_skip_post_task_synthesis`` marker on the tool context; ``emit_task_results`` copies it onto the task before the root predicate - runs, so the summary/reflection worker is never dispatched and no open + runs, so the reflection worker is never dispatched and no open ``root_phase_checkpoint`` is seeded for the boot reconciler to re-pay. A positive control (the same turn, not stopped) proves the recording model - would have seen the paid summary + reflection calls.""" + would have seen the paid reflection call. Neither turn buys a narrative: + the dispatched worker writes one free host facts row, the stopped turn none.""" import ouroboros.llm as llm_mod from ouroboros.outcomes import REASON_OWNER_REQUESTED_FINALIZATION from supervisor.owner_stop import REASON_OWNER_STOPPED_DIRECT_TURN @@ -501,26 +502,28 @@ def test_stopped_direct_turn_pays_no_post_task_synthesis(tmp_path, monkeypatch): owed_rows = [row for row in pending_deliveries(root) if row.get("task_id") == "stopped1"] assert owed_rows and owed_rows[0]["progress_meta"]["task_terminal_status"] == "failed", owed_rows - def _summary_rows(task_id): + def _summary_kinds(task_id): chat_log = root / "logs" / "chat.jsonl" - rows = chat_log.read_text(encoding="utf-8").splitlines() if chat_log.exists() else [] - return [row for row in rows if "authored_root_summary" in row and task_id in row] + rows = [json.loads(line) for line in chat_log.read_text(encoding="utf-8").splitlines()] if chat_log.exists() else [] + return [row.get("summary_kind") for row in rows if row.get("type") == "task_summary" and row.get("task_id") == task_id] - assert _summary_rows("stopped1") == [] + stopped_kinds = _summary_kinds("stopped1") # no post-task worker was dispatched at all + assert "host_task_facts" not in stopped_kinds and "authored_root_summary" not in stopped_kinds _task, _events, control_calls = _turn("control1", stopped=False) - assert len(control_calls) >= 2, control_calls - assert len(_summary_rows("control1")) == 1 # the reader sees the phase when it does run + assert len(control_calls) >= 1, control_calls + control_kinds = _summary_kinds("control1") + assert control_kinds.count("host_task_facts") == 1 and "authored_root_summary" not in control_kinds # --- "Stop now" while the paid synthesis is ALREADY in flight (audit point 4, G18) --- -_STAGES = ("chat_consolidation", "scratchpad_consolidation", "summary", "reflection", "promotion") +_STAGES = ("chat_consolidation", "scratchpad_consolidation", "reflection", "promotion") def _stubbed_stages(monkeypatch, calls, *, on_first=None): - """Record the five paid stages in order; ``on_first`` runs INSIDE stage 1 - (the Stop lands after the synthesis has begun, past the entry snapshot).""" + """Record the free facts row and the four paid stages in order; ``on_first`` runs + INSIDE stage 1 (the Stop lands after the synthesis has begun, past the entry snapshot).""" import ouroboros.llm as llm_mod import ouroboros.post_task_evolution as pte @@ -536,7 +539,7 @@ def _stubbed_stages(monkeypatch, calls, *, on_first=None): monkeypatch.setattr(pipeline, "_run_chat_consolidation", _stage("chat_consolidation")) monkeypatch.setattr(pipeline, "_run_scratchpad_consolidation", _stage("scratchpad_consolidation")) - monkeypatch.setattr(pipeline, "_run_task_summary", _stage("summary")) + monkeypatch.setattr(pipeline, "_record_task_facts", _stage("facts")) monkeypatch.setattr(pipeline, "_run_reflection", _stage( "reflection", {"reflection": "x", "backlog_candidates": [], "memory_actions": []})) monkeypatch.setattr(pipeline, "_update_improvement_backlog", _stage("promotion")) @@ -561,7 +564,7 @@ def test_stop_now_during_inflight_synthesis_skips_the_remaining_paid_stages(tmp_ """The Stop lands AFTER stage 1 began (the loop has returned, the entry snapshot saw no marker): the durable immediate cancel intent every stop ingress mints — or the live task marker re-read — trips the per-stage gate, - so stages 2..5 never run, the checkpoint settles ``degraded`` and the typed + so stages 2..4 never run, the checkpoint settles ``degraded`` and the typed ``post_task_stop_reason`` NAMES the skipped stages, riding the result row and the ``task_cost_finalized`` event alike. The in-flight key is gone.""" from ouroboros.cancel_intents import STOP_POLICY_IMMEDIATE, request_cancel @@ -583,11 +586,11 @@ def test_stop_now_during_inflight_synthesis_skips_the_remaining_paid_stages(tmp_ pipeline._run_post_task_processing_async( env, task, {"rounds": 20, "cost": 0.02}, {"tool_calls": [], "reasoning_notes": []}, {}, root / "logs", blocking=True) - assert calls == ["chat_consolidation"], calls + assert calls == ["facts", "chat_consolidation"], calls checkpoint = (pipeline.load_task_result(root, task_id) or {}).get("root_phase_checkpoint") or {} assert checkpoint.get("post_task_synthesis") == "degraded", checkpoint assert checkpoint.get("post_task_stop_reason") == ( - "owner_stopped:skipped=scratchpad_consolidation,summary,reflection,promotion"), checkpoint + "owner_stopped:skipped=scratchpad_consolidation,reflection,promotion"), checkpoint finalized = _finalized_events(root, task_id) assert len(finalized) == 1 and finalized[0]["post_task_status"] == "degraded", finalized assert finalized[0]["post_task_stop_reason"] == checkpoint["post_task_stop_reason"], finalized @@ -595,9 +598,9 @@ def test_stop_now_during_inflight_synthesis_skips_the_remaining_paid_stages(tmp_ def test_no_stop_runs_every_paid_stage_and_finalizes_without_a_stop_reason(tmp_path, monkeypatch): - """Positive control for the gate: an un-stopped synthesis runs all five - stages in order and the checkpoint is byte-identical to before (``completed``, - no ``post_task_stop_reason`` anywhere).""" + """Positive control for the gate: an un-stopped synthesis records its free + facts row, then runs all four paid stages in order, and the checkpoint is + byte-identical to before (``completed``, no ``post_task_stop_reason`` anywhere).""" root, env = _synthesis_root(tmp_path) task_id = "unstopped1" task = {"id": task_id, "type": "task", "chat_id": 1, "_is_direct_chat": True, "text": "keep listing"} @@ -607,7 +610,7 @@ def test_no_stop_runs_every_paid_stage_and_finalizes_without_a_stop_reason(tmp_p pipeline._run_post_task_processing_async( env, task, {"rounds": 20, "cost": 0.02}, {"tool_calls": [], "reasoning_notes": []}, {}, root / "logs", blocking=True) - assert calls == list(_STAGES), calls + assert calls == ["facts", *_STAGES], calls checkpoint = (pipeline.load_task_result(root, task_id) or {}).get("root_phase_checkpoint") or {} assert checkpoint.get("post_task_synthesis") == "completed", checkpoint assert "post_task_stop_reason" not in checkpoint, checkpoint @@ -619,7 +622,7 @@ def test_entry_marker_still_skips_every_paid_stage_and_seeds_no_checkpoint(tmp_p """The rc.14 path is unchanged: a Stop that landed inside the loop (marker on the task at entry) pays nothing and leaves no open checkpoint for the boot reconciler — the gate trips before stage 1 and the marker keeps the - checkpoint writer's root predicate False.""" + checkpoint writer's root predicate False. The free facts row still lands.""" root, env = _synthesis_root(tmp_path) task_id = "marked1" task = {"id": task_id, "type": "task", "chat_id": 1, "_is_direct_chat": True, @@ -630,7 +633,7 @@ def test_entry_marker_still_skips_every_paid_stage_and_seeds_no_checkpoint(tmp_p pipeline._run_post_task_processing_async( env, task, {"rounds": 2, "cost": 0.01}, {"tool_calls": [], "reasoning_notes": []}, {}, root / "logs", blocking=True) - assert calls == [], calls + assert calls == ["facts"], calls stored = pipeline.load_task_result(root, task_id) or {} assert "root_phase_checkpoint" not in stored, stored assert not (root / "logs" / "events.jsonl").exists() or _finalized_events(root, task_id) == [] diff --git a/tests/test_author_finish_invents_nothing.py b/tests/test_author_finish_invents_nothing.py new file mode 100644 index 000000000..b1be31910 --- /dev/null +++ b/tests/test_author_finish_invents_nothing.py @@ -0,0 +1,98 @@ +"""TZ-2 C4: the host never writes an author's disposition for it. + +An author finish is the author's act. The host records that act; it invents no +stance ("accepted") and derives no grade ("solved") from it. Blocking review +stays mandatory, a local preparation failure stays a failure, the plan/commit +stance envelopes keep their contract, and reviewer tiers keep their meaning. +""" +from __future__ import annotations + +import pytest + +from ouroboros.outcomes import _objective_axis, normalize_outcome_axes +from ouroboros.review_records import ( + build_author_disposition, + build_author_disposition_from_mapping, + validate_author_disposition, +) + + +def _record(**overrides): + fields = {"disposition": "", "action": "finish", "rationale": "Main submitted this complete response.", + "subject_hash": "subject-1", "reviewer_signal": "FAIL", "enforcement": "advisory", + "source": "author_final_response"} + fields.update(overrides) + return build_author_disposition(**fields) + + +def _author_finish(record, *, review_status="fail", **decision): + return {"status": review_status, "acceptance_decision": { + "status": "finalized_unaccepted", "reason": "author_finish", "author_disposition": record, **decision}} + + +def test_an_act_alone_is_a_valid_author_record_with_no_stance(): + record = _record() + assert record["disposition"] == "" and record["action"] == "finish" + assert validate_author_disposition(record, subject_hash="subject-1") == record + assert validate_author_disposition(record, subject_hash="another-subject") is None # still subject-bound + + +@pytest.mark.parametrize("fields", [ + {"action": ""}, # neither a stance nor an act + {"action": "ship"}, # an unknown act + {"disposition": "maybe"}, # an unknown stance beside a known act + {"subject_hash": ""}, # an unbound act + {"rationale": " "}, # an unexplained act +]) +def test_a_record_still_needs_a_known_stance_or_act_bound_and_explained(fields): + with pytest.raises(ValueError): + _record(**fields) + stored = {"disposition": "", "action": "finish", "rationale": "r", "subject_hash": "h", **fields} + assert validate_author_disposition(stored) is None + + +def test_stance_envelopes_of_plan_and_commit_still_require_a_disposition(): + with pytest.raises(ValueError): + build_author_disposition_from_mapping({"disposition": "", "rationale": "r"}, subject_hash="h") + assert build_author_disposition_from_mapping( + {"disposition": "deferred", "rationale": "r"}, subject_hash="h")["disposition"] == "deferred" + + +@pytest.mark.parametrize("record", [ + _record(), # Cyber's submitted final: the act only + _record(disposition="partial", source="author"), # an explicit author stance + {k: v for k, v in _record(disposition="accepted").items() if k != "action"}, # a historical record +]) +def test_author_finish_is_an_author_pass_with_no_host_tier(record): + objective = _objective_axis(_author_finish(record)) + assert objective == {"status": "pass", "source": "author_acceptance", "review_status": "fail", + "reason": "author_finish"} + + +@pytest.mark.parametrize("review", [ + _author_finish(_record(enforcement="blocking")), # Blocking needs a fresh critic approval + _author_finish(_record(action="stop")), # an honest stop is not a finish + _author_finish(None), # the host needs the author's own record + _author_finish(_record(), acceptance_incident={"status": "failed", "stage": "preparation"}), +]) +def test_everything_that_is_not_an_advisory_author_finish_keeps_its_own_outcome(review): + objective = _objective_axis(review) + assert objective["source"] != "author_acceptance" + assert objective["status"] == "fail" + + +def test_reviewer_tiers_keep_their_meaning(): + clean = _objective_axis({"status": "pass", "outcome_tier": "solved", + "acceptance_decision": {"status": "accepted", "reason": "clean_pass"}}) + assert (clean["status"], clean["source"], clean["outcome_tier"]) == ("pass", "task_acceptance_review", "solved") + + +def test_normalized_history_keeps_both_new_and_old_author_finish_rows(): + review = _author_finish(_record()) + new = normalize_outcome_axes({"status": "completed", "outcome_axes": { + "objective": _objective_axis(review), "review": review}}) + assert new["objective"]["source"] == "author_acceptance" and "outcome_tier" not in new["objective"] + # A row written before this change keeps its stored words: history is not rewritten. + old_objective = {**_objective_axis(review), "outcome_tier": "solved"} + old = normalize_outcome_axes({"status": "completed", "outcome_axes": {"objective": old_objective, "review": review}}) + assert old["objective"]["status"] == "pass" and old["objective"]["outcome_tier"] == "solved" diff --git a/tests/test_autonomy_review_fixes.py b/tests/test_autonomy_review_fixes.py index ae0718794..879c1e1ef 100644 --- a/tests/test_autonomy_review_fixes.py +++ b/tests/test_autonomy_review_fixes.py @@ -192,18 +192,18 @@ def test_pre_loop_checkpoint_failure_reports_unknown_counts_without_reading_corr assert stored["owner_wait"]["source_ref"] == case.wait["source_ref"] -def test_unknown_exception_summary_does_not_invent_zero_rounds_or_buy_a_model_call(tmp_path, monkeypatch): - from ouroboros.post_task_synthesis import _run_task_summary +def test_unknown_exception_facts_row_does_not_invent_zero_rounds_or_buy_a_model_call(tmp_path, monkeypatch): + from ouroboros.post_task_synthesis import _record_task_facts rows = [] - monkeypatch.setattr("ouroboros.project_dialogue.append_authored_task_summary", - lambda _root, _result_root, row, **_: rows.append(row)) + monkeypatch.setattr("ouroboros.project_dialogue.append_canonical_task_summary", + lambda _root, row: rows.append(row)) monkeypatch.setattr("ouroboros.llm_observability.chat_observed", lambda *_a, **_kw: pytest.fail("no new paid summary")) - _run_task_summary(SimpleNamespace(drive_root=tmp_path), None, {"id": "unknown", "text": "Recover work"}, - {"loop_evidence_unavailable": True}, - {"loop_evidence_unavailable": True, "tool_calls": []}, tmp_path / "logs") + _record_task_facts(SimpleNamespace(drive_root=tmp_path), {"id": "unknown", "text": "Recover work"}, + {"loop_evidence_unavailable": True}, + {"loop_evidence_unavailable": True, "tool_calls": []}, tmp_path / "logs") assert rows[0]["tool_calls"] is None and rows[0]["rounds"] is None - assert "round count unknown" in rows[0]["text"] + assert rows[0]["summary_kind"] == "host_task_facts" and rows[0]["text"] == "" def test_failed_exception_attachment_keeps_the_original_error_and_unknown_projection(tmp_path): diff --git a/tests/test_completion_observations.py b/tests/test_completion_observations.py index a607e17b3..993844c82 100644 --- a/tests/test_completion_observations.py +++ b/tests/test_completion_observations.py @@ -1,7 +1,6 @@ """Synthesis sees task observations without mistaking tool success for receipt.""" import hashlib -import gzip import json import shutil from types import SimpleNamespace @@ -131,33 +130,24 @@ def test_empty_and_legacy_observations_remain_unknown_without_an_extra_artifact( assert "not evidence that no action occurred" in legacy -def test_packet_summary_receives_inline_facts_and_custodies_its_prompt(tmp_path, monkeypatch): +def test_facts_row_buys_no_packet_and_persists_no_narrative(tmp_path, monkeypatch): + """Owner decision 2=A: the paid task narrative is gone. The host facts row + keeps the exact tool census for history with empty text, custodies no + model prompt, and leaves no continuation narrative on the result.""" from ouroboros.observability import OBSERVABILITY_DIR - import ouroboros.post_task_synthesis as synthesis monkeypatch.setattr("ouroboros.skill_readiness.acceptance_skill_lifecycle", lambda *_a, **_k: []) - monkeypatch.setattr("ouroboros.consolidator._consolidation_route", lambda: ("test-model", False)) - # The existing trace builder has already applied its 4000-character bound; - # a second 3000-character head cap used to discard this later fact again. - monkeypatch.setattr(synthesis, "build_trace_summary", lambda trace: "x" * 3200 + " LATE_TRACE_FACT") + monkeypatch.setattr("ouroboros.llm_observability.chat_observed", + lambda *_a, **_k: pytest.fail("the facts row buys no model call")) task = {"id": "summary", "root_task_id": "summary", "chat_id": 1, "type": "task", "text": "show it"} - observations = build_completion_observations(tmp_path, task, _trace()) - sealed = build_sealed_final_package({"completion_observations": observations}, "") - prompts = [] - class CapturingLLM: - def chat(self, **kwargs): - prompts.append(kwargs["messages"][0]["content"]) - return {"content": "Photo submission was recorded; owner receipt is unknown."}, {"cost": 0} - pipeline._run_task_summary(SimpleNamespace(drive_root=tmp_path), CapturingLLM(), task, - {"rounds": 4, "cost": 0}, _trace(), tmp_path / "logs", sealed_final=sealed) - assert len(prompts) == 1 and "LATE_TRACE_FACT" in prompts[0] - assert "send_photo" in prompts[0] and "not a chat receipt" in prompts[0] + env = SimpleNamespace(drive_root=tmp_path, repo_dir=tmp_path) + pipeline._store_task_result(env, task, "Photo sent.", {"rounds": 4, "cost": 0}, _trace()) + pipeline._record_task_facts(env, task, {"rounds": 4, "cost": 0}, _trace(), tmp_path / "logs") + [row] = [json.loads(line) for line in (tmp_path / "logs" / "chat.jsonl").read_text(encoding="utf-8").splitlines()] + assert row["summary_kind"] == "host_task_facts" and row["text"] == "" + assert row["tool_calls"] == 41 and row["tool_errors"] == 0 and row["routing_tool_calls"] == 0 + assert row["tool_call_counts"] == {"send_photo": 1, "send_user_message": 40} + assert "continuation_narrative" not in load_task_result(tmp_path, "summary") calls_root = tmp_path / OBSERVABILITY_DIR / "calls" manifests = [json.loads(path.read_text(encoding="utf-8")) for path in calls_root.glob("*/*.json")] - observed = [row for row in manifests if row.get("call_type") == "task_summary"] - request = next(row for row in observed if row["call_id"].endswith("_request")) - response = next(row for row in observed if row["call_id"].endswith("_response")) - with gzip.open(request["full_payload_ref"]["path"], "rt", encoding="utf-8") as source: - assert json.load(source)["kwargs"]["messages"][0]["content"] == prompts[0] - with gzip.open(response["full_payload_ref"]["path"], "rt", encoding="utf-8") as source: - assert "owner receipt is unknown" in json.load(source)["message"]["content"] + assert [row for row in manifests if row.get("call_type") == "task_summary"] == [] diff --git a/tests/test_consciousness_initiator_label.py b/tests/test_consciousness_initiator_label.py index 2910dc151..bf39d7482 100644 --- a/tests/test_consciousness_initiator_label.py +++ b/tests/test_consciousness_initiator_label.py @@ -95,8 +95,8 @@ def test_the_initiator_survives_the_chat_row_the_summary_row_and_replay(tmp_path "direction": "out", "chat_id": 1, "user_id": 0, "text": "💬 thinking", "content": "💬 thinking", "format": "", "initiator": "consciousness", }) - pipeline._run_task_summary( - env=None, llm=None, + pipeline._record_task_facts( + env=None, task={"id": "w1", "type": "task", "text": "wake", "chat_id": 1, "_is_direct_chat": True, "metadata": dict(WAKE_META)}, usage={"rounds": 1, "cost": 0.0}, llm_trace={"tool_calls": [], "reasoning_notes": []}, drive_logs=tmp_path / "logs", diff --git a/tests/test_conversation_activity_block.py b/tests/test_conversation_activity_block.py index 82c5c3794..50e34f0e0 100644 --- a/tests/test_conversation_activity_block.py +++ b/tests/test_conversation_activity_block.py @@ -6,7 +6,7 @@ chat block's chrome follows the work it holds, never this fact. The block represents an addressing-only turn by the typed routing action, never by a client tool-name list. Live frames learn the fact from the activity census and the rebuilt ``task_done``; replay learns it -from the terminal truth annotation and the authored summary row. These tests +from the terminal truth annotation and the free host facts row. These tests pin each producer of that one fact. """ @@ -52,17 +52,18 @@ def test_summary_row_copies_direct_fact_and_typed_routing_action(): assert "_is_direct_chat" not in plain and "addressing_only" not in plain -def test_authored_summary_row_writes_the_facts_and_history_replays_them(tmp_path): +def test_facts_row_writes_the_facts_and_history_replays_them(tmp_path): drive_logs = tmp_path / "logs" drive_logs.mkdir(parents=True) - pipeline._run_task_summary( - env=None, llm=None, + pipeline._record_task_facts( + env=None, task={"id": "direct-1", "type": "task", "text": "hi", "chat_id": 1, "_is_direct_chat": True}, - usage={"rounds": 1, "cost": 0.0, "typed_routing_action": "promote_chat_to_task"}, - llm_trace={"tool_calls": [], "reasoning_notes": []}, + usage={"rounds": 3, "cost": 0.0, "typed_routing_action": "promote_chat_to_task"}, + llm_trace={"tool_calls": [{"tool": "promote_chat_to_task"}], "reasoning_notes": []}, drive_logs=drive_logs, ) row = json.loads((drive_logs / "chat.jsonl").read_text(encoding="utf-8").splitlines()[0]) + assert row["summary_kind"] == "host_task_facts" and row["text"] == "" assert row["_is_direct_chat"] is True assert row["typed_routing_action"] == "promote_chat_to_task" # No task_results file: the history summary row falls back to the row's copy. @@ -73,6 +74,7 @@ def test_authored_summary_row_writes_the_facts_and_history_replays_them(tmp_path summary = next(item for item in payload if item.get("system_type") == "task_summary") assert summary["_is_direct_chat"] is True assert summary["addressing_only"] == "promote_chat_to_task" + assert summary["tool_calls"] == 1 and summary["routing_tool_calls"] == 1 and summary["rounds"] == 3 def _dispatch_task_done(tmp_path, *, evt_flag, stored_flag): diff --git a/tests/test_core_extraction.py b/tests/test_core_extraction.py index 74fc03e64..b8520a928 100644 --- a/tests/test_core_extraction.py +++ b/tests/test_core_extraction.py @@ -124,7 +124,7 @@ def test_core_catalog_schema_bytes_and_handler_owners_are_stable(): # await_messages companion (395 -> 698 bytes); the `message` parameter description states # the bound. Diffing the whole catalog base to head shows exactly those edits and nothing else. assert hashlib.sha256(schema_bytes).hexdigest() == ( - "7195a7459f276f4bdf8758864b079218f518ae446613cc8351967692f31472cd" + "700568a32d56b3e931eb12b6644c2d38a5d09e294b8a568b8884bf0b5d9b3e7d" ) assert { entry.name: (entry.handler.__module__, entry.handler.__name__) diff --git a/tests/test_dialogue_provenance.py b/tests/test_dialogue_provenance.py index 703da13c1..ff5c1e3b1 100644 --- a/tests/test_dialogue_provenance.py +++ b/tests/test_dialogue_provenance.py @@ -88,15 +88,14 @@ def test_normalized_presence_task_provenance_uses_event_and_ceiling_facts(): assert presence_provenance_from_task({"type": "user"}) == {} -def test_presence_task_summary_row_persists_normalized_provenance(tmp_path): - from ouroboros.agent_task_pipeline import _run_task_summary +def test_presence_task_facts_row_persists_normalized_provenance(tmp_path): + from ouroboros.agent_task_pipeline import _record_task_facts logs = tmp_path / "logs" logs.mkdir() env = SimpleNamespace(drive_root=tmp_path, repo_dir=tmp_path) - _run_task_summary( + _record_task_facts( env, - object(), _presence_task(), {"rounds": 1, "cost": 0.0}, {"tool_calls": []}, diff --git a/tests/test_direct_chat_turn_owner_control.py b/tests/test_direct_chat_turn_owner_control.py index 2b065547a..6a4e5dc39 100644 --- a/tests/test_direct_chat_turn_owner_control.py +++ b/tests/test_direct_chat_turn_owner_control.py @@ -541,7 +541,7 @@ def _inflight_synthesis(monkeypatch, tmp_path, task_id=TURN_ID): monkeypatch.setattr(llm_mod, "LLMClient", lambda *a, **k: object()) for name, attr in (("chat_consolidation", "_run_chat_consolidation"), ("scratchpad_consolidation", "_run_scratchpad_consolidation"), - ("summary", "_run_task_summary"), ("reflection", "_run_reflection"), + ("facts", "_record_task_facts"), ("reflection", "_run_reflection"), ("promotion", "_update_improvement_backlog")): monkeypatch.setattr(pipeline, attr, _stage(name)) monkeypatch.setattr(pipeline, "_apply_reflection_memory_actions", lambda *a, **k: None) @@ -612,13 +612,13 @@ def test_stop_now_during_the_inflight_synthesis_is_accepted_and_stops_the_remain finally: synthesis.release.set() assert second.status_code == 404, second.text - assert synthesis.calls == ["chat_consolidation"], synthesis.calls + assert synthesis.calls == ["facts", "chat_consolidation"], synthesis.calls stored = load_task_result(tmp_path, TURN_ID) assert stored["status"] == "completed" checkpoint = stored["root_phase_checkpoint"] assert checkpoint["post_task_synthesis"] == "degraded", checkpoint assert checkpoint["post_task_stop_reason"] == ( - "owner_stopped:skipped=scratchpad_consolidation,summary,reflection,promotion"), checkpoint + "owner_stopped:skipped=scratchpad_consolidation,reflection,promotion"), checkpoint assert not active_intent(tmp_path, TURN_ID) @@ -654,11 +654,11 @@ def test_cascade_stop_now_during_the_inflight_synthesis_keeps_the_intent_open_an assert intent.get("scope") == SCOPE_CASCADE, intent assert "post-task synthesis" in str(intent.get("last_error") or ""), intent assert _mailbox_kinds(tmp_path) == [] # no control for a loop that is gone - assert synthesis.calls == ["chat_consolidation"], synthesis.calls + assert synthesis.calls == ["facts", "chat_consolidation"], synthesis.calls synthesis.release.set() synthesis.thread.join(30) assert not synthesis.thread.is_alive() - assert synthesis.calls == ["chat_consolidation"], synthesis.calls + assert synthesis.calls == ["facts", "chat_consolidation"], synthesis.calls second = client.post(f"/api/tasks/{TURN_ID}/cancel", json=body) finally: synthesis.release.set() @@ -668,5 +668,5 @@ def test_cascade_stop_now_during_the_inflight_synthesis_keeps_the_intent_open_an checkpoint = stored["root_phase_checkpoint"] assert checkpoint["post_task_synthesis"] == "degraded", checkpoint assert checkpoint["post_task_stop_reason"] == ( - "owner_stopped:skipped=scratchpad_consolidation,summary,reflection,promotion"), checkpoint + "owner_stopped:skipped=scratchpad_consolidation,reflection,promotion"), checkpoint assert not active_intent(tmp_path, TURN_ID) diff --git a/tests/test_lc2_owner_facades.py b/tests/test_lc2_owner_facades.py index 23ed4463c..920f3b8b7 100644 --- a/tests/test_lc2_owner_facades.py +++ b/tests/test_lc2_owner_facades.py @@ -36,7 +36,7 @@ LC2_LEAF_OWNERS: dict[str, tuple[str, str]] = { "ouroboros.agent_task_pipeline", "build_trace_summary _update_improvement_backlog _apply_reflection_memory_actions " "_child_task_evidence _pre_synthesis_usage_snapshot _compact_review_projection " - "_TASK_SUMMARY_PROMPT _run_task_summary _run_chat_consolidation " + "_record_task_facts _run_chat_consolidation " "_run_scratchpad_consolidation _run_reflection" ), "ouroboros.usage_legacy_import": ( diff --git a/tests/test_max_tokens_constants.py b/tests/test_max_tokens_constants.py index 95fad2039..955333614 100644 --- a/tests/test_max_tokens_constants.py +++ b/tests/test_max_tokens_constants.py @@ -81,12 +81,15 @@ def test_summary_and_background_token_budgets(): "ouroboros/tools/review_synthesis.py": "max_tokens=16384", "ouroboros/consolidator.py": "max_tokens=16384", "ouroboros/reflection.py": "max_tokens=16384", - "ouroboros/post_task_synthesis.py": "max_tokens=16384", "ouroboros/tools/skill_publish.py": "max_tokens=8192", } for path, needle in expectations.items(): src = Path(path).read_text(encoding="utf-8").replace(" ", "") assert needle in src, f"{path} must contain {needle}" + # Owner decision 2=A (TZ-2 C5): the paid task narrative is gone; the free + # facts row buys no model call, so the synthesis leaf holds no summary call. + synthesis = Path("ouroboros/post_task_synthesis.py").read_text(encoding="utf-8") + assert "chat_observed" not in synthesis and 'call_type="task_summary"' not in synthesis assert context_compaction._SUMMARY_OUTPUT_TOKENS == 32_768 assert context_compaction._summarizer_spec()["output_budget"] == 32_768 @@ -453,10 +456,11 @@ def test_repository_index_collapses_junk_dirs(): def test_summary_and_reflection_callers_use_bounded_evidence(): - """Summary and reflection prompt builders must call format_review_evidence_for_prompt with max_chars.""" + """The reflection prompt builder (the one paid synthesis reader since the paid + summary's removal) must call format_review_evidence_for_prompt with max_chars.""" from pathlib import Path - for filename in ("ouroboros/post_task_synthesis.py", "ouroboros/reflection.py"): + for filename in ("ouroboros/reflection.py",): src = Path(filename).read_text(encoding="utf-8") assert "format_review_evidence_for_prompt(" in src # Must pass max_chars argument (not rely on default 0) @@ -810,8 +814,7 @@ def test_acceptance_panels_reach_the_synthesis_prompts(): def test_summary_and_reflection_callers_pass_the_acceptance_panels(): from pathlib import Path - # v7 relocated the summary/reflection synthesis callers out of - # agent_task_pipeline.py into ouroboros/post_task_synthesis.py. - for filename in ("ouroboros/post_task_synthesis.py", "ouroboros/reflection.py"): + # The paid summary caller is gone (TZ-2 C5); reflection is the one caller. + for filename in ("ouroboros/reflection.py",): src = Path(filename).read_text(encoding="utf-8") assert "acceptance_panels=" in src, filename diff --git a/tests/test_native_conversation_activity.py b/tests/test_native_conversation_activity.py index a902cc69a..919712a30 100644 --- a/tests/test_native_conversation_activity.py +++ b/tests/test_native_conversation_activity.py @@ -194,7 +194,7 @@ def test_native_post_task_retains_activity_and_delivers_answer_early(monkeypatch monkeypatch.setattr(pipeline, "_run_chat_consolidation", consolidate) for name in ( - "_run_scratchpad_consolidation", "_run_task_summary", "_run_reflection", + "_run_scratchpad_consolidation", "_record_task_facts", "_run_reflection", "_update_improvement_backlog", "_apply_reflection_memory_actions", ): monkeypatch.setattr(pipeline, name, lambda *a, **kw: None) diff --git a/tests/test_native_owner_wait.py b/tests/test_native_owner_wait.py index b8e87efe4..533e6101b 100644 --- a/tests/test_native_owner_wait.py +++ b/tests/test_native_owner_wait.py @@ -40,7 +40,7 @@ def test_native_wait_retains_source_and_leaves_answer_delivery_to_loop(tmp_path, wait_after_tools(ctx, messages, {}, {}, 4, [], set()) assert len(waits) == 1 and waits[0]["state"] == "waiting" after = load_task_result(tmp_path, ctx.task_id)["owner_wait"] - assert after == {**waits[0], "state": "resumed"} + assert after == {**waits[0], "state": "resumed", "resume_reason": "owner_input"} assert after["source_ref"] and "restart_transaction_id" not in after assert ctx.pending_events == [] and ctx._owner_wait_requested == "" assert ctx._loop_mailbox_seen_ids == set() @@ -48,6 +48,66 @@ def test_native_wait_retains_source_and_leaves_answer_delivery_to_loop(tmp_path, assert drain_owner_entries(tmp_path, ctx.task_id, set())[0]["text"] == "Continue with that form" +def test_peer_mail_wakes_but_does_not_answer_an_open_question(tmp_path, monkeypatch): + from ouroboros.owner_mailbox import write_task_message + from ouroboros.owner_quiz import quiz_states, record_asked + import supervisor.message_bus as mb + + ctx = native_context(tmp_path) + ctx.current_chat_id = 1 + record_asked(tmp_path, ctx.task_id, quiz_id="q1", question="Which path?", + options=[], wait_for_answer=True, chat_id=1) + frames = [] + bridge = mb.LocalChatBridge() + bridge._broadcast_fn = frames.append + monkeypatch.setattr(mb, "get_bridge", lambda: bridge) + + def wake(_seconds): + assert write_task_message(tmp_path, "Here is context", task_id=ctx.task_id, + source_task_id="peer-1", provenance="independent_task") + + monkeypatch.setattr("ouroboros.owner_wait.time.sleep", wake) + messages = [] + wait_after_tools(ctx, messages, {}, {}, 1, [], set()) + assert load_task_result(tmp_path, ctx.task_id)["owner_wait"]["resume_reason"] == "mail" + assert quiz_states(tmp_path, ctx.task_id)["q1"]["state"] == "open" + assert "wait_for_answer" not in quiz_states(tmp_path, ctx.task_id)["q1"] + assert len([frame for frame in frames if frame.get("type") == "quiz_state"]) == 1 + assert len(messages) == 1 and "peer-1" in messages[0]["content"] + assert "not a confirmed owner answer" in messages[0]["content"] + + +def test_answer_racing_a_wait_end_never_reopens_a_settled_card(tmp_path, monkeypatch): + from ouroboros.owner_quiz import quiz_states, record_asked, record_answered + from ouroboros.owner_wait import announce_wait_ended + import supervisor.message_bus as mb + + record_asked(tmp_path, "root-1", quiz_id="q1", question="Which path?", options=[], + wait_for_answer=True, chat_id=1) + assert record_answered(tmp_path, "root-1", quiz_id="q1", option_index=None, + request_id="r1", comment="Proceed")['ok'] + frames = [] + bridge = mb.LocalChatBridge() + bridge._broadcast_fn = frames.append + monkeypatch.setattr(mb, "get_bridge", lambda: bridge) + announce_wait_ended(tmp_path, "root-1", "q1", 1) + assert quiz_states(tmp_path, "root-1")["q1"]["state"] == "answered" + assert not [frame for frame in frames if frame.get("type") == "quiz_state"] + + +def test_answer_before_capacity_grant_replaces_a_stale_timeout(tmp_path): + from ouroboros.owner_mailbox import KIND_QUIZ_ANSWER + from ouroboros.owner_wait import _fresh_wake, classify_wake + + ctx = native_context(tmp_path) + assert classify_wake([{"kind": "hurry", "msg_id": "h"}], "q1") == "mail" + assert classify_wake([{"kind": "task_message", "provenance": "ancestor_task"}], "q1") == "mail" + assert classify_wake([{"kind": KIND_QUIZ_ANSWER, "msg_id": "quiz_answer:other"}], "q1") == "owner_input" + assert write_owner_message(tmp_path, "Owner answered", ctx.task_id, + msg_id="quiz_answer:q1", kind=KIND_QUIZ_ANSWER) + assert _fresh_wake(ctx, "q1", "timeout") == "answer" + + def _spent_bound(ctx, minutes=5): """An absolute stamp already in the past — the bound of a wait that ended.""" import datetime @@ -152,7 +212,7 @@ def test_an_answer_before_the_bound_resumes_without_a_timeout_notice(tmp_path, m messages = [] wait_after_tools(ctx, messages, {}, {}, 1, [], set()) assert messages == [] # the owner answered; the loop delivers it as usual - assert "resume_reason" not in load_task_result(tmp_path, ctx.task_id)["owner_wait"] + assert load_task_result(tmp_path, ctx.task_id)["owner_wait"]["resume_reason"] == "owner_input" @pytest.mark.parametrize("reason", ["cancelled", "finalize_requested", "deadline", "absolute_ceiling"]) @@ -166,7 +226,7 @@ def test_a_control_reason_outranks_a_spent_bound(tmp_path, monkeypatch, reason): messages = [] wait_after_tools(ctx, messages, {}, {}, 1, [], set()) row = load_task_result(tmp_path, ctx.task_id)["owner_wait"] - assert row["state"] == "resumed" and "resume_reason" not in row + assert row["state"] == "resumed" and row["resume_reason"] == f"control:{reason}" assert messages == [] diff --git a/tests/test_native_terminal_history.py b/tests/test_native_terminal_history.py index 45bbe5488..04beecc53 100644 --- a/tests/test_native_terminal_history.py +++ b/tests/test_native_terminal_history.py @@ -1,4 +1,4 @@ -"""Native task summaries preserve unknown counters in history.""" +"""Native task facts rows preserve unknown counters in history.""" import asyncio import json from types import SimpleNamespace @@ -9,18 +9,18 @@ from ouroboros.gateway.history import make_chat_history_endpoint from ouroboros.cost_projection import carry_cost_meta -def test_unknown_task_summary_counts_remain_readable_in_history(tmp_path, monkeypatch): - from ouroboros.post_task_synthesis import _run_task_summary +def test_unknown_task_facts_counts_remain_readable_in_history(tmp_path, monkeypatch): + from ouroboros.post_task_synthesis import _record_task_facts monkeypatch.setattr("ouroboros.llm_observability.chat_observed", lambda *_a, **_k: pytest.fail("unknown evidence must not buy a model call")) - _run_task_summary(SimpleNamespace(drive_root=tmp_path), None, - {"id": "uncaptured-summary", "chat_id": 1, "text": "Inspect current work"}, - {"loop_evidence_unavailable": True}, - {"loop_evidence_unavailable": True, "tool_calls": []}, tmp_path / "logs") + _record_task_facts(SimpleNamespace(drive_root=tmp_path), + {"id": "uncaptured-summary", "chat_id": 1, "text": "Inspect current work"}, + {"loop_evidence_unavailable": True}, + {"loop_evidence_unavailable": True, "tool_calls": []}, tmp_path / "logs") response = asyncio.run(make_chat_history_endpoint(tmp_path)(SimpleNamespace(query_params={"chat_id": "1"}))) [row] = json.loads(response.body)["messages"] - assert "round count unknown" in row["text"] + assert row["text"] == "" # unknown stays a typed null below, never prose assert row["tool_calls"] is None and row["rounds"] is None diff --git a/tests/test_owner_wait.py b/tests/test_owner_wait.py index a3f9d8f9f..a8908a9cc 100644 --- a/tests/test_owner_wait.py +++ b/tests/test_owner_wait.py @@ -90,7 +90,7 @@ def test_deferred_question_flush_and_resume_requires_pool_grant(tmp_path): write_owner_message(tmp_path, "Change the requested destination", task_id="root-1") resume = events.get(timeout=3) # An owner answer is never labelled a bound expiry. - assert resume["phase"] == "resume" and "resume_reason" not in resume + assert resume["phase"] == "resume" and resume["resume_reason"] == "owner_input" assert not ended.is_set() commands.put({**identity, "phase": "resume_granted"}) thread.join(timeout=2) diff --git a/tests/test_packaged_runtime_and_lifecycle.py b/tests/test_packaged_runtime_and_lifecycle.py index 8020f12af..6541bb4d6 100644 --- a/tests/test_packaged_runtime_and_lifecycle.py +++ b/tests/test_packaged_runtime_and_lifecycle.py @@ -489,7 +489,7 @@ def test_supervisor_grace_toast_is_not_the_task_answering_it(monkeypatch, tmp_pa assert meta["finalization_requested_at"] == 2000.0 tick(2000.5) # the toast is dispatched here, exactly as the loop does it - assert any("reached idle_timeout" in line for line in tick.delivered), ( + assert any("Task t-narration: The task made no progress for too long." in line for line in tick.delivered), ( "the harness never drained the bus — the toast under test was never dispatched" ) assert meta["last_progress_at"] == 1000.0, "host narration counted as the task's work" diff --git a/tests/test_phase2_cognition.py b/tests/test_phase2_cognition.py index 4fa22d5f6..282089181 100644 --- a/tests/test_phase2_cognition.py +++ b/tests/test_phase2_cognition.py @@ -12,7 +12,7 @@ def _chat_rows(root): return [json.loads(line) for line in path.read_text(encoding="utf-8").splitlines()] -def test_split_project_root_summary_lands_canonically_once_before_child_gc(tmp_path): +def test_split_project_root_facts_row_lands_canonically_once_before_child_gc(tmp_path): from ouroboros import agent_task_pipeline as pipeline canonical = tmp_path / "canonical" @@ -30,8 +30,8 @@ def test_split_project_root_summary_lands_canonically_once_before_child_gc(tmp_p "budget_drive_root": str(canonical), } - pipeline._run_task_summary( - env, object(), task, {"rounds": 1, "cost": 0}, + pipeline._record_task_facts( + env, task, {"rounds": 1, "cost": 0}, {"tool_calls": []}, child / "logs", ) @@ -50,7 +50,7 @@ def test_split_project_root_summary_lands_canonically_once_before_child_gc(tmp_p assert len([row for row in _chat_rows(canonical) if row.get("task_id") == "project-root"]) == 1 -def test_root_checkpoint_prevents_a_second_paid_authored_summary(tmp_path, monkeypatch): +def test_root_checkpoint_prevents_a_second_facts_row_and_buys_no_summary(tmp_path, monkeypatch): from ouroboros import agent_task_pipeline as pipeline from ouroboros.task_results import STATUS_COMPLETED, write_task_result @@ -60,7 +60,7 @@ def test_root_checkpoint_prevents_a_second_paid_authored_summary(tmp_path, monke def chat(self, **_kwargs): self.calls += 1 - return {"content": "Authored once"}, {"cost": 0} + return {"content": "No paid narrative exists"}, {"cost": 0} import ouroboros.llm as llm_mod import ouroboros.memory as memory_mod @@ -98,12 +98,14 @@ def test_root_checkpoint_prevents_a_second_paid_authored_summary(tmp_path, monke (tmp_path / "logs" / "chat.jsonl").replace(archive / "chat_rotated.jsonl") pipeline._run_post_task_processing_async(*args, blocking=True) - assert llm.calls == 1 + assert llm.calls == 0 assert not (tmp_path / "logs" / "chat.jsonl").exists() - assert "root-llm" in (archive / "chat_rotated.jsonl").read_text(encoding="utf-8") + rotated = [json.loads(line) for line in (archive / "chat_rotated.jsonl").read_text(encoding="utf-8").splitlines()] + kinds = [row.get("summary_kind") for row in rotated if row.get("task_id") == "root-llm"] + assert kinds.count("host_task_facts") == 1 and "authored_root_summary" not in kinds -def test_authored_summary_hot_path_never_scans_rotated_biography(tmp_path, monkeypatch): +def test_facts_row_hot_path_never_scans_rotated_biography(tmp_path, monkeypatch): from ouroboros import agent_task_pipeline as pipeline import ouroboros.project_dialogue as dialogue @@ -123,22 +125,14 @@ def test_authored_summary_hot_path_never_scans_rotated_biography(tmp_path, monke monkeypatch.setattr(dialogue, "iter_jsonl_objects", forbidden_scan) - class Llm: - calls = 0 - - def chat(self, **_kwargs): - self.calls += 1 - return {"content": "One paid narrative"}, {"cost": 0} - - llm = Llm() - pipeline._run_task_summary( - SimpleNamespace(drive_root=tmp_path), llm, + pipeline._record_task_facts( + SimpleNamespace(drive_root=tmp_path), {"id": "new-root", "root_task_id": "new-root", "type": "task", "text": "work"}, {"rounds": 2, "cost": 0}, {"tool_calls": [{"tool": "read_file"}]}, tmp_path / "logs", ) - assert llm.calls == 1 + assert [row["summary_kind"] for row in _chat_rows(tmp_path)] == ["host_task_facts"] assert scans == 0 @@ -323,7 +317,7 @@ def test_running_async_root_truth_survives_restart_degradation_once(tmp_path): assert len([row for row in _chat_rows(tmp_path) if row.get("task_id") == "restart-root"]) == 1 -def test_authored_narrative_never_suppresses_final_artifact_failure_truth(tmp_path): +def test_facts_row_never_suppresses_final_artifact_failure_truth(tmp_path): from ouroboros import agent_task_pipeline as pipeline from ouroboros.project_dialogue import append_terminal_task_projection from ouroboros.task_results import STATUS_COMPLETED, write_task_result @@ -333,8 +327,8 @@ def test_authored_narrative_never_suppresses_final_artifact_failure_truth(tmp_pa root_task_id="artifact-root", result="Built output", outcome_axes={"execution": {"status": "ok"}, "artifacts": {"status": "ready"}}, ) - pipeline._run_task_summary( - SimpleNamespace(drive_root=tmp_path), object(), + pipeline._record_task_facts( + SimpleNamespace(drive_root=tmp_path), {"id": "artifact-root", "root_task_id": "artifact-root", "text": "build", "type": "task", "chat_id": 1}, {"rounds": 1, "cost": 0, "outcome_axes": initial["outcome_axes"]}, @@ -353,14 +347,14 @@ def test_authored_narrative_never_suppresses_final_artifact_failure_truth(tmp_pa rows = [row for row in _chat_rows(tmp_path) if row.get("task_id") == "artifact-root"] assert [row["summary_kind"] for row in rows] == [ - "authored_root_summary", "terminal_root_projection", + "host_task_facts", "terminal_root_projection", ] assert rows[-1]["outcome"] == "Failed" assert rows[-1]["outcome_axes"]["artifacts"]["status"] == "failed" assert rows[-1]["outcome_final"] is True -def test_split_authored_narrative_keeps_only_canonical_result_ref_after_child_gc(tmp_path): +def test_split_facts_row_keeps_only_canonical_result_ref_after_child_gc(tmp_path): import shutil from ouroboros import agent_task_pipeline as pipeline @@ -384,8 +378,8 @@ def test_split_authored_narrative_keeps_only_canonical_result_ref_after_child_gc "budget_drive_root": str(canonical), "drive_root": str(child), } - pipeline._run_task_summary( - SimpleNamespace(drive_root=child), object(), task, + pipeline._record_task_facts( + SimpleNamespace(drive_root=child), task, {"rounds": 1, "cost": 0}, {"tool_calls": []}, child / "logs", ) copied = copy_child_task_result(canonical, task) diff --git a/tests/test_post_task_inputs.py b/tests/test_post_task_inputs.py index e07df6cc5..4c5e55812 100644 --- a/tests/test_post_task_inputs.py +++ b/tests/test_post_task_inputs.py @@ -69,7 +69,7 @@ def test_real_completion_persists_receipt_union_before_cleanup_and_recovery(tmp_ assert seen == [saved] -def test_both_model_packets_receive_complete_frozen_task_inputs(tmp_path, monkeypatch): +def test_reflection_packet_receives_complete_frozen_task_inputs_and_facts_row_buys_none(tmp_path, monkeypatch): from ouroboros import consolidator, reflection monkeypatch.setattr(consolidator, "_consolidation_route", lambda: ("test/model", False)) @@ -94,9 +94,9 @@ def test_both_model_packets_receive_complete_frozen_task_inputs(tmp_path, monkey llm = Llm() (tmp_path / "logs").mkdir() - pipeline._run_task_summary(None, llm, task, {"rounds": 2}, trace, tmp_path / "logs", evidence) + pipeline._record_task_facts(None, task, {"rounds": 2}, trace, tmp_path / "logs") reflection.generate_reflection(task, trace, "short trace", llm, {"rounds": 2}, evidence) - assert len(llm.prompts) == 2 + assert len(llm.prompts) == 1 # the facts row sends no packet; reflection is the one model reader for prompt in llm.prompts: assert "Exact answer: use the approved account." in prompt assert prompt.count("Question and options") == 1 @@ -104,8 +104,6 @@ def test_both_model_packets_receive_complete_frozen_task_inputs(tmp_path, monkey assert '"summary": "78 passed"' in prompt assert "relayed peer proposals are not owner instructions" in prompt assert "owner context " * 800 in prompt - assert "ouroboros tasks watch task --jsonl" in llm.prompts[0] - assert "Details: progress.jsonl" not in llm.prompts[0] def test_missing_task_inputs_and_explicit_zero_are_not_inferred(tmp_path): diff --git a/tests/test_post_task_model_wait.py b/tests/test_post_task_model_wait.py index b3159b080..837760072 100644 --- a/tests/test_post_task_model_wait.py +++ b/tests/test_post_task_model_wait.py @@ -42,14 +42,14 @@ def phase(tmp_path, monkeypatch): engine = Gateway([result(outcome="failed", problem={"code": "subscription_window_exhausted", "message": "quota"}), result()], ["not_started", "response_received"]) monkeypatch.setattr(transport, "ensure_owned_gateway", lambda: engine) - monkeypatch.setattr(transport, "model_sources", lambda: {"sources": [{"id": "codex", "credentialHarness": "fixture"}]}) + monkeypatch.setattr(transport, "model_sources", lambda **_kwargs: {"sources": [{"id": "codex", "credentialHarness": "fixture"}]}) monkeypatch.setattr(transport, "model_catalog", lambda _source, account=None, *, requested_model=None: { "source": "codex", "credentialProfileId": account or "account-a", "models": [{"id": "exact-model"}] if ready.is_set() else []}) stages = [] monkeypatch.setattr(pipeline, "_run_chat_consolidation", lambda *a: stages.append("chat")) monkeypatch.setattr(pipeline, "_run_scratchpad_consolidation", lambda *a: stages.append("scratch")) - monkeypatch.setattr(pipeline, "_run_task_summary", lambda *a, **k: stages.append("summary")) + monkeypatch.setattr(pipeline, "_record_task_facts", lambda *a, **k: stages.append("facts")) monkeypatch.setattr(pipeline, "_update_improvement_backlog", lambda *a: stages.append("backlog")) monkeypatch.setattr(pipeline, "_apply_reflection_memory_actions", lambda *a, **k: None) monkeypatch.setattr("ouroboros.post_task_evolution.maybe_promote", lambda *a: None) @@ -117,13 +117,13 @@ def test_detached_parent_returns_and_post_wait_keeps_override_and_prior_stages(p until(lambda: active(f)) owner = active(f) assert owner is not parent and not owner.closed and owner.worker_slot_held is False - assert f.stages == ["chat", "scratch", "summary", "reflection"] and not f.done.is_set() + assert f.stages == ["facts", "chat", "scratch", "reflection"] and not f.done.is_set() assert load_task_result(f.root, f.task["id"])["root_phase_checkpoint"]["post_task_synthesis"] == "running" assert f.engine.uploads[0][0]["account"] == {"mode": "pin", "profileId": "original-choice"} f.ready.set() assert f.done.wait(5) until(lambda: post_task_model_wait(f.root, f.task["id"]) is None) - assert f.stages == ["chat", "scratch", "summary", "reflection", "backlog"] + assert f.stages == ["facts", "chat", "scratch", "reflection", "backlog"] assert len(f.engine.creates) == 2 and f.engine.uploads[0][0]["messages"] == f.engine.uploads[1][0]["messages"] until(lambda: owner.closed) assert load_task_result(f.root, f.task["id"])["status"] == main_status @@ -295,12 +295,12 @@ def _controlled_worker(input_queue, output_queue, data_root, repo_root, resume): "supervisor.worker_process._prepare_worker_task_runtime": lambda: None, "supervisor.worker_process._adopt_published_extensions": lambda *_: None, "ouroboros.llm_claudexor.ensure_owned_gateway": lambda: engine, - "ouroboros.llm_claudexor.model_sources": lambda: {"sources": [{"id": "codex", "credentialHarness": "fixture"}]}, + "ouroboros.llm_claudexor.model_sources": lambda **_kwargs: {"sources": [{"id": "codex", "credentialHarness": "fixture"}]}, "ouroboros.llm_claudexor.model_catalog": lambda source, account=None, **kwargs: { "source": source, "credentialProfileId": account or "account-a", "models": [{"id": "exact-model"}] if resume.is_set() else []}, "ouroboros.agent_task_pipeline._run_chat_consolidation": lambda *_: None, "ouroboros.agent_task_pipeline._run_scratchpad_consolidation": lambda *_: None, - "ouroboros.agent_task_pipeline._run_task_summary": lambda *args, **kwargs: None, + "ouroboros.agent_task_pipeline._record_task_facts": lambda *args, **kwargs: None, "ouroboros.agent_task_pipeline._run_reflection": reflect, "ouroboros.agent_task_pipeline._update_improvement_backlog": lambda *_: None, "ouroboros.agent_task_pipeline._apply_reflection_memory_actions": lambda *args, **kwargs: None, diff --git a/tests/test_post_task_reflection.py b/tests/test_post_task_reflection.py index 2bb948bab..c053212c4 100644 --- a/tests/test_post_task_reflection.py +++ b/tests/test_post_task_reflection.py @@ -18,7 +18,7 @@ def test_project_scoped_post_task_processing_feeds_global_backlog_but_project_me calls = [] reflection = {"backlog_candidates": [{"summary": "tool friction"}], "memory_actions": [{"kind": "note"}]} - monkeypatch.setattr(pipeline, "_run_task_summary", lambda *args, **kwargs: calls.append(("summary",))) + monkeypatch.setattr(pipeline, "_record_task_facts", lambda *args, **kwargs: calls.append(("facts",))) monkeypatch.setattr(pipeline, "_run_reflection", lambda *args, **kwargs: reflection) monkeypatch.setattr(pipeline, "_update_improvement_backlog", lambda _env, entry: calls.append(("backlog", entry)) or 1) monkeypatch.setattr( diff --git a/tests/test_presence_post_task.py b/tests/test_presence_post_task.py index 97f44c5f9..2886d94fd 100644 --- a/tests/test_presence_post_task.py +++ b/tests/test_presence_post_task.py @@ -23,8 +23,8 @@ def test_presence_keeps_own_memory_but_skips_evolution_effects(tmp_path, monkeyp ) monkeypatch.setattr( pipeline, - "_run_task_summary", - lambda *args, **kwargs: calls.append("summary"), + "_record_task_facts", + lambda *args, **kwargs: calls.append("facts"), ) monkeypatch.setattr( pipeline, @@ -71,9 +71,9 @@ def test_presence_keeps_own_memory_but_skips_evolution_effects(tmp_path, monkeyp assert result == {"reflection": "ok"} assert calls == [ + "facts", # free, before every paid stage "chat_consolidation", "scratchpad_consolidation", - "summary", "reflection", "memory_actions", ] @@ -91,7 +91,7 @@ def test_ordinary_task_retains_global_post_task_effects(tmp_path, monkeypatch): monkeypatch.setattr(memory_module, "Memory", lambda **kwargs: object()) monkeypatch.setattr(pipeline, "_run_chat_consolidation", lambda *args, **kwargs: None) monkeypatch.setattr(pipeline, "_run_scratchpad_consolidation", lambda *args, **kwargs: None) - monkeypatch.setattr(pipeline, "_run_task_summary", lambda *args, **kwargs: None) + monkeypatch.setattr(pipeline, "_record_task_facts", lambda *args, **kwargs: None) monkeypatch.setattr( pipeline, "_run_reflection", @@ -138,7 +138,7 @@ def test_presence_post_task_applies_own_experience_with_background_off(tmp_path, from ouroboros.knowledge import read_knowledge_note, resolve_knowledge_address monkeypatch.setattr(llm_module, "LLMClient", lambda: object()) - for name in ("_run_chat_consolidation", "_run_scratchpad_consolidation", "_run_task_summary"): + for name in ("_run_chat_consolidation", "_run_scratchpad_consolidation", "_record_task_facts"): monkeypatch.setattr(pipeline, name, lambda *a, **k: None) entry = {"reflection": "A useful shared moment.", "memory_actions": [{ "type": "knowledge_write", "topic": "shared experience", "scope": "global", diff --git a/tests/test_presence_terminal_latency.py b/tests/test_presence_terminal_latency.py index f75c24bf6..5b60571af 100644 --- a/tests/test_presence_terminal_latency.py +++ b/tests/test_presence_terminal_latency.py @@ -45,7 +45,7 @@ def test_native_agent_returns_durable_result_before_synthesis_and_retry_reuses_i monkeypatch.setattr(pipeline, "_run_chat_consolidation", consolidate) monkeypatch.setattr(pipeline, "_run_scratchpad_consolidation", lambda *_a: stages.append("scratchpad")) - monkeypatch.setattr(pipeline, "_run_task_summary", lambda *_a, **_k: stages.append("summary")) + monkeypatch.setattr(pipeline, "_record_task_facts", lambda *_a, **_k: stages.append("facts")) monkeypatch.setattr(pipeline, "_run_reflection", lambda *_a, **_k: stages.append("reflection")) monkeypatch.setattr(pipeline, "_apply_reflection_memory_actions", lambda *_a, **_k: stages.append("memory")) monkeypatch.setattr(loop, "call_llm_with_retry", lambda *_a, **_k: (_call(outcome, "Reply" if outcome == "message" else ""), 0.0)) @@ -100,7 +100,7 @@ def test_native_agent_returns_durable_result_before_synthesis_and_retry_reuses_i release.set() for thread in set(threading.enumerate()) - threads_before: thread.join(timeout=5) - assert stages.count("summary") == stages.count("reflection") == 2 + assert stages.count("facts") == stages.count("reflection") == 2 for task_id in (first.task_id, second.task_id): row = load_task_result(tmp_path, task_id) assert row["root_phase_checkpoint"]["post_task_synthesis"] == "completed" diff --git a/tests/test_project_plain_rows.py b/tests/test_project_plain_rows.py index f21398ae8..100df9488 100644 --- a/tests/test_project_plain_rows.py +++ b/tests/test_project_plain_rows.py @@ -657,7 +657,9 @@ def test_a_host_salvage_row_is_never_a_bare_headline_and_reason(tmp_path): f"{SALVAGE_EXCERPT_LABEL}: Applied Rewrote the atlas builder and reran the suite." in row["text"] ) - assert "context_overflow." in row["text"] + # The execution cause is the rail's owner sentence (TASK_CAUSE_PHRASES); the code stays on the result. + assert "The task outgrew its context before it could finish cleanly." in row["text"] + assert "context_overflow" not in row["text"] assert "get_task_result" not in row["text"] for marker in ("#", "**", "`"): assert marker not in row["text"] diff --git a/tests/test_quiz_answer.py b/tests/test_quiz_answer.py index f17555058..5ac1d791f 100644 --- a/tests/test_quiz_answer.py +++ b/tests/test_quiz_answer.py @@ -431,7 +431,7 @@ def test_escalate_settled_parent_is_a_typed_dead_end(tmp_path, monkeypatch): def test_escalate_invalid_payload_is_typed(tmp_path): ctx = _tool_ctx(tmp_path) - out = _escalate(ctx, question="?", options=["only-one"], assumption="a") + out = _escalate(ctx, question="?", options=["a"] * 7, assumption="a") assert out.startswith("⚠️ QUIZ_OPTIONS_INVALID") out = _escalate(ctx, question="?", options=["a", "b"], assumption="") assert out.startswith("⚠️ QUIZ_ASSUMPTION_REQUIRED") @@ -693,7 +693,7 @@ def test_own_answer_needs_no_option_index(tmp_path, monkeypatch): entries = drain_owner_entries(tmp_path, "task-1", set()) frame_text = [e for e in entries if e.get("kind") == KIND_QUIZ_ANSWER][0]["text"] - assert ("The owner rejected all offered options and answered verbatim: " + assert ("No option was selected; the owner wrote verbatim: " "neither — use duckdb") in frame_text assert "chose option" not in frame_text diff --git a/tests/test_quiz_display.py b/tests/test_quiz_display.py index 5d1691344..fef124185 100644 --- a/tests/test_quiz_display.py +++ b/tests/test_quiz_display.py @@ -31,7 +31,7 @@ class TestValidateQuizPayload: assert payload["assumption"] == "continuing with the merge" @pytest.mark.parametrize("options", [ - [], ["only-one"], ["a"] * (_MAX_QUIZ_OPTIONS + 1), "not-a-list", + ["a"] * (_MAX_QUIZ_OPTIONS + 1), "not-a-list", [{"detail": "no label"}], ]) def test_bad_options_are_refused_atomically(self, options): @@ -44,6 +44,14 @@ class TestValidateQuizPayload: validate_quiz_payload("q", ["a", "b"], "", " ") assert err.value.code == "QUIZ_ASSUMPTION_REQUIRED" + @pytest.mark.parametrize("options", [None, [], ["Confirm"]]) + def test_open_and_single_choice_questions_keep_the_answer_contract(self, options): + payload = validate_quiz_payload("What should change?", options, "", "", wait_for_answer=True) + assert payload["options"] == ([] if options is None else [{"label": x} for x in options]) + with pytest.raises(QuizValidationError) as error: + validate_quiz_payload("What should change?", options, "", "") + assert error.value.code == "QUIZ_ASSUMPTION_REQUIRED" + def test_wait_bound_is_whole_minutes_capped_by_the_task_ceiling(self, monkeypatch): """The bound belongs to a REQUIRED wait and can never promise more time than the absolute wall-clock ceiling already allows (beyond it the @@ -139,7 +147,7 @@ def test_send_quiz_refuses_invalid_payload_and_missing_ids(monkeypatch, tmp_path ok, error = bridge.send_quiz(1, quiz_id="qz", question="q", options=[{"label": "a"}, {"label": "b"}], assumption="x") assert not ok and "task_id" in error ok, error = bridge.send_quiz(1, quiz_id="qz", question="q", options=[{"label": "a"}], assumption="x", task_id="t") - assert not ok + assert (ok, error) == (True, "ok") ok, error = bridge.send_quiz(-5, quiz_id="qz", question="q", options=[{"label": "a"}, {"label": "b"}], assumption="x") assert (ok, error) == (True, "ok") # A2A chats: silent no-op, like links @@ -171,9 +179,12 @@ def test_handle_send_quiz_prefers_bound_project_chat(monkeypatch): assert sent[0][1]["quiz_id"] == "qz-2" assert sent[0][1]["assumption"] == "path A meanwhile" - # No options -> typed drop, no bridge call. + # An explicitly empty options list is an open question; absence is not. sent.clear() _handle_send_quiz({**evt, "options": []}, ctx) + assert sent and sent[0][1]["options"] == [] + sent.clear() + _handle_send_quiz({key: value for key, value in evt.items() if key != "options"}, ctx) assert sent == [] # Headless exception: an interactive card in the hidden chat-0 panel can diff --git a/tests/test_rail_cause_coverage.py b/tests/test_rail_cause_coverage.py new file mode 100644 index 000000000..ab23d2a3b --- /dev/null +++ b/tests/test_rail_cause_coverage.py @@ -0,0 +1,224 @@ +"""TZ-2 C1: every typed reason the runtime can stamp on a task row has one owner +sentence in BOTH cause twins, and the rails that end a task say it on every +owner surface — the card line, the reaper's grace toast, kill notice and salvage +line, and the loop's own fallback text — while the typed code stays on the +record and on the ``task_incident`` key. + +Coverage is IMPORT-based: the sets below are the modules' own typed +vocabularies (``outcomes.BEST_EFFORT_REASON_CODES``, both sides of +``outcomes.ACCEPTANCE_BYPASS_REASON_BY_RAIL``, the ``REASON_*`` constants and +``queue_timeouts.TIMEOUT_TERMINAL_REASONS``), never a regex over source text, +so a reason minted outside a registry is a registry bug rather than a silent +raw code on the card. A code with no sentence still stays raw by design +(docs/DESIGN.md §4); the exemptions below are the codes that render nothing on +purpose. +""" + +from __future__ import annotations + +import pathlib +import time +import types +from types import SimpleNamespace + +import pytest + +from ouroboros import outcomes +from ouroboros.project_dialogue import ( + OUTCOME_PHASE_HEADLINE, + TASK_CAUSE_PHRASES, + _completion_verdict, + outcome_phase, +) +from supervisor.queue_timeouts import TIMEOUT_TERMINAL_REASONS +from tests._cancel_intents_shared import qenv # noqa: F401 - shared reaper fixture + +REPO = pathlib.Path(__file__).resolve().parents[1] +TWIN = REPO / "web" / "modules" / "log_events.js" + +# The codes that state nothing by design: a clean delivery has no cause, and the +# owner's own stop carries its marker instead of a sentence, on both surfaces. +SILENT_BY_DESIGN = frozenset({ + outcomes.REASON_FINAL_MESSAGE, + outcomes.REASON_OWNER_REQUESTED_FINALIZATION, + outcomes.ACCEPTANCE_BYPASS_REASON_BY_RAIL[outcomes.REASON_OWNER_REQUESTED_FINALIZATION], +}) + + +def _imported_reason_codes() -> set: + named = {value for name, value in vars(outcomes).items() + if name.startswith("REASON_") and isinstance(value, str)} + return ( + named + | set(outcomes.BEST_EFFORT_REASON_CODES) + | set(outcomes.ACCEPTANCE_BYPASS_REASON_BY_RAIL) + | set(outcomes.ACCEPTANCE_BYPASS_REASONS) + | set(TIMEOUT_TERMINAL_REASONS) + ) - SILENT_BY_DESIGN + + +# ---------------------------------------------------------------- the vocabulary + +def test_every_imported_reason_code_has_one_sentence_in_both_twins(): + codes = _imported_reason_codes() + assert {"round_limit", "finalization_grace", "deadline_local", "children_unabsorbed", + "context_overflow", "absolute_ceiling", "deadline", "idle_timeout", + "provider_failure", "empty_final_text"} <= codes + twin = TWIN.read_text(encoding="utf-8") + for code in sorted(codes): + assert code in TASK_CAUSE_PHRASES, f"no owner sentence for typed reason {code!r}" + assert f'{code}: "{TASK_CAUSE_PHRASES[code]}",' in twin, f"the browser twin lacks {code!r}" + + +def test_the_twins_carry_byte_identical_sentences_for_every_key(): + twin = TWIN.read_text(encoding="utf-8") + for code, sentence in TASK_CAUSE_PHRASES.items(): + assert f'{code}: "{sentence}",' in twin, code + + +def test_the_timeout_rails_are_one_typed_tuple_in_priority_order(): + from supervisor import queue_timeouts + + assert TIMEOUT_TERMINAL_REASONS == ("absolute_ceiling", "deadline", "idle_timeout") + assert (queue_timeouts.REASON_ABSOLUTE_CEILING, queue_timeouts.REASON_DEADLINE, + queue_timeouts.REASON_IDLE_TIMEOUT) == TIMEOUT_TERMINAL_REASONS + + +def test_c1_adds_no_status_word(): + assert list(OUTCOME_PHASE_HEADLINE.values()) == ["Working", "Done", "Done with warnings", "Failed", "Cancelled"] + + +# ---------------------------------------------------------------- the card line + +def _reaped(reason: str) -> dict: + return {"status": "failed", "reason_code": reason, + "outcome_axes": outcomes.terminal_outcome_axes( + lifecycle="failed", execution=outcomes.EXECUTION_INFRA_FAILED, + reason_code=reason, review_trigger="supervisor_terminal")} + + +@pytest.mark.parametrize("reason", TIMEOUT_TERMINAL_REASONS) +def test_a_reaped_row_states_the_rail_in_owner_words(reason): + record = _reaped(reason) + assert outcome_phase(record, {}) == "error" + expected = f"{TASK_CAUSE_PHRASES[reason]}." + assert _completion_verdict(record, {}) == _completion_verdict({}, record) == expected + + +def test_a_railed_best_effort_row_states_the_rail_in_owner_words(): + railed = {"status": "completed", "reason_code": "round_limit", + "outcome_axes": {"execution": {"status": "best_effort", "reason_code": "round_limit"}}} + assert outcome_phase(railed, {}) == "warn" + assert _completion_verdict(railed, {}) == "The task hit its round limit before it could finish cleanly." + + +def test_an_unknown_rail_code_still_stays_raw_on_the_row(): + assert _completion_verdict(_reaped("some_future_rail"), {}) == "some_future_rail." + + +# ---------------------------------------------------------------- the reaper's surfaces + +def test_the_grace_toast_speaks_the_sentence_and_keeps_the_typed_incident(tmp_path, monkeypatch): + from supervisor import task_reaper, workers + + put: list = [] + monkeypatch.setattr(workers, "get_event_q", lambda: types.SimpleNamespace(put=put.append), raising=False) + assert task_reaper.request_finalization_grace(tmp_path, "t1", "idle_timeout", chat_id=7, stamp=1000) + (toast,) = put + assert toast["text"].startswith( + "⏳ Task t1: The task made no progress for too long. Finalize artifacts/results now; ") + assert "idle_timeout" not in toast["text"] + assert toast["progress_meta"] == {"task_incident": "idle_timeout", "toast_once": "t1:idle_timeout:1000"} + # An unknown rail keeps its raw code on the toast rather than borrowing a sentence. + put.clear() + task_reaper.request_finalization_grace(tmp_path, "t2", "some_future_rail", chat_id=7, stamp=1000) + assert put[0]["text"].startswith("⏳ Task t2: some_future_rail. Finalize") + assert put[0]["progress_meta"]["task_incident"] == "some_future_rail" + + +def test_the_salvage_line_names_the_cause_after_the_supervisor_stop(tmp_path, monkeypatch): + from ouroboros import observability + from supervisor import task_reaper, terminal_delivery + + delivered: list = [] + monkeypatch.setattr(observability, "latest_llm_response_text", lambda *_a, **_k: "partial work") + monkeypatch.setattr(observability, "preserved_salvage_path", lambda *_a, **_k: "") + monkeypatch.setattr(terminal_delivery, "deliver_unreviewed_salvage", + lambda drive, task, tid, **kw: delivered.append({"task_id": tid, **kw})) + q = types.SimpleNamespace(DRIVE_ROOT=tmp_path, _task_drive_for_task=lambda task, tid: tmp_path) + task = {"id": "r1", "chat_id": 4} + task_reaper._deliver_reap_salvage(q, task, "r1", "absolute_ceiling") + (call,) = delivered + assert call["outcome"] == "stopped by the supervisor. The task reached its maximum running time" + # The real builder frames it as one owner line, the code nowhere in it. + event = terminal_delivery.build_unreviewed_salvage_event( + tmp_path, task, "r1", outcome=call["outcome"], salvaged_text=call["salvaged_text"]) + assert event["text"].startswith( + "⚠️ Task r1 was stopped by the supervisor. The task reached its maximum running time. " + "Below is the last persisted intermediate model message") + assert "absolute_ceiling" not in event["text"] + + +def test_the_kill_notice_leads_with_the_cause_and_keeps_the_typed_task_done(qenv, monkeypatch): # noqa: F811 + import ouroboros.tools.services as services_mod + from supervisor import task_reaper as tr + from supervisor import workers as workers_mod + + events: list = [] + sent: list = [] + monkeypatch.setattr(tr, "_kill_and_confirm_worker_dead", lambda *_a, **_kw: True) + monkeypatch.setattr(tr, "_deliver_reap_salvage", lambda *_a, **_kw: None) + monkeypatch.setattr(tr, "send_with_budget", lambda cid, text, **kw: sent.append((cid, text, kw))) + monkeypatch.setattr(services_mod, "archive_task_service_logs", lambda *a, **k: None) + monkeypatch.setattr(workers_mod, "get_event_q", lambda: types.SimpleNamespace(put=events.append)) + monkeypatch.setattr(qenv.q, "reconstruct_task_cost", + lambda tid, fields=True, **_kw: {"cost_accounting_status": "available", + "cost_final": True, "cost_usd": 0.0}) + tr.reap_timed_out_task({ + "worker_id": 0, "proc": None, "task_id": "reap1", "task": {"id": "reap1", "chat_id": 4}, + "task_type": "chat", "terminal_reason": "absolute_ceiling", "attempt": 3, "owner_chat_id": 0, + "runtime_sec": 10.0, "will_retry": False, "ceiling_reached": True, + }) + (chat_id, text, kw), = sent + assert chat_id == 4 + assert text.startswith("🛑 The task reached its maximum running time: task reap1 killed after 10s.\n") + assert "Absolute ceiling reached; task stopped." in text + assert "absolute_ceiling" not in text + assert kw["progress_meta"]["task_incident"] == "task_reaper_stopped" + (done,) = [e for e in events if e.get("type") == "task_done"] + assert done["reason_code"] == "absolute_ceiling" + assert done["outcome_axes"]["execution"]["reason_code"] == "absolute_ceiling" + + +# ---------------------------------------------------------------- the loop's fallback + +def _limit_ctx(tmp_path): + import ouroboros.loop as loop_mod + + return loop_mod._RoundLimitContext( + messages=[], llm=object(), active_model="model", active_effort="high", + max_retries=1, drive_logs=tmp_path / "logs", task_id="task", round_idx=1, + event_queue=None, accumulated_usage={}, task_type="task", + active_use_local=False, max_rounds=1, deadline_ts=time.time() - 1, + tools=SimpleNamespace(_ctx=SimpleNamespace()), llm_trace={}, + ) + + +def test_the_loop_fallback_speaks_the_rail_and_keeps_an_unknown_rail_raw(tmp_path, monkeypatch): + import ouroboros.loop as loop_mod + + monkeypatch.setattr(loop_mod, "_finalize_forced_services", lambda *_args: None) + monkeypatch.setattr(loop_mod, "_call_forced_model_once", + lambda *_a, **_k: (_ for _ in ()).throw(AssertionError("paid final call"))) + monkeypatch.setattr(loop_mod, "_forced_fallback_result", + lambda _ctx, _trace, text, reason, **_kw: (text, _ctx.accumulated_usage, _trace)) + for rail, expected in ( + ("absolute_ceiling", "⚠️ The task reached its maximum running time; finalization grace produced no answer."), + ("deadline", "⚠️ The task reached its deadline; finalization grace produced no answer."), + ("", "⚠️ The task reached its deadline; finalization grace produced no answer."), + ("some_future_rail", "⚠️ Task reached some_future_rail; finalization grace produced no answer."), + ): + ctx = _limit_ctx(tmp_path) + text, usage, _trace = loop_mod._handle_forced_finalization(ctx, rail) + assert text == expected, rail + assert usage["reason_code"] == "finalization_grace" diff --git a/tests/test_root_post_task_synthesis.py b/tests/test_root_post_task_synthesis.py index aed139513..da895d935 100644 --- a/tests/test_root_post_task_synthesis.py +++ b/tests/test_root_post_task_synthesis.py @@ -5,9 +5,11 @@ by theme; every moved block is verbatim. Covers the durable `root_phase_checkpoint` state machine and its exact-subtree cost reconciliation, startup recovery of pending/indeterminate synthesis, the shared pre-synthesis usage snapshot taken once before worker dispatch, and -that snapshot reaching (or staying out of) the summary and reflection prompts. +that snapshot reaching (or staying out of) the free facts row and the +reflection prompt. """ +import json from types import SimpleNamespace import ouroboros.agent_task_pipeline as pipeline @@ -255,9 +257,9 @@ def test_root_synthesis_uses_one_shared_nonfinal_subtree_cost_snapshot(tmp_path, ) monkeypatch.setattr( pipeline, - "_run_task_summary", - lambda _env, _llm, _task, usage, *_args, **_kwargs: ( - order.append("summary"), snapshots.append(usage) + "_record_task_facts", + lambda _env, _task, usage, *_args, **_kwargs: ( + order.append("facts"), snapshots.append(usage) ), ) monkeypatch.setattr( @@ -293,8 +295,8 @@ def test_root_synthesis_uses_one_shared_nonfinal_subtree_cost_snapshot(tmp_path, assert reads == [(tmp_path, "root-synthesis", "")] assert order[:5] == [ - "snapshot", "chat_consolidation", "scratchpad_consolidation", - "summary", "reflection", + "snapshot", "facts", "chat_consolidation", "scratchpad_consolidation", + "reflection", ] assert len(snapshots) == 2 and snapshots[0] is snapshots[1] snapshot = snapshots[0] @@ -386,9 +388,10 @@ def test_pre_synthesis_cost_failure_is_unavailable_not_zero(tmp_path, monkeypatc assert pipeline._synthesis_cost_text(snapshot) == "cost unavailable (non-final)" -def _capture_summary_and_reflection_prompts( +def _capture_facts_row_and_reflection_prompt( tmp_path, monkeypatch, usage, *, task_overrides=None, ): + """The shared snapshot reaches the free facts row's flat fields and the one paid prompt.""" import ouroboros.consolidator as consolidator monkeypatch.setattr( @@ -425,15 +428,14 @@ def _capture_summary_and_reflection_prompts( "reasoning_notes": [], } - summary_llm = CapturingLlm() - pipeline._run_task_summary( + pipeline._record_task_facts( env=None, - llm=summary_llm, task=task, usage=usage, llm_trace=trace, drive_logs=drive_logs, ) + [row] = [json.loads(line) for line in (drive_logs / "chat.jsonl").read_text(encoding="utf-8").splitlines()] reflection_llm = CapturingLlm() entry = pipeline._run_reflection( @@ -446,12 +448,11 @@ def _capture_summary_and_reflection_prompts( ) assert entry is not None - assert len(summary_llm.prompts) == 1 assert len(reflection_llm.prompts) == 1 - return summary_llm.prompts[0], reflection_llm.prompts[0] + return row, reflection_llm.prompts[0] -def test_shared_cost_snapshot_reaches_summary_and_reflection_prompts(tmp_path, monkeypatch): +def test_shared_cost_snapshot_reaches_facts_row_and_reflection_prompt(tmp_path, monkeypatch): snapshot = { "rounds": 8, "cost": 1.25, @@ -472,9 +473,14 @@ def test_shared_cost_snapshot_reaches_summary_and_reflection_prompts(tmp_path, m }, } - prompts = _capture_summary_and_reflection_prompts( + row, prompt = _capture_facts_row_and_reflection_prompt( tmp_path, monkeypatch, snapshot, ) + for key in ("accounted_upper_bound_usd_with_children", "reserved_usd", "unresolved_upper_bound_usd", + "unknown_unmetered", "cost_final", "cost_with_children_partial", "cost_accounting_status"): + assert row[key] == snapshot[key], key + assert row["reason_code"] == "child_results_deferred" + assert row["outcome_axes"]["execution"]["status"] == "degraded" snapshot_text = pipeline._synthesis_usage_snapshot_text(snapshot) expected_fragments = ( '"accounted_upper_bound_usd_with_children": 4.75', @@ -489,18 +495,17 @@ def test_shared_cost_snapshot_reaches_summary_and_reflection_prompts(tmp_path, m '"reason_code": "child_results_deferred"', '"status": "best_effort"', ) - for prompt in prompts: - assert snapshot_text in prompt - assert "accounted subtree cost only" in prompt - assert "separate non-final exposure fields" in prompt - assert "including the reserved" not in prompt - assert "outcome_axes` is canonical task truth" in prompt - assert '"review": {' in prompt - for fragment in expected_fragments: - assert fragment in prompt + assert snapshot_text in prompt + assert "accounted subtree cost only" in prompt + assert "separate non-final exposure fields" in prompt + assert "including the reserved" not in prompt + assert "outcome_axes` is canonical task truth" in prompt + assert '"review": {' in prompt + for fragment in expected_fragments: + assert fragment in prompt -def test_unavailable_cost_snapshot_is_null_not_zero_in_both_prompts(tmp_path, monkeypatch): +def test_unavailable_cost_snapshot_is_null_not_zero_in_facts_row_and_prompt(tmp_path, monkeypatch): snapshot = { "rounds": 8, "cost": 1.25, @@ -515,7 +520,7 @@ def test_unavailable_cost_snapshot_is_null_not_zero_in_both_prompts(tmp_path, mo "cost_accounting_status": "unavailable", } - prompts = _capture_summary_and_reflection_prompts( + row, prompt = _capture_facts_row_and_reflection_prompt( tmp_path, monkeypatch, snapshot, ) snapshot_text = pipeline._synthesis_usage_snapshot_text(snapshot) @@ -525,20 +530,22 @@ def test_unavailable_cost_snapshot_is_null_not_zero_in_both_prompts(tmp_path, mo "unresolved_upper_bound_usd", "unknown_unmetered", ) - for prompt in prompts: - assert snapshot_text in prompt - for field in null_fields: - assert f'"{field}": null' in prompt - assert '"ledger_integrity": "unavailable"' in prompt - assert '"cost_snapshot_at": "2026-07-15T12:35:00+00:00"' in prompt - assert '"cost_final": false' in prompt - assert '"cost_with_children_partial": true' in prompt - assert '"cost_accounting_status": "unavailable"' in prompt - assert "$0" not in prompt + for field in null_fields: + assert row[field] is None, field + assert row["cost_accounting_status"] == "unavailable" + assert snapshot_text in prompt + for field in null_fields: + assert f'"{field}": null' in prompt + assert '"ledger_integrity": "unavailable"' in prompt + assert '"cost_snapshot_at": "2026-07-15T12:35:00+00:00"' in prompt + assert '"cost_final": false' in prompt + assert '"cost_with_children_partial": true' in prompt + assert '"cost_accounting_status": "unavailable"' in prompt + assert "$0" not in prompt def test_child_legacy_usage_does_not_claim_a_subtree_snapshot(tmp_path, monkeypatch): - prompts = _capture_summary_and_reflection_prompts( + row, prompt = _capture_facts_row_and_reflection_prompt( tmp_path, monkeypatch, {"rounds": 8, "cost": 1.25}, @@ -550,8 +557,7 @@ def test_child_legacy_usage_does_not_claim_a_subtree_snapshot(tmp_path, monkeypa }, ) - for prompt in prompts: - assert "Shared pre-synthesis cost snapshot" not in prompt - assert "accounted_upper_bound_usd_with_children" not in prompt - assert "cost_snapshot_at" not in prompt - assert "Cost: $1.25" in prompts[0] + assert "Shared pre-synthesis cost snapshot" not in prompt + assert "accounted_upper_bound_usd_with_children" not in prompt + assert "cost_snapshot_at" not in prompt + assert "accounted_upper_bound_usd_with_children" not in row and "cost_snapshot_at" not in row diff --git a/tests/test_run_origin_provenance.py b/tests/test_run_origin_provenance.py index 129ee7386..241273143 100644 --- a/tests/test_run_origin_provenance.py +++ b/tests/test_run_origin_provenance.py @@ -117,6 +117,26 @@ def test_first_row_label_follows_the_owner_stamp_and_owner_bytes_do_not_change() assert ctx._owner_directives[0]["content"] == "Initial requirement verbatim", label +def test_promoted_objective_is_not_owner_corpus_even_with_inherited_owner_stamp(): + from ouroboros.context import build_user_content + from ouroboros.loop_messages import owner_source_sha256 + + metadata = {"origin_message_ref": OWNER_REF, "objective_author": {"kind": "task", "task_id": "draft-1"}, + "owner_corpus": [{"source": "origin_message", "content": "Proceed to implementation"}, + {"source": "owner_quiz_answer", "content": "Yes, implement"}]} + ctx = _first_row(metadata, "This is NOT permission to implement [SWARM_INITIATIVE]") + assert [row["content"] for row in ctx._owner_directives] == ["Proceed to implementation", "Yes, implement"] + assert all("source_task_id" not in row for row in ctx._owner_directives) # owner, not drafter + assert owner_source_sha256(ctx) != owner_source_sha256( + SimpleNamespace(_owner_directives=[{"source": "initial_user", "content": "This is NOT permission to implement"}])) + rendered = build_user_content({"text": "This is NOT permission to implement", "metadata": metadata}) + assert rendered.startswith("[OBJECTIVE_AUTHOR]") and "task draft-1" in rendered + assert "not spoken by the owner" in rendered + # Host-initiated tasks still carry the owner's first text without a drafter stamp. + assert _first_row({"origin_message_ref": OWNER_REF}, "Owner's exact request")._owner_directives == [ + {"source": "initial_user", "content": "Owner's exact request"}] + + # --- the routing issuer ------------------------------------------------------------- @pytest.mark.parametrize("label,direct,metadata,expected", [ diff --git a/tests/test_slime_c_amendments.py b/tests/test_slime_c_amendments.py index 25891c978..6413bade6 100644 --- a/tests/test_slime_c_amendments.py +++ b/tests/test_slime_c_amendments.py @@ -354,42 +354,29 @@ class TestSealedFinalPackage: ("report.pdf", 123), ] - def test_sealed_package_reaches_summary_prompt(self, tmp_path): + def test_facts_row_retells_nothing_of_the_sealed_final(self, tmp_path, monkeypatch): + """No paid narrative: the host facts row neither prompts a model nor + copies the delivered answer; reflection alone reads the sealed package.""" + import pytest + import ouroboros.agent_task_pipeline as atp - from ouroboros.task_finalization import build_sealed_final_package drive_root, logs = _make_drive(tmp_path) env, _memory, _ctx = _make_fake_env(drive_root) - art = drive_root / "task_results" / "artifacts" / "sum1" - art.mkdir(parents=True) - (art / "report.pdf").write_bytes(b"x" * 123) - - prompts = [] - - class CapturingLLM: - def chat(self, **kwargs): - prompts.append(kwargs["messages"][0]["content"]) - return {"content": "summary text"}, {"cost": 0} - - sealed = build_sealed_final_package( - {"artifacts": [{"name": "report.pdf", "path": str(art / "report.pdf"), - "size": 123, "status": "ready"}]}, - "Delivered the 52-page PDF.") - atp._run_task_summary( - env, CapturingLLM(), + monkeypatch.setattr("ouroboros.llm_observability.chat_observed", + lambda *_a, **_k: pytest.fail("the facts row buys no model call")) + atp._record_task_facts( + env, {"id": "sum1", "type": "task", "chat_id": 1, "text": "make a pdf"}, {"cost": 0.01, "rounds": 3}, {"tool_calls": [{"tool": "shell", "error": ""}], "reasoning_notes": []}, logs, - review_evidence={}, - sealed_final=sealed, ) - assert len(prompts) == 1 - assert "Sealed final outcome (host-attested ground truth)" in prompts[0] - assert "Delivered the 52-page PDF." in prompts[0] - assert "report.pdf (123 bytes)" in prompts[0] - assert "OVERRIDE impressions from the error trace" in prompts[0] + [row] = [json.loads(line) for line in (logs / "chat.jsonl").read_text(encoding="utf-8").splitlines()] + assert row["summary_kind"] == "host_task_facts" and row["text"] == "" + assert row["tool_calls"] == 1 and row["rounds"] == 3 + assert "Delivered" not in json.dumps(row) and "make a pdf" not in json.dumps(row) def test_sealed_package_reaches_reflection_prompt(self, tmp_path, monkeypatch): import ouroboros.llm_observability as obs diff --git a/tests/test_task_summary.py b/tests/test_task_summary.py index c79222325..d933e31f0 100644 --- a/tests/test_task_summary.py +++ b/tests/test_task_summary.py @@ -1,91 +1,136 @@ -"""The task-summary synthesis of ``ouroboros.agent_task_pipeline``. +"""The free host facts row of ``ouroboros.agent_task_pipeline`` (no paid narrative). -Split out of ``tests/test_agent_task_pipeline.py`` when that module was divided -by theme; every moved block is verbatim. Covers `_run_task_summary` model -routing and its chat-row payload (chat_id, flat snapshot cost fields, outcome -axes), the trivial-task LLM bypass, the multi-round zero-tool prompt, the -review-evidence prompt section and `build_trace_summary` failure facts. +Owner decision 2=A (TZ-2 C5) removed the paid "task summary" narrative. What +survives is `_record_task_facts`: one ``task_summary`` chat row of kind +``host_task_facts`` with the facts its readers need (chat_id, flat snapshot cost +fields, outcome axes, tool metrics, routing) and no prose, never labelled as an +authored narrative. Also covers the Light consolidation route the remaining +chat consolidation uses and `build_trace_summary` failure facts. """ import json +from types import SimpleNamespace + +import pytest import ouroboros.agent_task_pipeline as pipeline -def test_task_summary_prefers_direct_model_when_openrouter_missing(tmp_path, monkeypatch): - monkeypatch.delenv("OPENROUTER_API_KEY", raising=False) - monkeypatch.setenv("OPENAI_API_KEY", "test-openai-key") - monkeypatch.setenv("OUROBOROS_MODEL_LIGHT", "openai::gpt-5.5-mini") - monkeypatch.setenv("OUROBOROS_MODEL_FALLBACKS", "openai::gpt-5.5-mini") - monkeypatch.setenv("OUROBOROS_MODEL", "openai::gpt-5.5") - monkeypatch.setenv("OUROBOROS_MODEL_HEAVY", "openai::gpt-5.5") +def _rows(drive_logs): + return [json.loads(line) for line in (drive_logs / "chat.jsonl").read_text(encoding="utf-8").splitlines() if line.strip()] - captured = {} - class FakeLlm: - def chat(self, *, messages, model, reasoning_effort, max_tokens, use_local, model_role=""): - assert model_role == "light" - captured["messages"] = messages - captured["model"] = model - captured["reasoning_effort"] = reasoning_effort - captured["max_tokens"] = max_tokens - captured["use_local"] = use_local - return {"content": "direct summary ok"}, {"cost": 0} +@pytest.fixture +def no_model_calls(monkeypatch): + monkeypatch.setattr("ouroboros.llm_observability.chat_observed", + lambda *_a, **_k: pytest.fail("the facts row buys no model call")) + monkeypatch.setattr("ouroboros.llm.LLMClient", lambda *_a, **_k: pytest.fail("the facts row needs no client")) + +def test_facts_row_buys_no_model_call_and_carries_the_facts(tmp_path, no_model_calls): drive_logs = tmp_path / "logs" drive_logs.mkdir(parents=True) - # Use rounds > 1 so the task is non-trivial and the LLM summary path is taken - pipeline._run_task_summary( + # Non-trivial (several rounds, a tool call): the old paid narrative path. + pipeline._record_task_facts( env=None, - llm=FakeLlm(), task={"id": "task-123", "type": "task", "text": "Reply with exactly OK."}, usage={"rounds": 3, "cost": 0.01, "result_status": "failed", "reason_code": "empty_final_text"}, llm_trace={"tool_calls": [{"tool": "read_file", "args": {}}], "reasoning_notes": []}, drive_logs=drive_logs, ) - assert captured["model"] == "openai::gpt-5.5-mini" - assert captured["use_local"] is False - chat_lines = (drive_logs / "chat.jsonl").read_text(encoding="utf-8").splitlines() - assert len(chat_lines) == 1 - payload = json.loads(chat_lines[0]) + [payload] = _rows(drive_logs) assert payload["type"] == "task_summary" - assert payload["text"] == "direct summary ok" - # Non-trivial task metadata is persisted - assert payload["tool_calls"] == 1 + assert payload["summary_kind"] == "host_task_facts" + assert payload["summary_id"] == "task-facts:task-123" + assert payload["text"] == "" # no prose: the task text is not retold either + assert payload["outcome_final"] is False + assert payload["tool_calls"] == 1 and payload["tool_call_counts"] == {"read_file": 1} assert payload["rounds"] == 3 assert payload["outcome_axes"]["execution"]["status"] == "failed" assert payload["outcome_axes"]["objective"]["status"] == "not_evaluated" assert payload["reason_code"] == "empty_final_text" + assert "source_coverage" not in payload -def test_task_summary_row_carries_chat_id_for_trivial_task(tmp_path): - """A trivial task (no tools, <=1 round) skips the LLM summary but still - stamps the project chat_id, so the summary row routes to its project + +def test_facts_row_is_never_an_authored_narrative_while_legacy_rows_still_resolve(tmp_path, no_model_calls): + from ouroboros.main_context_authority import project_main_task_authority + from ouroboros.project_dialogue import append_canonical_task_summary + from ouroboros.task_results import load_task_result, write_task_result + + ref = {"kind": "task_result", "task_id": "", "reader": "get_task_result"} + for task_id in ("new-root", "legacy-root"): + write_task_result(tmp_path, task_id, "completed", result="R" * 200001) + pipeline._record_task_facts( + SimpleNamespace(drive_root=tmp_path), + {"id": "new-root", "root_task_id": "new-root", "type": "task", "chat_id": 1, "text": "work"}, + {"rounds": 4, "cost": 0.0}, {"tool_calls": [{"tool": "read_file"}]}, tmp_path / "logs", + ) + # A historical paid narrative keeps resolving through the unchanged reader. + legacy_ref = {**ref, "task_id": "legacy-root"} + assert append_canonical_task_summary(tmp_path, { + "type": "task_summary", "summary_kind": "authored_root_summary", + "summary_id": "task-narrative:legacy-root", "task_id": "legacy-root", + "result_ref": legacy_ref, "source_coverage": {"task_result": legacy_ref}, + "text": "Legacy authored account", + }) + + assert "continuation_narrative" not in load_task_result(tmp_path, "new-root") + + def projected(task_id): + authority = {"task_id": task_id, "result": "R" * 200001, "task_contract": {"objective": "old"}, + "source": {**ref, "task_id": task_id, "arguments": {"task_id": task_id, "include_authority": True}}} + return project_main_task_authority( + {"id": "next", "predecessor_authority": authority}, drive_root=tmp_path, + )["predecessor_authority"]["result"] + + new = projected("new-root") + assert new["narrative_status"] == "unavailable" + assert new["narrative_gap"]["kind"] == "continuation_narrative_unavailable" + legacy = projected("legacy-root") + assert legacy["narrative_status"] == "available" + assert legacy["narrative"]["text"] == "Legacy authored account" + + +def test_facts_row_has_no_visible_summary_even_for_a_project(tmp_path, no_model_calls): + drive_logs = tmp_path / "logs" + pipeline._record_task_facts( + None, {"id": "bound", "type": "task", "text": "Ship it", "chat_id": 1, "project_id": "launch"}, + {"rounds": 5, "cost": 0.0}, {"tool_calls": []}, drive_logs, + ) + pipeline._record_task_facts( + None, {"id": "unbound", "type": "task", "text": "Ship it", "chat_id": 1}, + {"rounds": 5, "cost": 0.0}, {"tool_calls": []}, drive_logs, + ) + texts = {row["task_id"]: row["text"] for row in _rows(drive_logs)} + assert texts == {"bound": "", "unbound": ""} + + +def test_facts_row_carries_chat_id_for_trivial_task(tmp_path, no_model_calls): + """The facts row stamps the project chat_id, so it routes to its project thread on history reload instead of defaulting to the main chat.""" drive_logs = tmp_path / "logs" drive_logs.mkdir(parents=True) - pipeline._run_task_summary( + pipeline._record_task_facts( env=None, - llm=None, task={"id": "p1", "type": "task", "text": "hi", "chat_id": 1234}, - usage={"rounds": 1, "cost": 0.0}, + usage={"rounds": 1, "cost": 0.0, "result_status": "infra_failed", "reason_code": "llm_api_error"}, llm_trace={"tool_calls": [], "reasoning_notes": []}, drive_logs=drive_logs, ) - rows = [ - json.loads(line) - for line in (drive_logs / "chat.jsonl").read_text(encoding="utf-8").splitlines() - if line.strip() - ] - summaries = [r for r in rows if r.get("type") == "task_summary"] - assert summaries and summaries[0]["chat_id"] == 1234 + [summary] = [r for r in _rows(drive_logs) if r.get("type") == "task_summary"] + assert summary["chat_id"] == 1234 + assert summary["text"] == "" # the former trivial-task host line is gone + assert summary["tool_calls"] == 0 and summary["rounds"] == 1 + assert summary["outcome_axes"]["execution"]["status"] == "infra_failed" + assert summary["reason_code"] == "llm_api_error" -def test_task_summary_row_carries_flat_snapshot_cost_fields(tmp_path): + +def test_facts_row_carries_flat_snapshot_cost_fields(tmp_path): """v6.82 P1: the task_summary chat row carries the pre-synthesis snapshot's - flat cost fields (previously discarded into prose) so history replay can - show honest card cost. Fields absent from the snapshot (cost_usd, - cost_accounting_error) are never fabricated.""" + flat cost fields so history replay can show honest card cost. Fields absent + from the snapshot (cost_usd, cost_accounting_error) are never fabricated.""" drive_logs = tmp_path / "logs" drive_logs.mkdir(parents=True) snapshot_usage = { @@ -102,20 +147,14 @@ def test_task_summary_row_carries_flat_snapshot_cost_fields(tmp_path): "ledger_integrity": "ok", "cost_accounting_status": "available", } - pipeline._run_task_summary( + pipeline._record_task_facts( env=None, - llm=None, task={"id": "p2", "type": "task", "text": "hi", "chat_id": 1}, usage=snapshot_usage, llm_trace={"tool_calls": [], "reasoning_notes": []}, drive_logs=drive_logs, ) - rows = [ - json.loads(line) - for line in (drive_logs / "chat.jsonl").read_text(encoding="utf-8").splitlines() - if line.strip() - ] - row = next(r for r in rows if r.get("type") == "task_summary") + row = next(r for r in _rows(drive_logs) if r.get("type") == "task_summary") assert row["cost_final"] is False assert row["cost_with_children_partial"] is True # ABI-3 fix-round-2: the snapshot producer emits the honest name only @@ -129,7 +168,20 @@ def test_task_summary_row_carries_flat_snapshot_cost_fields(tmp_path): assert "cost_usd" not in row assert "cost_accounting_error" not in row -def test_task_summary_uses_configured_light_model_when_openrouter_present(monkeypatch): + +def test_facts_row_failure_is_contained(tmp_path, monkeypatch, caplog): + """A failed append names the task and never raises into post-task work.""" + import ouroboros.project_dialogue as dialogue + + monkeypatch.setattr(dialogue, "append_canonical_task_summary", + lambda *_a, **_k: (_ for _ in ()).throw(OSError("disk full"))) + with caplog.at_level("WARNING"): + pipeline._record_task_facts(None, {"id": "full-disk", "chat_id": 1}, {"rounds": 2}, {"tool_calls": []}, + tmp_path / "logs") + assert "full-disk" in caplog.text + + +def test_consolidation_route_uses_configured_light_model_when_openrouter_present(monkeypatch): from ouroboros.consolidator import _consolidation_route monkeypatch.setenv("OPENROUTER_API_KEY", "test-openrouter-key") @@ -144,7 +196,8 @@ def test_task_summary_uses_configured_light_model_when_openrouter_present(monkey assert _consolidation_route() == ("openai/gpt-5.5-mini", False) -def test_task_summary_accepts_openai_compatible_when_legacy_base_url_is_present(monkeypatch): + +def test_consolidation_route_accepts_openai_compatible_when_legacy_base_url_is_present(monkeypatch): from ouroboros.consolidator import _consolidation_route monkeypatch.delenv("OPENROUTER_API_KEY", raising=False) @@ -159,6 +212,7 @@ def test_task_summary_accepts_openai_compatible_when_legacy_base_url_is_present( assert _consolidation_route() == ("openai-compatible::custom-model", False) + def test_build_trace_summary_shows_structured_failure_facts(): trace = { "tool_calls": [{ @@ -192,104 +246,3 @@ def test_build_trace_summary_shows_structured_failure_facts(): "reasoning_notes": ["note" * 2000], } assert "OMISSION NOTE" in pipeline.build_trace_summary(long_trace) - -def test_task_summary_prompt_includes_review_evidence(tmp_path, monkeypatch): - monkeypatch.setenv("OPENAI_API_KEY", "test-openai-key") - monkeypatch.setenv("OUROBOROS_MODEL_LIGHT", "openai::gpt-5.5-mini") - - captured = {} - - class FakeLlm: - def chat(self, *, messages, model, reasoning_effort, max_tokens, use_local, model_role=""): - assert model_role == "light" - captured["prompt"] = messages[0]["content"] - return {"content": "summary with review evidence"}, {"cost": 0} - - drive_logs = tmp_path / "logs" - drive_logs.mkdir(parents=True) - - pipeline._run_task_summary( - env=None, - llm=FakeLlm(), - task={"id": "task-review", "type": "task", "text": "Fix commit flow"}, - usage={"rounds": 4, "cost": 0.02}, - llm_trace={"tool_calls": [{"tool": "commit_reviewed", "args": {}}], "reasoning_notes": []}, - drive_logs=drive_logs, - review_evidence={ - "has_evidence": True, - "recent_attempts": [{ - "status": "blocked", - "critical_findings": [{ - "severity": "critical", - "item": "tests_affected", - "reason": "broken", - }], - }], - }, - ) - - assert "Structured review evidence" in captured["prompt"] - assert "tests_affected" in captured["prompt"] - assert "critical" in captured["prompt"] - assert "meta-reflection" in captured["prompt"].lower() - assert "What friction, errors, or weak assumptions slowed the work?" in captured["prompt"] - assert "What should Ouroboros change in its own process or prompts" in captured["prompt"] - assert "keep it to 1-2 sentences and DO NOT add meta-reflection" in captured["prompt"] - -def test_trivial_task_summary_bypasses_llm_and_uses_short_format(tmp_path): - class FailIfCalledLlm: - def chat(self, *args, **kwargs): # pragma: no cover - should never be called - raise AssertionError("LLM summary path must be skipped for trivial tasks") - - drive_logs = tmp_path / "logs" - drive_logs.mkdir(parents=True) - - pipeline._run_task_summary( - env=None, - llm=FailIfCalledLlm(), - task={"id": "task-trivial", "type": "task", "text": "Say hi"}, - usage={"rounds": 1, "cost": 0.0, "result_status": "infra_failed", "reason_code": "llm_api_error"}, - llm_trace={"tool_calls": [], "reasoning_notes": []}, - drive_logs=drive_logs, - ) - - payload = json.loads((drive_logs / "chat.jsonl").read_text(encoding="utf-8").splitlines()[0]) - assert payload["type"] == "task_summary" - assert payload["task_id"] == "task-trivial" - assert payload["text"] == "Task task-trivial (task): Say hi. 1r, $0.00." - assert payload["tool_calls"] == 0 - assert payload["rounds"] == 1 - assert payload["outcome_axes"]["execution"]["status"] == "infra_failed" - assert payload["outcome_axes"]["objective"]["status"] == "not_evaluated" - assert payload["reason_code"] == "llm_api_error" - -def test_multi_round_zero_tool_task_uses_llm_summary_prompt(tmp_path, monkeypatch): - monkeypatch.setenv("OPENAI_API_KEY", "test-openai-key") - monkeypatch.setenv("OUROBOROS_MODEL_LIGHT", "openai::gpt-5.5-mini") - - captured = {} - - class FakeLlm: - def chat(self, *, messages, model, reasoning_effort, max_tokens, use_local, model_role=""): - assert model_role == "light" - captured["prompt"] = messages[0]["content"] - return {"content": "multi-round summary"}, {"cost": 0} - - drive_logs = tmp_path / "logs" - drive_logs.mkdir(parents=True) - - pipeline._run_task_summary( - env=None, - llm=FakeLlm(), - task={"id": "task-zero-tool-multi-round", "type": "task", "text": "Think carefully"}, - usage={"rounds": 3, "cost": 0.01}, - llm_trace={"tool_calls": [], "reasoning_notes": ["note"]}, - drive_logs=drive_logs, - ) - - assert "0 tool calls and ≤1 round" in captured["prompt"] - assert "DO NOT add meta-reflection" in captured["prompt"] - payload = json.loads((drive_logs / "chat.jsonl").read_text(encoding="utf-8").splitlines()[0]) - assert payload["text"] == "multi-round summary" - assert payload["tool_calls"] == 0 - assert payload["rounds"] == 3 diff --git a/tests/test_task_tool_metrics.py b/tests/test_task_tool_metrics.py index 2f525c7b9..7e37c8e7a 100644 --- a/tests/test_task_tool_metrics.py +++ b/tests/test_task_tool_metrics.py @@ -1,4 +1,4 @@ -"""Exact generic call evidence survives metrics, authored summary and history.""" +"""Exact generic call evidence survives metrics, the free facts row and history.""" import asyncio import json import time @@ -8,7 +8,7 @@ import pytest from ouroboros.agent_task_pipeline import emit_task_results from ouroboros.gateway.history import make_chat_history_endpoint -from ouroboros.post_task_synthesis import _run_task_summary, task_tool_metrics +from ouroboros.post_task_synthesis import _record_task_facts, task_tool_metrics from ouroboros.task_results import load_task_result from ouroboros.utils import append_jsonl from supervisor.events_worker_reports import _handle_task_metrics @@ -38,15 +38,11 @@ def test_actual_metrics_summary_and_history_keep_complete_tool_census(tmp_path, env = SimpleNamespace(drive_root=tmp_path, repo_dir=tmp_path) logs = tmp_path / "logs" logs.mkdir() - # Only the authored summary's model response is substituted. All writers, - # metric forwarding, stored result and history projection execute normally. + # The facts row buys no model call. All writers, metric forwarding, + # stored result and history projection execute normally. model_calls = [] - - def summary_response(*args, **kwargs): - model_calls.append(kwargs) - return {"content": "Exact authored summary."}, {} - - monkeypatch.setattr("ouroboros.llm_observability.chat_observed", summary_response) + monkeypatch.setattr("ouroboros.llm_observability.chat_observed", + lambda *args, **kwargs: model_calls.append(kwargs)) usage = {"rounds": 2 if names else 0} pending = [] emit_task_results(env, None, None, pending, task, "Done.", usage, trace, time.time(), logs) @@ -58,27 +54,24 @@ def test_actual_metrics_summary_and_history_keep_complete_tool_census(tmp_path, bridge=SimpleNamespace(push_log=forwarded.append)) _handle_task_metrics(metric, ctx) stored_before = load_task_result(tmp_path, task["id"]) - _run_task_summary(env, None, task, {**usage, "outcome_axes": metric["outcome_axes"], - "reason_code": metric["reason_code"]}, trace, logs) - authored = next(json.loads(line) for line in (logs / "chat.jsonl").read_text().splitlines() - if json.loads(line).get("summary_kind") == "authored_root_summary") + _record_task_facts(env, task, {**usage, "outcome_axes": metric["outcome_axes"], + "reason_code": metric["reason_code"]}, trace, logs) + facts = next(json.loads(line) for line in (logs / "chat.jsonl").read_text().splitlines() + if json.loads(line).get("summary_kind") == "host_task_facts") replay = next(row for row in _history(tmp_path) if row["system_type"] == "task_summary") addressing = sum(1 for name in names if name.strip() in ("promote_chat_to_task", "route_to_project", "steer_task")) - for row in (metric, evaluation, forwarded[0], authored, replay): + for row in (metric, evaluation, forwarded[0], facts, replay): assert row["tool_calls"] == len(names) assert row["tool_errors"] == len(failed) assert row["routing_tool_calls"] == addressing assert row["tool_call_counts"] == expected assert row["outcome_axes"]["execution"] == metric["outcome_axes"]["execution"] - assert len(model_calls) == bool(names) - if names: - assert "Owner objective" in model_calls[0]["messages"][0]["content"] - assert authored["text"] == "Exact authored summary." + assert model_calls == [] and facts["text"] == "" stored_after = load_task_result(tmp_path, task["id"]) assert stored_after["result"] == stored_before["result"] == "Done." for key in ("status", "outcome_axes", "accounted_upper_bound_usd", "cost_final"): assert stored_after[key] == stored_before[key] - assert replay["text"] == authored["text"] + assert replay["text"] == facts["text"] @pytest.mark.parametrize("trace, expected_total, expected_errors", [ @@ -106,9 +99,9 @@ def test_unknown_and_legacy_evidence_keep_absence_distinct_from_zero(tmp_path, m monkeypatch.setattr("ouroboros.llm_observability.chat_observed", lambda *a, **k: pytest.fail("unavailable trace buys no summary call")) task = {"id": "unknown", "chat_id": 1, "text": "Original objective"} - _run_task_summary(SimpleNamespace(drive_root=tmp_path), None, task, - {"loop_evidence_unavailable": True}, - {"loop_evidence_unavailable": True, "tool_calls": []}, tmp_path / "logs") + _record_task_facts(SimpleNamespace(drive_root=tmp_path), task, + {"loop_evidence_unavailable": True}, + {"loop_evidence_unavailable": True, "tool_calls": []}, tmp_path / "logs") [unknown] = _history(tmp_path) assert all(unknown[key] is None for key in ("tool_calls", "tool_errors", "routing_tool_calls", "tool_call_counts")) append_jsonl(tmp_path / "logs/chat.jsonl", {"type": "task_summary", "task_id": "legacy", diff --git a/tests/test_telegram_quiz_answers.py b/tests/test_telegram_quiz_answers.py index 0cd6a1cb7..c4a06a574 100644 --- a/tests/test_telegram_quiz_answers.py +++ b/tests/test_telegram_quiz_answers.py @@ -132,6 +132,29 @@ def test_quiz_card_carries_one_button_per_option_and_remembers_the_card(tmp_path assert api.logs == [] +def test_open_question_is_a_plain_message_with_a_reply_address(tmp_path, monkeypatch): + plugin = _load_plugin() + _settings(tmp_path) + monkeypatch.setattr(plugin, "TelegramClient", Client) + api = Api(tmp_path) + asyncio.run(plugin._make_quiz(api)({**_EVENT, "options": [], "wait_for_answer": True})) + first = _LAST_CLIENT[-1] + assert first.panels == [] + assert first.sent == [(42, "Question: Which db?\nWaiting for your answer; Stop and the task deadline still apply." + "\nReply to this message with your answer.")] + record = plugin.telegram_quiz.quiz_for_message(api, 42, 1) + assert record and record["options"] == [] + Client.updates = [{"update_id": 100, "message": { + "message_id": 1001, "chat": {"id": 42, "type": "private"}, "from": {"id": 42}, + "text": "Try a different database", "reply_to_message": {"message_id": 1}, + }}] + posts = [] + assert _run_poller(plugin, api, monkeypatch, posts) == [] + assert posts == [("/chat/decision", { + "request_id": "tg:100", "decision_id": "quiz:task-1:q1", "comment": "Try a different database", + })] + + def test_tapped_option_reaches_the_decision_ingress_and_settles_the_card(tmp_path, monkeypatch): plugin = _load_plugin() api = _send_card(plugin, tmp_path, monkeypatch) diff --git a/tests/test_terminal_file_boundary.py b/tests/test_terminal_file_boundary.py index 60cbdf50b..63b4f2dd3 100644 --- a/tests/test_terminal_file_boundary.py +++ b/tests/test_terminal_file_boundary.py @@ -160,7 +160,7 @@ def test_post_task_cleanup_waits_for_pooled_files_only(tmp_path, monkeypatch, po monkeypatch.setattr(pipeline, "_pre_synthesis_usage_snapshot", lambda *_a: {}) monkeypatch.setattr("ouroboros.llm.LLMClient", lambda: object()) monkeypatch.setattr("ouroboros.memory.Memory", lambda **_kw: object()) - for name in ("_run_chat_consolidation", "_run_scratchpad_consolidation", "_run_task_summary", + for name in ("_run_chat_consolidation", "_run_scratchpad_consolidation", "_record_task_facts", "_run_reflection", "_update_improvement_backlog", "_apply_reflection_memory_actions"): monkeypatch.setattr(pipeline, name, lambda *_a, **_kw: None) monkeypatch.setattr("ouroboros.post_task_evolution.maybe_promote", lambda *_a: None) @@ -213,7 +213,7 @@ def test_pooled_post_work_preserves_real_followup_until_copyback(tmp_path, monke monkeypatch.setattr(pipeline, "_pre_synthesis_usage_snapshot", lambda *_a: {}) monkeypatch.setattr("ouroboros.llm.LLMClient", lambda: object()) monkeypatch.setattr("ouroboros.memory.Memory", lambda **_kw: object()) - for name in ("_run_chat_consolidation", "_run_scratchpad_consolidation", "_run_task_summary", + for name in ("_run_chat_consolidation", "_run_scratchpad_consolidation", "_record_task_facts", "_run_reflection", "_update_improvement_backlog", "_apply_reflection_memory_actions"): monkeypatch.setattr(pipeline, name, lambda *_a, **_kw: None) monkeypatch.setattr("ouroboros.post_task_evolution.maybe_promote", lambda *_a: None) diff --git a/tests/test_timeout_policy.py b/tests/test_timeout_policy.py index 6dbc25f44..20d24c6f6 100644 --- a/tests/test_timeout_policy.py +++ b/tests/test_timeout_policy.py @@ -1025,7 +1025,7 @@ def test_expired_supervisor_grace_does_not_dispatch_a_paid_final_call(monkeypatc monkeypatch.setattr(loop_mod, "_forced_fallback_result", fake_fallback) result = loop_mod._handle_forced_finalization(ctx, "idle_timeout") - assert result[0].startswith("⚠️ Task reached idle_timeout") + assert result[0].startswith("⚠️ The task made no progress for too long; finalization grace produced no answer.") assert observed["source"] == "finalization_grace_window_elapsed" assert ctx.accumulated_usage == { "execution_status": "failed", diff --git a/tests/test_ui_coherence_browser.py b/tests/test_ui_coherence_browser.py index 1b68730f8..27568f80c 100644 --- a/tests/test_ui_coherence_browser.py +++ b/tests/test_ui_coherence_browser.py @@ -709,7 +709,7 @@ def test_question_mirrors_full_form_settle_and_reload(subscription_ui, width, he assert len(decisions) == 2, 'a reload and a navigation never answer anything' -@pytest.mark.parametrize('width', [1100, 320]) +@pytest.mark.parametrize('width', [1100, 375]) def test_short_question_and_routing_cards_keep_their_width_floor(subscription_ui, width): """The shared card floor survives shrink-to-fit but yields to a narrow column.""" page = subscription_ui['page'] @@ -726,6 +726,8 @@ def test_short_question_and_routing_cards_keep_their_width_floor(subscription_ui const column = document.querySelector('#chat-messages'); column.append(decision.buildQuizCard({ type: 'quiz', task_id: 'width-proof', quiz_id: 'short', state: 'answered', answered_index: 0, question: 'Ок?', options: ['Да', 'Нет'] })); + column.append(decision.buildQuizCard({ type: 'quiz', task_id: 'width-proof', quiz_id: 'open', + state: 'open', question: 'What matters most here?', options: [], wait_for_answer: true })); const owner = media.bubbleFrameNode({ role: 'user' }, document.createElement('span')); owner.dataset.clientMessageId = 'width-route'; decision.renderRoutingDecision(owner, { status: 'needs_manual_target', routing_token: 'width-token', @@ -745,7 +747,9 @@ def test_short_question_and_routing_cards_keep_their_width_floor(subscription_ui inside: box.right <= rect.right - parseFloat(cs.paddingRight) + 1}; })}; }""") - assert len(metrics['cards']) == 2, metrics + assert len(metrics['cards']) == 3, metrics + assert page.locator('.chat-quiz-card[data-quiz-id="open"] .chat-quiz-option').count() == 0 + assert page.locator('.chat-quiz-card[data-quiz-id="open"] .chat-quiz-comment').count() == 1 assert metrics['overflow'] <= 1, metrics assert all(card['width'] >= card['minimum'] - 1 and card['inside'] for card in metrics['cards']), metrics setup_browser.capture(page, f'question-routing-width-{width}') diff --git a/tests/test_ui_failed_finalizing_browser.py b/tests/test_ui_failed_finalizing_browser.py index 537071871..5bfc32145 100644 --- a/tests/test_ui_failed_finalizing_browser.py +++ b/tests/test_ui_failed_finalizing_browser.py @@ -5,9 +5,11 @@ real WebSocket reconnect and on a narrow light page — then settles once. Only model judgment is a fixture. The provider refusal is a real HTTP 401 on the model wire, the host's own provider-unavailable rail salvages the answer, -and the one event-held call is the post-task summary (identified by the first -line of the production prompt), so synthesis is provably open while the -browser looks. Nothing is timed: every wait is an event or a bounded DOM poll. +and the one event-held call is the post-task reflection (identified by the first +line of the production prompt; the fixture's one tool round reads a missing +file, so the typed reflection trigger fires — there is no paid summary), so +synthesis is provably open while the browser looks. Nothing is timed: every +wait is an event or a bounded DOM poll. """ import json @@ -15,7 +17,7 @@ import pytest from devtools.benchmarks.common.server_runner import _api from ouroboros.contracts.chat_id_policy import WEB_UI_CHAT_ID -from ouroboros.post_task_synthesis import _TASK_SUMMARY_PROMPT +from ouroboros.reflection import _REFLECTION_PROMPT_HEAD from tests.test_owner_wait_integration import wait_clone as clone_fixture from tests.system_e2e.harness import ( ArtifactOracle, KeylessIsolatedServer, ModelGate, ScriptedStubModel, body_text, @@ -26,10 +28,9 @@ wait_clone = clone_fixture pytestmark = [pytest.mark.serial, pytest.mark.browser] MARKER = "FAILED_FINALIZING_REAL_ACTOR" -SUMMARY_MARKER = _TASK_SUMMARY_PROMPT.splitlines()[0] +REFLECTION_MARKER = _REFLECTION_PROMPT_HEAD.splitlines()[0] SALVAGE_MARKER = "[PROVIDER_UNAVAILABLE]" # ouroboros/loop.py::_provider_unavailable_result -SALVAGE = "The provider refused the next request; the VERSION read before the outage is retained." -SUMMARY = "Episodic summary: read VERSION, then the provider refused the next round with HTTP 401." +SALVAGE = "The provider refused the next request; the work before the outage is retained." NEUTRAL = "Nothing further to record for this fixture." TITLE = "Provider outage proof" CARD = '#chat-messages .chat-live-card[data-task-id="{}"]' @@ -67,13 +68,13 @@ OBSERVE_JS = """tid => { const card = document.querySelector(`#chat-messages .ch class _OutageModel(ScriptedStubModel): - """One real tool round, then a provider refusal (HTTP 401, a permanent class, - so no backoff retries) on every later tool round of the marked task. The host's - salvage call and the post-task summary get their own answers; the summary is - the one call the gate holds.""" + """One real (failing) tool round, then a provider refusal (HTTP 401, a permanent + class, so no backoff retries) on every later tool round of the marked task. The + host's salvage call gets its own answer; the post-task reflection is the one + call the gate holds.""" def __init__(self, gate): - super().__init__([{"tool": "read_file", "arguments": {"path": "VERSION"}}], + super().__init__([{"tool": "read_file", "arguments": {"path": "MISSING_BEFORE_OUTAGE.txt"}}], final_answer=NEUTRAL, gate=gate) self.refused = 0 outer, base = self, self._server.RequestHandlerClass @@ -110,8 +111,6 @@ class _OutageModel(ScriptedStubModel): text = body_text(body) if SALVAGE_MARKER in text: return "salvage", {"role": "assistant", "content": SALVAGE} - if SUMMARY_MARKER in text: - return "summary", {"role": "assistant", "content": SUMMARY} return super()._answer(body, seq) @@ -164,7 +163,7 @@ def test_failed_root_reads_failed_then_finalizing(wait_clone, tmp_path, monkeypa # only and must itself be a temp-root path, so the test never writes elsewhere. shots = tmp_path / "screenshots" shots.mkdir(parents=True, exist_ok=True) - gate = ModelGate(lambda body: SUMMARY_MARKER in body_text(body) and MARKER in body_text(body), timeout=300) + gate = ModelGate(lambda body: REFLECTION_MARKER in body_text(body) and MARKER in body_text(body), timeout=300) facts = {} with _OutageModel(gate) as model: settings_path = root / "settings.json" @@ -181,12 +180,12 @@ def test_failed_root_reads_failed_then_finalizing(wait_clone, tmp_path, monkeypa # An owner's open Main: the admission name frame must reach a live socket. desk.wait_for_function("() => window.__ouroWs?.ws?.readyState === 1", timeout=60000) created = _api(server.base_url, "POST", "/api/tasks", { - "description": f"{MARKER}: read VERSION, then report what it says.", + "description": f"{MARKER}: read MISSING_BEFORE_OUTAGE.txt, then report what it says.", "title": TITLE, "chat_id": WEB_UI_CHAT_ID, "source": "web", "memory_mode": "forked", "metadata": {"delegation_role": "root"}}) task_id = str(created.get("task_id") or "") assert task_id, created - # The held summary IS the open synthesis: the early final already left. + # The held reflection IS the open synthesis: the early final already left. assert gate.arrived.wait(180), (model.kinds(), oracle.task_result(task_id)) # The worker's own (forked) row already settled Failed; the canonical # row stays live until task_done, and carries the open checkpoint. diff --git a/tests/test_ui_result_browser.py b/tests/test_ui_result_browser.py index 507a95af2..b0fcfeb6f 100644 --- a/tests/test_ui_result_browser.py +++ b/tests/test_ui_result_browser.py @@ -65,19 +65,16 @@ def _seed_history(root): start_time=0.0, drive_logs=logs) final = next(row for row in pending if row["type"] == "send_message") log_chat("out", 1, 1, final["text"], task_id=direct["id"], message_meta=final.get("progress_meta", {}), drive_root=root) - # Ordinary native history receives counts from its normal authored summary, - # not from the removed ephemeral final-frame metadata producer. Only the - # summary model answer is a fixture; all summary/history writers are real. - from ouroboros.post_task_synthesis import _run_task_summary + # Ordinary native history receives counts from its free host facts row, not + # from the removed ephemeral final-frame metadata producer. This stopped turn + # dispatched no post-task worker, so the fixture calls the real writer (no + # model call, no narrative); all history readers are real. + from ouroboros.post_task_synthesis import _record_task_facts from ouroboros.gateway.history import _assemble_history_response - with pytest.MonkeyPatch.context() as summary_model: - summary_model.setattr("ouroboros.llm_observability.chat_observed", lambda *_a, **_k: ( - {"content": "Read the evidence and routed the follow-up into its Project."}, {}, - )) - _run_task_summary(env, None, direct, {"rounds": 2}, direct_trace, logs) + _record_task_facts(env, direct, {"rounds": 2}, direct_trace, logs) summaries = [row for row in json.loads(_assemble_history_response(root, 1, 50, 200))["messages"] if row.get("task_id") == direct["id"] and row.get("system_type") == "task_summary"] - assert len(summaries) == 1 and summaries[0]["tool_calls"] == 2 + assert len(summaries) == 1 and summaries[0]["tool_calls"] == 2 and summaries[0]["text"] == "" return {"preserved_path": str(preserved), "preserved_bytes": preserved.read_bytes()} diff --git a/web/modules/chat_decision.js b/web/modules/chat_decision.js index ec5aa5638..4a027b988 100644 --- a/web/modules/chat_decision.js +++ b/web/modules/chat_decision.js @@ -165,7 +165,7 @@ export function createChatDecision({ // What the copy can offer: the form needs the question and its options; only a known open or // finished question takes an answer. Either may arrive later than the first delivery. const mirrorShape = (row) => { - const complete = Boolean(row.question) && (row.options?.length || 0) >= 2; + const complete = Boolean(row.question) && Array.isArray(row.options) && row.options.length <= MAX_QUIZ_OPTIONS; return { complete, answerable: complete && ANSWERABLE_QUIZ_STATES.includes(row.quiz_state) }; }; @@ -185,7 +185,9 @@ export function createChatDecision({ const current = view.needsValidation && state !== 'answered' ? { ...frame, state: 'unknown' } : observe({ ...frame, state }, live); for (const field of MIRROR_FIELDS) - if (field in current && (current[field] == null || current[field] === '' || current[field]?.length === 0)) delete current[field]; + if (field in current && (current[field] == null || current[field] === '' + || (current[field]?.length === 0 + && (field !== 'options' || view.row.options?.length > 0)))) delete current[field]; view.row = { ...view.row, ...current, quiz_state: current.state }; return view.row; } @@ -366,11 +368,13 @@ export function createChatDecision({ ...(src.recommended_index === index ? { recommended: true } : {}) } : option)); const corrupt = normalized.some( (option) => !option || typeof option !== 'object' || !String(option.label || '').trim()); - const options = corrupt ? [] : normalized.slice(0, MAX_QUIZ_OPTIONS); + const optionsKnown = Array.isArray(src.options) && !corrupt && normalized.length <= MAX_QUIZ_OPTIONS; + const options = optionsKnown ? normalized : []; return { quizId: String(src.quiz_id || ''), question: String((nested ? msg.text : src.question) || ''), options, + optionsKnown, stake: String(src.stake || ''), assumption: String(src.assumption || ''), // The wait facts the header and the signature line read (waitFacts): the @@ -575,7 +579,7 @@ export function createChatDecision({ const quiz = normalizeQuiz(msg); // A Main mirror keeps its way to the Project even while its row cannot carry the whole // form yet: then it shows what is known and takes no answer (never a guessed one). - const complete = Boolean(quiz.question) && quiz.options.length >= 2; + const complete = Boolean(quiz.question) && quiz.optionsKnown; if (!quiz.quizId || !quiz.taskId || !(complete || mirror)) return null; const key = questionKey(quiz.taskId, quiz.quizId); const frame = { task_id: quiz.taskId, quiz_id: quiz.quizId, state: quiz.state, @@ -689,8 +693,8 @@ export function createChatDecision({ }); optionsBox.append(btn); }); - if (complete) card.append(optionsBox); - if (complete && quiz.detailsUnavailable) { + if (complete && quiz.options.length) card.append(optionsBox); + if (complete && quiz.options.length && quiz.detailsUnavailable) { const note = document.createElement('div'); note.className = 'chat-quiz-stake chat-quiz-details-unavailable'; note.textContent = 'Option details were not retained for this older question.'; diff --git a/web/modules/log_events.js b/web/modules/log_events.js index eb3255a0b..ca612fe93 100644 --- a/web/modules/log_events.js +++ b/web/modules/log_events.js @@ -458,6 +458,21 @@ const TASK_CAUSE_PHRASES = { host_child_status_suffix: "A child task had not settled when the answer was delivered", invalid_delivery_control_after_repair: "Ouroboros's final delivery instruction could not be read even after repair, so the answer stands as delivered.", budget_exhausted: "The task ran out of budget before it could finish cleanly", + // The other forced-finalization rails (outcomes.BEST_EFFORT_REASON_CODES / ACCEPTANCE_BYPASS_REASON_BY_RAIL keys). + round_limit: "The task hit its round limit before it could finish cleanly", + finalization_grace: "The task hit a time limit and had to wrap up before it could finish cleanly", + deadline_local: "The task reached its deadline before it could finish cleanly", + context_overflow: "The task outgrew its context before it could finish cleanly", + children_unabsorbed: "Some sub-task results were never folded in, so the task had to wrap up", + // The supervisor's timeout rails (queue_timeouts.TIMEOUT_TERMINAL_REASONS): the reaper's task_done reason_code. + absolute_ceiling: "The task reached its maximum running time", + deadline: "The task reached its deadline", + idle_timeout: "The task made no progress for too long", + // The reason codes outcomes.derive_loop_outcome stamps from typed terminal facts. + provider_failure: "The model provider failed to answer, so the task could not finish", + empty_final_text: "The task ended without a final answer", + deep_self_review_unavailable: "The deep self-review could not run", + deep_self_review_error: "The deep self-review stopped on an error", // #869: the provider-death rail's terminal words (twin of project_dialogue.TASK_CAUSE_PHRASES). provider_unavailable: "The model provider stopped answering, so the task could not finish", delivery_control_degraded: "Ouroboros's final delivery instruction could not be applied, so the answer stands as delivered.", diff --git a/web/tests/awaited_review_not_degraded.test.js b/web/tests/awaited_review_not_degraded.test.js index 5bf267832..6ac9da0d9 100644 --- a/web/tests/awaited_review_not_degraded.test.js +++ b/web/tests/awaited_review_not_degraded.test.js @@ -60,7 +60,7 @@ test('the awaited fact never outranks a decision sentence, a rail reason or an o }); assert.equal(taskReasonDetail(decided), 'The one allowed improvement pass was already used.'); const railed = { ...done({ execution: { ...awaited, status: 'best_effort' } }), reason_code: 'round_limit' }; - assert.equal(taskReasonDetail(railed), 'round_limit'); + assert.equal(taskReasonDetail(railed), 'The task hit its round limit before it could finish cleanly'); assert.equal(taskOutcomeSeverity(railed), 'warn'); const stopped = { ...done({ execution: awaited }), reason_code: 'owner_requested_finalization' }; assert.equal(taskReasonDetail(stopped), ''); diff --git a/web/tests/chat_decision.test.js b/web/tests/chat_decision.test.js index e66cb0c7f..e2e4b182b 100644 --- a/web/tests/chat_decision.test.js +++ b/web/tests/chat_decision.test.js @@ -193,13 +193,35 @@ test('an accepted answer marks the chosen option; degenerate cards refuse to ren assert.ok(buttons[1].classList.contains('chosen')); assert.ok(buttons.every((btn) => btn.disabled)); - assert.equal(fx.decision.buildQuizCard({ ...WS_MSG, options: [{ label: 'only' }] }), null); + assert.equal(fx.decision.buildQuizCard({ ...WS_MSG, quiz_id: 'one', options: [{ label: 'only' }] }) + .querySelectorAll('.chat-quiz-option').length, 1); assert.equal(fx.decision.buildQuizCard({ ...WS_MSG, quiz_id: '' }), null); // An anonymous quiz has no answer address: refuse to render buttons. assert.equal(fx.decision.buildQuizCard({ ...WS_MSG, task_id: '' }), null); } finally { fx.restore(); } }); +test('an open question offers a free-text answer without fabricated options', async () => { + const fx = fixture(); + try { + const card = fx.decision.buildQuizCard({ ...WS_MSG, quiz_id: 'open', options: [] }); + assert.ok(card); + assert.equal(card.querySelectorAll('.chat-quiz-option').length, 0); + const answer = card.querySelector('.chat-quiz-comment'); + assert.ok(answer); + answer.value = 'I would take a different route.'; + answer.listeners.get('input')(); + card.querySelector('.chat-quiz-send').click(); + await new Promise((resolve) => setTimeout(resolve, 0)); + const body = JSON.parse(fx.calls[0].init.body); + assert.equal(body.comment, 'I would take a different route.'); + assert.equal(Object.hasOwn(body, 'option_index'), false); + assert.equal(fx.decision.buildQuizCard({ ...WS_MSG, quiz_id: 'absent', options: undefined }), null); + assert.equal(fx.decision.buildQuizCard({ ...WS_MSG, quiz_id: 'too-many', + options: Array.from({ length: 7 }, (_, i) => ({ label: `Choice ${i}` })) }), null); + } finally { fx.restore(); } +}); + test('one corrupt option refuses that card only, preserving index integrity', () => { const fx = fixture(); try { diff --git a/web/tests/fixtures/outcome_phase_parity.json b/web/tests/fixtures/outcome_phase_parity.json index 1fbc83471..b1b907f3d 100644 --- a/web/tests/fixtures/outcome_phase_parity.json +++ b/web/tests/fixtures/outcome_phase_parity.json @@ -719,6 +719,90 @@ "phase": "warn", "headline": "Done with warnings", "acceptance_clause": "Ouroboros delivered this answer on its own judgement; the reviewers had not signed it off." + }, + { + "name": "the round-limit rail says its own sentence on a best-effort finish", + "record": {"status": "completed", "reason_code": "round_limit", "outcome_axes": {"execution": {"status": "best_effort", "reason_code": "round_limit"}}}, + "phase": "warn", + "headline": "Done with warnings", + "acceptance_clause": "The task hit its round limit before it could finish cleanly." + }, + { + "name": "the deadline_local rail says its own sentence on a best-effort finish", + "record": {"status": "completed", "reason_code": "deadline_local", "outcome_axes": {"execution": {"status": "best_effort", "reason_code": "deadline_local"}}}, + "phase": "warn", + "headline": "Done with warnings", + "acceptance_clause": "The task reached its deadline before it could finish cleanly." + }, + { + "name": "the children_unabsorbed rail says its own sentence on a best-effort finish", + "record": {"status": "completed", "reason_code": "children_unabsorbed", "outcome_axes": {"execution": {"status": "best_effort", "reason_code": "children_unabsorbed"}}}, + "phase": "warn", + "headline": "Done with warnings", + "acceptance_clause": "Some sub-task results were never folded in, so the task had to wrap up." + }, + { + "name": "a grace window that produced no answer names the time limit, not the code", + "record": {"status": "failed", "reason_code": "finalization_grace", "outcome_axes": {"execution": {"status": "failed", "reason_code": "finalization_grace"}}}, + "phase": "error", + "headline": "Failed", + "acceptance_clause": "The task hit a time limit and had to wrap up before it could finish cleanly." + }, + { + "name": "the context-overflow rail says its own sentence", + "record": {"status": "failed", "reason_code": "context_overflow", "outcome_axes": {"execution": {"status": "infra_failed", "reason_code": "context_overflow"}}}, + "phase": "error", + "headline": "Failed", + "acceptance_clause": "The task outgrew its context before it could finish cleanly." + }, + { + "name": "a reaped task names its absolute ceiling in owner words", + "record": {"status": "failed", "reason_code": "absolute_ceiling", "outcome_axes": {"lifecycle": {"status": "failed"}, "execution": {"status": "infra_failed", "reason_code": "absolute_ceiling"}, "review": {"status": "skipped", "trigger": "supervisor_terminal"}}}, + "phase": "error", + "headline": "Failed", + "acceptance_clause": "The task reached its maximum running time." + }, + { + "name": "a reaped task names its deadline in owner words", + "record": {"status": "failed", "reason_code": "deadline", "outcome_axes": {"lifecycle": {"status": "failed"}, "execution": {"status": "infra_failed", "reason_code": "deadline"}, "review": {"status": "skipped", "trigger": "supervisor_terminal"}}}, + "phase": "error", + "headline": "Failed", + "acceptance_clause": "The task reached its deadline." + }, + { + "name": "a reaped task names its idle timeout in owner words", + "record": {"status": "failed", "reason_code": "idle_timeout", "outcome_axes": {"lifecycle": {"status": "failed"}, "execution": {"status": "infra_failed", "reason_code": "idle_timeout"}, "review": {"status": "skipped", "trigger": "supervisor_terminal"}}}, + "phase": "error", + "headline": "Failed", + "acceptance_clause": "The task made no progress for too long." + }, + { + "name": "a provider failure says its own sentence", + "record": {"status": "failed", "reason_code": "provider_failure", "outcome_axes": {"execution": {"status": "infra_failed", "reason_code": "provider_failure"}}}, + "phase": "error", + "headline": "Failed", + "acceptance_clause": "The model provider failed to answer, so the task could not finish." + }, + { + "name": "an empty final text says its own sentence", + "record": {"status": "failed", "reason_code": "empty_final_text", "outcome_axes": {"execution": {"status": "failed", "reason_code": "empty_final_text"}}}, + "phase": "error", + "headline": "Failed", + "acceptance_clause": "The task ended without a final answer." + }, + { + "name": "an unavailable deep self-review says its own sentence", + "record": {"status": "failed", "reason_code": "deep_self_review_unavailable", "outcome_axes": {"execution": {"status": "infra_failed", "reason_code": "deep_self_review_unavailable"}}}, + "phase": "error", + "headline": "Failed", + "acceptance_clause": "The deep self-review could not run." + }, + { + "name": "a deep self-review error says its own sentence", + "record": {"status": "failed", "reason_code": "deep_self_review_error", "outcome_axes": {"execution": {"status": "infra_failed", "reason_code": "deep_self_review_error"}}}, + "phase": "error", + "headline": "Failed", + "acceptance_clause": "The deep self-review stopped on an error." } ] }